Gemini's Multimodality

2 Jul 2025 · 44 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Google AI: Release Notes - Gemini's Multimodality

Episode Overview

In this episode of "Google AI

Release Notes," host Logan Kilpatrick speaks with Ani Baddepudi, the Gemini Model Behavior Product Lead. They delve into the multimodal capabilities of Gemini, exploring the model's foundational design, proactive AI assistants, and the vision for future developments in AI technology.

Key Themes Discussed

  • Multimodal Model Design: Understanding why Gemini was built as a multimodal model from the beginning.
  • Video Understanding: The advancements in video processing and understanding through Gemini.
  • Future of Proactive AI Assistants: How AI can evolve to become more intuitive and interactive.
  • Challenges in AI Development: Information loss in image and video processing and how it impacts model performance.
  • Collaborative Development: The teamwork and synergies behind building advanced AI capabilities.

Episode Breakdown 0:00 - Intro

  • Introduction of Ani Baddepudi and the episode's focus on Gemini's multimodal capabilities.

1:12 - Why Gemini is Natively Multimodal

  • Discussion on the importance of vision in AI and why Gemini was designed to be multimodal.
  • Vision as a core component of human experience, essential for tasks in various fields.

2:23 - The Technology Behind Multimodal Models

  • Explanation of how Gemini processes text, images, audio, and video into token representations.
  • Importance of a single trained model for efficient multimodal understanding.

5:15 - Video Understanding with Gemini 2.5

  • Enhancements in video processing capabilities, including improved robustness and understanding broad contexts.
  • The ability to generate useful outputs from video, such as converting video content into code or interactive applications.

9:25 - Deciding What to Build Next

  • Insights into prioritizing use cases based on user demands, long-term goals, and unexpected capabilities from model scaling.

13:23 - Building New Product Experiences with Multimodal AI

  • Discussion on how new experiences can emerge from improved multimodal capabilities, such as using video for educational materials.

17:15 - The Vision for Proactive Assistants

  • Envisioning a future where AI leverages visual cues to provide proactive assistance in real-time.

24:13 - Improving Video Usability with Variable FPS and Frame Tokenization

  • Exploration of challenges related to sampling rates and how they affect information retention.
  • Strategies for optimizing frame representation in video understanding.

27:35 - What’s Next for Gemini’s Multimodal Development

  • Future directions for enhancing Gemini’s capabilities, focusing on interaction design and user experience.

31:47 - Deep Dive on Gemini’s Document Understanding Capabilities

  • Discussion on how Gemini processes documents, leveraging its multimodal understanding to enhance document analysis.

37:56 - Teamwork and Collaboration Behind Gemini

  • Insights into the collaborative environment fostering innovation within the Gemini development team.

40:56 - What’s Next with Model Behavior

  • Sneak peek into upcoming developments related to model personality and user interaction.

Key Takeaways

  • Multimodal Interactions are the Future: Gemini's design reflects the need for models that can understand and respond to various forms of input simultaneously.
  • Proactivity and Intuition: The future of AI involves creating systems that can act proactively, anticipating user needs based on visual cues and context.
  • Document and Video Understanding: These areas are pivotal for unlocking new use cases, emphasizing the importance of improving video processing capabilities.
  • Collaboration is Key: Success in AI development heavily relies on effective collaboration across teams, bringing together diverse expertise to enhance model capabilities.

Final Thoughts The episode provides an insightful look into the future of AI technology, emphasizing the potential for models like Gemini to transform user interaction and the overall landscape of AI applications. With continuous advancements in multimodal capabilities, the podcast sparks excitement for the evolving role of AI in everyday life.

Watch the Episode [YouTube Link](https://www.youtube.com/watch?v=K4vXvaRV0dw)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Gemini from the beginning was built to be a multi model model. If we want to build AGI and like powerful AI systems that can perform these general human tasks, vision is a core component of the human experience. These models should be able to see and perceive the world like we do. I think vision still feels like one of the areas where there's the biggest gap between the model capability and the products that people are building. I think we're like very early in that world because it takes time for people to build intuition for what these models can do. This productivity bit is what I'm most excited about.

0:31We see a world where everything is vision, and these models can see your screen, see the world, just like we can. But they're also domain experts in every field, which I think is a future that I'm super excited by.

1:12Hey, everyone. Welcome back to Release Notes. Today, we're chatting with Ani Badapudi, who's the multimodal vision product lead for Gemini and also newly the product lead for Gemini Model Behavior. So thanks for coming on to talk all things Gemini multimodal. Great to be here. Thanks for having me. I think Gemini from the beginning was built to be a multimodal model. What does that mean in practice? Why was that the case actually going back to Gemini 1.0 when we sort of planted the flag that we were going to build a language model that was multimodal? Can you sort of share that context? Yeah, totally.

1:44So, yeah, DeepMind has been working on multimodal capabilities for a long time. And the reason for this is like if we want to build AGI and like powerful AI systems that can perform these general human tasks, vision is like a core component of the human experience. So tasks in various domains like medicine, finance and so on have like a strong visual component. So like the vision with Gemini from the Gemini 1.0 days was to have a model that can see and perceive the world like people can, which enables these models to perform these tasks in that manner. I think about this all the time because it feels like if you look at AI products in a lot of ways, they're like they're sort of screaming to be multimodal.

2:33And it's like so many of these like weird product experiences you have to build if you're like just building for this tax world. And like in a lot of ways, like the solution is the is it making a multimodal and like showing the model something visually. and we you know we've pushed on the multimodal live uh live api thread and like that's a great example of this playing out where like people can actually go and build that stuff um what is it like back to this thread of gemini being multimodal what does it actually mean from a model perspective for it to be like a quote-unquote natively multimodal model what does that mean actually yeah so we have a single model that's trained to be multimodal from the ground up um at like a high level, what this means is like text, images, video, audio, like all these modalities are turned into like a token representation and the model is trained on all of this information together.

3:23What this results in is a model that can like understand not just text, but text with images and audio and video and so on. So the abstraction that I like to think of these models at is they should be able to see and perceive the world like um like we do and that's the goal behind like training these models to be yeah like natively multimodal and like trained this way from from ground up yeah is there some uh like do you get like a compression loss effect when you do that like i think about like how i think of tokens um as just like numbers under the hood um and then i think of like an image and it's like you know the picture is worth a thousand words like how much do you but it also feels at the same time that the models are really good at multimodal so like what is there is there some magic happening there that like makes it so that you don't lose all of the like nuance of of things yeah totally so a couple of things the first is yeah i mean like losing information from images is like a big research problem um when we turn images into token representations, we like inherently lose some information from the image.

4:31This is a constant research question of like, how do we make our like image representations less lossy? The second is when we like extrapolate to video, we sample videos at one frame per second. And like during training, there are like other tricks we can use, but we like lose information because the model's not seeing the entire video stream. So there is some information loss. I think the thing that's really surprising is these models generalize pretty well so once it sees enough images and like sees videos even if they're sampled at one frame per second these capabilities generalize pretty well and it's kind of mind-blowing what these models are then able to do um so yeah these are like constant research questions that that we're working on yeah i think if um for folks listening to this if they haven't seen we put out a blog post about uh gemini 2.5 pro having state-of-the-art model performance and video understanding and like it's a very visual use case and i think it's like go read the blog post because there's you know a ton of different a ton of like really really great visual use cases and applets that we built in ai studio and others to like showcase this capability but how much is that like the the video to like image related is it like there's a lot more complexity that happens from a model perspective on making videos work well or is it like actually just like pass a bunch of images behind the scene yeah so video with the 2.5 models is like pretty mind-blowing um and i'd say there are like a few things the first is uh previous gemini models they were pretty good at video but like robustness was a bit of an issue so like one of the issues that we had for example was like if you fed in like an hour-long video to a model the model would focus in on the first five and ten minutes and then and trail off for the rest of the video.

6:19So there are some of these quality aspects that the team has worked a ton on. And these are very video specific, especially long context video, which is really the upshot of these capabilities. The second is just core vision improvements. And that generalizes to video as well. So one really cool example that we highlight in the blog post is the ability to turn videos into code. And this enables a ton of cool things. You can turn videos into animations. you can turn videos into like websites. So I fed in like a YouTube video of like a recipe and turned that into like a step-by-step recipe. A use case that we see people using a lot is videos of like lectures and like turning that into like lecture web pages and lecture notes and stuff.

7:03Makes college sound fun. Yeah, yeah. Just take a bunch of boring lectures and feed them to AI and it like makes it all interactive and like custom learning experience. Exactly, yeah. And like these things turn into like interactive apps that you can like learn with um so i think the the really cool thing about gemini 2.5 is that it unlocks video as a medium of information to do really useful things um so yeah like i see it's a bit of both i mean you and i talk a bunch about how you know there's actually so many different vision use cases like bundled under this multimodal umbrella and there's like all the ocr stuff video understanding probably 50 other things that I don't even know about or think about that often.

7:45How do you think about this from a multimodal product Gemini side? And like, what's the relationship and the interplay between all these capabilities? Is it like, are they independent of each other? Do you see gain across all of them when one gets better? What's the relationship? Yeah, two things. The first is, I think we see like the pro of having like a single single multimodal model is that we see a ton of positive capability transfers. One of the cool things about this 2.5 launch is things like video to code work really well because the 2.5 models are just like a lot stronger at code. The second is like even within vision we see a ton of like capability transfer.

8:26Like in the past a ton of these like you would have had separate models for vision capabilities like a separate OCR model, a separate detection segmentation model and so on. The cool thing now is all of this is bundled into Gemini, and that results in a ton of cool use cases. So for example, say I'm transcribing a video that requires strong OCR, but that also requires strong temporal understanding for the model to be able to understand what happens in the video and then transcribe that. One use case that we're super excited about in Gemini is using Gemini as a pair programmer. So we stream in like a video of your IDE to Gemini, like ask it questions about your code base, get answers and so on.

9:07And this is like a use case that requires strong coding capabilities, strong just like core vision, which is like spatial understanding, OCR. But then also the ability to understand a video and information in a video across time horizons, which is that like temporal reasoning piece. I love this use case. I think I would not be, my bet is, you know, if we look a year from now, like every developer product has something and actually like maybe even more generic than that, like the OS is going to have something like that, like different products are going to have like custom versions of this because it's so powerful.

9:41We've seen this already from like a customer traction piece of people building with the live API. There's just like so many cool, like not run of the mill AI applications that are being built, which is exciting to see because it feels like there's, you know, lots of people doing the same stuff. And it's awesome when people go and build new things. um i'm also super curious about just like what like how you think about what to focus on from a model capability standpoint versus where to your point about just like as the base model gets better yeah it sort of raises the you know the tide lifts all the ships or whatever the expression is um are there certain areas where you're like you know ah we don't really need to make that bet from a multi-model perspective because like it's just gonna happen naturally or like are you having to like explicitly track that to like see oh no here we need to focus on this because you know it's not getting better as the base model goes up yeah this is a great question i'd like to split it out into um like three portions yeah the first are use cases that we see are like critical today for users and customers so folks using the apis like developers um like google products that that use Gemini for multimodal vision use cases.

11:00So these are things that feel like short-term capabilities that we need to make Gemini super strong at. The second piece, which I think is very critical, are some of these long-term aspirational capabilities. So these are things that people aren't asking us for Gemini to be able to do today, but we think are very critical for building like powerful ai systems and so on yeah so i mean like one of the cool examples is um like visual reasoning like this is like and we see early signs of this with the gemini 2.5 models is this ability to reason over pixels and we have like um a bunch of toy examples like i don't know you have you have like a pinball with a bunch of surfaces and you ask gemini like um about the the path that the ball would take and like which bucket it would fall into that's like capability that's super interesting because the model isn't reasoning over like text form, but it actually needs to reason over the image and like understand like what the trajectory of the ball would be in an image.

12:02That's like a very simple toy example, but you can extrapolate a future where this becomes critical for things like robotics. Like if robots and self-driving cars have AI systems like Gemini powering like embodied reasoning, that unlocks a ton of use cases. So like these are things that like customers aren't asking us for today, but like the team is super excited by and we think are like very important for building AGI. The third, like you say, are things that we get surprised by, like things we can't plan for. And this happens just from like scaling our models up. So 2.5 was a great example of this.

12:39Like we didn't plan particularly for these models to be like this amazing at like image to code and video to code, but this turned out to be like a super strong capability with 2.5. And I think the key is like when we see early signs of this happening, figuring out what the use cases are and like where these things can be really powerful. So an example of this is like UX to code. Like I think the workflows of like designers and product managers change completely with these capabilities. Like now I can sketch a UX, feed that into Gemini and it generates like a pretty good prototype using HTML or JavaScript React for that UX interface.

13:18I think these types of things are super cool and are capabilities that we get surprised by. This makes being a PM way more fun. Through this lens of stuff that you and I spend a bunch of time talking about, one of the bits is just like, what are people actually building with multimodal? And how do we help the people who are building interesting stuff now, but also help sort of try to convince builders and startup founders, like, here's all the next things you could be building. Here's what are the good ideas. Here are the capabilities. Through that lens, like, what is some of the stuff that you're most excited about from a vision standpoint?

13:52And like, the product experience people could be building with this stuff versus what I think they're building now, which is not that much. It's just some interesting stuff, but there's lots more capability to pull out of the model still, I think. Yeah. So I think we're trying to see people do some like really cool things with vision and some of this just happens as the models get strong enough to um yeah like make these make these things work um yeah i like to think of this in like three three buckets like the first are use cases that like um existing models or systems were like able to do so these are things like traditional ocr translate um image retrieval so like google lens does this really well for things like shopping like find a similar sweater and things like classification, like help me identify this plant or this animal and so on.

14:38So I think we're seeing a lot of usage in this sphere of things because people are used to using existing vision modules for these things. And like Gemini is a single model that's able to do all these things. I think it starts to get more interesting when we look at the second and third the second uh like the set of use cases that i like to think of these of like gemini being able to do are um tasks that like a human could do or say you had like an expert in a given domain with you tasks that they could do so these are things like we like travel to london a bunch for for work and like something i've really enjoyed doing is like taking taking gemini out and like walking around the city and like asking questions about things around me yeah previously i would have had to like figure out a question in text to ask Google, get a like get a response.

15:31But now I have like a completely lossless way of asking these questions using Vision. Another cool use case that I was trying the other day was I had a Google Doc with a bunch of comments and I took a screenshot of the doc with the comments, fed that into Gemini and I was like, hey, like help me rewrite this doc while like answering these comments. Gemini did a pretty good job. I think like 50 % of the comments were addressed perfectly. 30 % were like pretty good. I needed to make minor tweaks. 20 % I had to rewrite. But I think if we extrapolate this, we see like a world where like everything is vision and these models can like see your screen, see the world, like just like we can.

16:10But they're also domain experts in like every field, which I think is like a future that I'm super excited by. The third set are use cases that I think of as like beyond human or like beyond tasks that humans could do in like a feasible amount of time. So these are things like being able to watch a six hour long video and like find specific moments where things happen. So like, I don't know, you feed in like a very long sports game and you generate a highlights reel, like this takes a lot of time for a human to do. Or generate like generating fine grained segmentation masks on an image is like something that is like hard for humans to do.

16:46Like another example is like some of the video to code things where like you have a video and you can turn that into like an interactive like learning application. These are things that would take people like a long time to do, but you can just like zero shot Gemini with these things. And I think we're like very early in that world because it takes time for people to build intuition for what these models can do. And also takes time for us to build the interfaces for people that do these things smoothly. But I'm really excited for like a world where we can like really tackle like the second and third piece.

17:15How different do you think those products tackling the second and third bit look from today's bar? Is it like you, like, you know, imagine I have a AI chat app today and I want to sort of fully embrace this world of like, you know, everything is vision through that mantra, which I love that. We should make sure. I'm going to have shirts made that say everything is vision. What's the delta? Like what do folks actually have to do if they want to sort of buy into that world? Because I do think like builders are like trying to find the edge in today's world. And I think you and I, again, I've had this conversation a bunch of time that like there's so many interesting edges for people building and vision right now just because there's not that many products in the space.

17:52Yeah, so I don't know. Do you have any advice of like how to do the exploration to find the experience? Yeah, so something that I really like doing is anthropomorphizing these models as much as I can. So thinking of these models as expert humans at a given task and treating the interface in a manner that like a human would perform a given task in or like do something in i think as an industry we like defaulted to chat as an interface primarily because like humans are very used to using chat like we like message all the time we use search for retrieval and so on um but i actually think some of these like human modes of communication and like interaction are actually far more natural um part of this like Malt's getting good enough to do these things.

18:38And I think we're like getting there. So I think the thing that I really think about is like, can we make these models feel as natural as possible? And like the world was like built for humans. So I think it makes sense to like build these machines and these systems in the same way. But yeah, there's like still some, still some work to get there. Like a vision of the future that I'm really excited by is like today, like most AI products are term-based so you like query the model or this or like even a system you get back an answer query the model again you get back an answer and you like repeat that process um in this view of the world where like everything is vision products that like i'm super excited by are um like having a world where uh the interface of interacting with ai systems is like a bi-directional like audio video interface right so this results in some cool things like your model can understand audio and video um just like a human would um it can be proactive so based on visual cues it can like suggest for like this is what i have to do this productivity bit is what i'm most excited about because i think there's so many use cases where i'm like i could ask the model like if i'm showing if the model can see my screen on my computer for something i could ask the model do it i'm like i don't really want to like i'd like to just like write like here are the things that you know you could do for me like take action on these whenever this thing happens on the screen like that would be like i get an error in my terminal like i kind of want the model to just like go and you know find a bunch of stuff and give me suggestions to fix without me having to like actually talk to the model so it is this like interesting um yeah there's like so many interesting product directions and it actually feels like there's not it's like actually not that complicated to build that which is also really interesting like it's the live api with a bunch of the stuff that we just shipped in it and like out of the box it can do exactly so the crazy thing is these models are like pretty good at these things today i think one way that i like to think of what new products could look like is like imagine you had like an expert human looking over your shoulder yeah and seeing what you can see and like helping you with things i think we like the form factor where this works today is like screen share because you can like feed your screen into a model and like you're looking at your screen to perform tasks um one example that I've used Gemini Live for is like I was cooking and previously I would have had to like follow a step-by-step recipe and try and pattern match what I'm doing to the recipe and like more often than not like it doesn't turn out exactly like what's in the recipe.

21:16Something cool that Gemini can do is it like looks at what you're doing as you're doing it and then proactively based on visual cues in the video suggest for things to do right. So I don't know I was like boiling pasta. I was like, hey, add the pasta now and like things like that. Were you just like holding up your phone to do this? Exactly. This is why we need glasses. Exactly, exactly. So some other mechanism to like, you know. Right. Yeah, yeah. So I think the core problem is developing the interfaces. Yeah. And I think we've like moved towards this world of glasses and like we're working on these things at Google as well.

21:48So like that could be one way that we do this. But I think there might be others too. Like the phone isn't great because you like lose some amount of mobility. So it doesn't feel as natural. but to me like a necklace exactly yeah so people are trying necklaces as well so i think like thinking of these models as being able to or these systems as being able to like look over your shoulder and see what you see and help you with things in the real world i think it's like very powerful the question is how do we build interfaces to actually enable this in practice another piece like uh related to this discussion that i think is very exciting is like um along with proactivity these models being able to like at a high level multitask right so we have thinking models now imagine if i could like talk to the model it sees what it like has has visions it can either see me or like sees what i see and i can think at the same time while i'm talking to the model so it's able to like take in audio and video but then also like think at the same time in like some form or um we have things like project mariner where yeah uh gemini can actuate on a screen and perform actions.

22:53I think a cool world is where while I'm talking to Gemini, it's doing things on my screen, providing me with feedback and things like that. I mean, this is the thing that I'm really excited about and what I think a lot about is how can we make these models feel as human or even beyond that, superhuman as possible? And thinking of the interfaces as being as close to that as we can make it. I love that. Ani, we talked before about how the model actually understands on the back end, like what an image looks like from a token perspective. How does that happen on the video side? What's the delta between a video understanding use case and an image understanding use case behind the scenes?

23:34Yeah. So like firstly, Gemini is like one of the only foundation models that can take in video and like state of the art at like video understanding and like reasoning over videos. For Gemini to be able to understand video, it needs to be able to understand both the audio component and the visual component. And this is like a pretty tricky problem to solve because you need these things to line up and so on. But like the way this happens today is we like interleave audio and frames that correspond to that audio at each given time chunk. And what's really remarkable is this generalizes pretty well.

24:05So the model is able to understand videos pretty well using this approach. Yeah. And yeah, like fields, feels pretty natural. I part of this FPS conversation and we've been kicking around a bunch of threads on this for a while I don't have a good intuition as to like why at the model level to have multiple like different FPSs we have to like do something like why can we not just like take just like grab more images and does it just like make the audio bit like kind of like garbled or like less there's just like less context attached to each image so you just like lose reasoning capability as it as Is it processes or like why is it hard to add an FPS functionality?

24:46Yeah, like a couple of things. So the first is like part of this is just like a function of design. So something that like seemed to work pretty well and like one FPS did a pretty good job. That being said, there are a bunch of use cases that having higher frame sampling helps a ton for. So we've seen people come to Gemini to do things like feed in your golf swing and have Gemini rate your golf swing or critique your dance moves. So for these types of things, having higher FPS is super powerful and it's something that we're working on. We actually saw that this was a real need when we saw customers start to slow down videos.

25:26So folks would want, let's say, 5 FPS. So they'd slow down their videos by 5X to be able to support this. Part of the reason for 1 FPS is just the way we designed Gemini and like our tokenization sampling at one fps supported around an hour of video so it was like a pretty clean video length to support with these models yeah that being said like we've now come up we've now uh released more efficient tokenization so these models can do up to six hours of video with uh two million context this is with lower detail too this is with yeah yeah okay yeah yeah lower detail but performance is like surprisingly very high so we just like um we like represent each frame with uh 64 tokens instead of 256 previously and what does that mean is it just like there's like a less verbose description of what's happening when you use so like in a you know we're sitting in a library if you take a picture of the behind us and you did the 64 token representation you would just like see less titles of the book or like what what actually and maybe this is like a noisy example so it's harder to make yeah so this is like a very abstract idea like what we've actually seen is so like as our like tokenization methods have become uh stronger like we need fewer tokens to like represent frames in a video so say with gemini 1.0 like back then representing an image with 64 tokens was just like a very lossy representation what we see today is like actually 64 tokens performs like remarkably well um almost actually to the same quality level as 256 and like um back to your question earlier of like like why can't we just sample it like higher frames per second um part of the reasons like the models were trained um at the one fps got it um like sampling rate and like what this results in is the model learning to line up audio and video um at this like time frame or like sampling uh sampling method um but yeah there's something that we're working on and like we have a bunch of cool things to share coming soon yeah i'm excited for the higher fps bit to land so that gemini can tell me how horrible my golf swing is yeah because i don't i don't golf enough um you and i have talked a bunch about just like and and as also as you sort of transition to doing model personality stuff like what is the future for gemini multimodal stuff start to look like obviously we've we've launched a bunch of the native output modality capabilities now with audio with image uh you know some future version of the world will hopefully have video as well which would be awesome and a single model um but on the other like non-output mode which it feels like the maybe the frontier is on output modalities now and less on input modalities but do you see there's like still a bunch of places to to hill climb from a quality or capability standpoint from a multi-modal input perspective yeah there's a ton so um i think the first thing is like we want to get to a world where these models are amazing at uh multi-modal in multi-modal out so it can take in any modality generate any modality um and some of the generation stuff is super exciting on the vision side there's still a ton to do so one of the things that I'm really excited about is bringing some of these capabilities together to form a more cohesive system so for example Gemini is amazing at spatial understanding so it can generate 2d bounding boxes 3d bounding boxes point coordinates segmentation masks and so on this is really cool just for folks who haven't tried this before if you were to like screenshot if you're watching this video recording if you were to screenshot and look at behind us you can be like go to ai studio right now which also has the native image generation editing and i think it benefits from the spatial understanding you could say like move the couch and ani against the wall and it'll like it'll understand what that actually which is just have to make the plug for how cool it is to like actually you say spatial understanding and people are like uh that's anything but it's actually like so cool and yeah yeah i mean it's it's it's like super amazing like the cool thing about spatial understanding is models in the past were able to do detection um the cool thing about gemini being able to do detection is you have this reasoning backbone and world knowledge as well so some of the cool things gemini can do this is a very simple example but you ask gemini to detect the person that's like the furthest to the left in this in this image gemini is able to do that because it's able to reason over the image understand like that an object is uh or like the relative positioning of an object and then generate bounding box something that i tried before was i uh took an image of like our uh like the fridge in our micro kitchen and i was like which drink has the fewest calories and a generated bounding box around the bottle of water so these things are like super cool um we're still in like very early stages of these capabilities um so i think there's like a lot more that we can do there they still kind of feel like um they like still feel like a toy uh so yeah there's like a ton we can do there That being said, like a lot of these niche capabilities have very specific groups of power users.

30:35So on the spatial stuff, we have folks building models for robotics that are like using these models a ton. Yeah. Because spatial understanding is like a core building block for like embodied reasoning and perception for robots. So like what I'm excited about is like seeing some of these capabilities come together and like both on the understanding side, but also generation. I mean, like one of the cool things that you get with spatial as well is like some notion of thinking, right? If a model can like point to objects in an image, generate bounding boxes for things in an image, that improves the model to be able to reason and think over like visual like data formats.

31:15And that's something that I also think is like super cool. And yeah, like we have like tons of folks working on. Yeah, I think vision still feels like, to opine on a point that you've already made, like it feels like one of the areas where there's the biggest gap between the model capability and the products that people are building. Like it just feels like there's so many interesting things to be built and like there's just not that much stuff being built, which gets me kind of excited for, again, for people who are in the position of like going in, building companies around this stuff. one of the modalities that we talk a lot about is this like document understanding maybe not modality but use case um and we've seen a ton like there's a ton of positive traction around gemini and document understanding and ocr we want to talk about that and like what it takes for the models to be good at that like what that use case looks like why people are so excited about it yeah So, like, a ton of information is stored in documents.

32:14So it's, like, very clear that, like, documents is, like, a powerful, like, medium of information that, like, Gemini should be really good at analyzing and reasoning over. I think the reason why we see a ton of demand for documents as a vision use case is there were existing, like, vision models that were able to do OCR and translate and things like that. the cool thing that you get with Gemini though is you get these capabilities but you get them with like the reasoning backbone that like Gemini offers and some of the things that I'm really excited about with documents is being able to feed in a ton of documents as as context to the model to perform like a fairly complex like multi-step task that's something that like existing models weren't able to do pre-Gemini.

33:00And I think given that so much information at like, like so much personal information, information at companies and things like that are like stored in documents. That's like a very powerful visual use case. The other reason why using Gemini for like documents is interesting is because in the past with documents, the way these like workflows worked was users would like OCR a document and then feed that as text. into like AI model, like AI systems or like OCR these things and then store information in that medium, right? So like search over that information for retrieval and things like that. I think the cool thing about using vision to understand documents is like now you have like a really powerful system that can see a document just like a human can.

33:49And some of the cool things here is like documents are often not just plain text. like they have interesting formats they contain charts images diagrams things like that in the past these were very hard to like transcribe and use for like even use cases like search but but also some of these more complex tasks um the cool thing about gemini is you can just feed all this in gemini like reads all these documents like a human does and then is able to like do a ton of cool things i mean like something that i was trying the other day was i fed in earnings reports from companies over the last like 10 quarters with like a million token contacts uh which is like tens of thousands of pages with uh yeah like two million tokens and like got it to like do a bunch of analysis on these companies for me the cool thing here is like gem and like to be able to do this effectively you need to be able to like read very long intricate tables in these documents which um like previous uh like ocr modules like weren't weren't as great at so i think given that like documents are like such a massive like store of information it's uh uh it like yeah makes a ton of sense that like people are using gemini for this and it's something that like we care a lot about as well it also feels uh like very unique uniquely google in the sense that like you know through the I think the official Google mission is organize the world's information and make it universally accessible I think there's like so much of like data that even I think it's just thinking about myself I'm like I have a drawer somewhere in my in my apartment of a bunch of you know hard copies of paper that I'm never going to look at again and like it's not universally accessible in that medium even for the data that I have today so yes yeah I mean like that's an amazing point I think two things there um this is something that I'm super excited about on Vision is I think it like unlocks Vision as like a store of information and makes like visual information so much more accessible and useful to Google's mission.

35:47So with like documents, we see this like Gemini is amazing at like what we call layout preserving transcription. So it can transcribe a document and preserve layout style structure. The other really cool way that like vision makes information more accessible is video actually so something that like we see a ton of people do is like take videos of things around them feed these videos into gemini and then use that to catalog information um so yeah like a ton of people take videos of their bookshelf and like libraries and then like catalog that information yeah i do this in ai studio all the time where i'll take like long videos of like whatever topic someone talked about in a podcast that I was a part of or I listened to or whatever it was and like have them pull out you know interesting clips and stuff like that it's super super powerful to not have to especially with content where I don't want to hear myself have to talk you let the let the model take care of all the yeah do the work exactly it also makes like tasks far more efficient something that I did when I was yeah like at our office in London was like we have we have like a very nice library there and i took a video of like all the books in this library and asked gemini to catalog these books by genre by author and because gemini has this like world knowledge reasoning backbone but then also these visual capabilities and like a single model i was able to do this super well i did the same thing like at our micro kitchens like what's trying to hear after this yeah it was something that i've been trying since the gemini 1.0 days is like cataloging snacks in our like mks yeah and uh yeah these models are now like pretty much perfect at these use cases i think this like goes to some of the questions around user interfaces as well like this is something that doesn't feel like a natural use case because it hasn't been possible before like people are very used to using vision for like traditional ocr translate um classification these types of things um but what like gemini's multimodal capabilities gives us is the ability to like do so much more that we previously wouldn't have thought was possible um and i think like that'll take some time to to play out as well anny one of the sort of consistent threads of and i don't even maybe this is intentional maybe it's not i'd have to reflect on this of a lot of the conversations we have with people through this podcast is just about like how much the like team you know sort of team gemini like everyone is there's like deep collaboration there's all the different modalities getting better makes other people's modalities better like i'm curious like on the multimodal side like as someone who has a little bit less context of like what the team looks like and what the structure is like how um yeah how do you think about that like what's happening who are the research people all that stuff yeah there's a there's a ton of people but yeah yeah so there are like a ton of people firstly like i'm just a spokesperson for like a massive research team and yeah like the Gemini multimodal team is like the most amazing team um and we've like grown a ton since the Gemini 1.0 days which which also shows like how strong these capabilities are getting um and yeah I think the thing that's like really awesome is multimodal has like so many of these capabilities and like a ton of things that that uh uh that we've spoken about and like what you need to make this happen like if it's like what's a very hard problem is you need to like bring these capabilities together into a single model and make sure that like each capability performs super well and we have JB who like leads our multimodal team he's a rock star he's been working on vision from before Gemini like back in the Flamingo days and we have work stream leads image video and like all these things spatial and I think what's been like really remarkable it's like how like all of this has come together into a single model um that is like super strong at these uh yeah like at these multimodal capabilities yeah it's very interesting to reflect on like i don't think that's like the default outcome i think this is like i don't maybe it's maybe it's just we have great people who work really well together but like i think this it's it's hard to make that collaboration happen and it's so cool to see like time and time again like it actually plays out and it works and like we get the great results from a model perspective you know from everyone coming together so it's it's awesome to see yeah Yeah.

40:02And I think the other thing that's like really awesome is I think the team really thinks about things really deeply about a like how developers and consumers will use these vision capabilities. And I think we like really try and build strong intuition for this and bring that into our model. So we have this like very close product model feedback loop. And the second is we spend a lot of time thinking about and like chatting with each other about how people will use these capabilities in the future. So if we extrapolate to like these capabilities becoming much stronger and like coming together in like a cohesive way, like what are the ways that people are going to interact with these models a year from now, two years from now, five years from now?

40:47and a lot of the capabilities that go in today are like building blocks towards this vision that our team has which I also think is like super powerful yeah it's been awesome to see all the progress on multimodal for the last year it's been awesome to collaborate with you and JB and the multimodal team so I'm super appreciative of all the hard work that y 'all have done even outside the model stuff like you know if you if folks have complaints about the multimodal AI docs go to Ani and he'll help make them better but you're transitioning now to go start working on model behavior stuff. So we won't do a deep dive on model behavior, but just to sort of plant the seed, what is, I think this is definitely somewhat emergent right now.

Read the full transcript

41:29So like, what will you be thinking about next? Yeah. So I think related to some of the things we've spoken about, something that I think is like a very important problem is having these models feel like they're natural to interact with. I go back to this, like the like world today where we have these like very turn-based systems it feels kind of unnatural it feels a bit dated um and something that i'm passionate about is like building ai systems that feel likable that you can like interact naturally with so i mean by going into more detail on the model behavior stuff how this translate is translates is like giving the model skills like empathy and being able to understand the user understand implied intent giving the model like a personality while it's striking the balance of like um like all of these like uh yeah like amazing raw capabilities that that Gemini has I think the other piece to some of this is um a lot of the AI use cases today like these models just give you a ton of text something that I've been thinking a ton about is like are there interesting visual formats that we could use to be able to like communicate information and like a more information dense or like high calorie manner is uh like is uh yeah like the way that we like to think about these things and i think it's like a very critical problem for uh yeah like making making gemini uh like a nice model to to talk to and interact with i'm excited for this i think my my seed to play with you is that um people really like the way that the sort of notebook lm audio overview personality is and like the way that it sort of engages from a conversational set way is like super super relatable and folks really like it.

43:15So I think there's some interesting thread to pull on that, which I'm, yeah, we'll have to catch up more about this sometime in the future and see if there's any interesting outcomes. But to give you credit for folks who have been watching, you and the multimodal team have been, I think, like one of, on the AI Studio, Gemini API side, some of our strongest collaborators. And I think it's been a ton of fun to work with you and the team to like do that. Like, it feels like this very unique, like research to product acceleration story, which is like, you know, you all care about what the API looks like and what the capabilities are and all this stuff.

43:48So I'm appreciative of you and you and JB and everyone else for pushing so hard to make all that happen. Likewise. Yeah. Yeah. Yeah. Thanks so much for bringing these capabilities to life through a pretty nice video. Thanks for taking the time to sit down and chat everything multimodal. And thanks everyone for listening. And we'll see you in the next episode.

44:12you

From the publisher

Ani Baddepudi, Gemini Model Behavior Product Lead, joins host Logan Kilpatrick for a deep dive into Gemini's multimodal capabilities. Their conversation explores why Gemini was built as a natively multimodal model from day one, the future of proactive AI assistants, and how we are moving towards a world where "everything is vision." Learn about the differences between video and image understanding and token representations, higher FPS video sampling, and more.

 

Chapters:

0:00 - Intro
1:12 - Why Gemini is natively multimodal
2:23 - The technology behind multimodal models
5:15 - Video understanding with Gemini 2.5
9:25 - Deciding what to build next
13:23 - Building new product experiences with multimodal AI
17:15 - The vision for proactive assistants
24:13 - Improving video usability with variable FPS and frame tokenization
27:35 - What’s next for Gemini’s multimodal development
31:47 - Deep dive on Gemini’s document understanding capabilities
37:56 - The teamwork and collaboration behind Gemini
40:56 - What’s next with model behavior


Watch on YouTube: https://www.youtube.com/watch?v=K4vXvaRV0dw

More from Google AI: Release Notes

All 30 episodes
Gemini's MultimodalityGoogle AI: Release Notes · 44 min
Listen in VO