Behind the scenes of Google's state-of-the-art "nano-banana" image model

27 Aug 2025 · 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Google AI - Release Notes Episode on "Nano-Banana" Image Model

Episode Overview

In this episode of *Google AI

Release Notes*, host Logan Kilpatrick engages with key members of the Gemini team to discuss the groundbreaking image model, Gemini 2.5 Flash. The conversation delves into the advanced capabilities of the model, its application in image generation and editing, and the iterative processes behind its development.

Key Participants

  • Logan Kilpatrick: Host
  • Nicole Brichtova: Product Lead
  • Kaushik Shivakumar: Research Lead
  • Mostafa Dehghani: Research Lead
  • Robert Riachi: Product Lead

Episode Highlights

  1. Introduction to Gemini 2.5 Flash
  2. Quality Leap: The Gemini 2.5 Flash model is described as a significant advancement in image generation and editing.
  3. Mono-Banana Codename: The model was initially referred to as "nano-banana," sparking curiosity and speculation in the community.
  1. Demonstration of Image Editing
  2. Image Generation Demo: A live demonstration showcased the model’s ability to generate images based on vague prompts, maintaining character consistency and context.
  3. Natural Interaction: The model allows users to interact using natural language instead of complex prompts, enhancing the user experience.
  1. Text Rendering Capabilities
  2. Rendering Challenges: While the model shows promise in text rendering, there are still areas for improvement. The team is actively working on addressing these gaps.
  1. Evaluation and Feedback Mechanisms
  2. Human Preference Evaluations: The team discusses the challenges of using human preference evaluations for model training, highlighting the subjective nature of image quality.
  3. New Metrics: The introduction of new metrics, such as text rendering, has proven beneficial in evaluating model performance during training.
  1. Interleaved Generation and Context-Aware Edits
  2. Iterative Editing: The model supports multi-turn interactions where users can make incremental changes to images over time, facilitating a more intuitive creative process.
  3. Pixel-Perfect Editing: Emphasis on maintaining consistency across edits, ensuring that changes do not disrupt the integrity of the original image.
  1. Comparison with Specialized Models
  2. Imagine vs. Gemini: Discussion on the differences between the specialized Imagine model for text-to-image generation and the Gemini model, which incorporates multimodal capabilities and more interactive editing.
  1. Future Directions for Image Generation
  2. Smartness in AI: The team aims to enhance the model's ability to surprise users by producing outputs that exceed their expectations, fostering a sense of interaction with a “smarter” system.
  3. Upcoming Features: Insights into future enhancements, including improved factual accuracy and more sophisticated visual outputs for practical applications (e.g., presentations).

Key Takeaways

  • User-Centric Development: Continuous feedback from users is crucial in refining the model's capabilities and addressing failure modes identified in earlier versions.
  • Emphasis on Multi-modal Understanding: Gemini aims to integrate various forms of media (text, images, audio) to enrich the model's performance and versatility.
  • Future Potential: The conversation concludes with an optimistic outlook on the advancements in AI technology, with team members expressing excitement about continuous improvements and upcoming features.

Closing Remarks The episode encapsulates a vibrant discussion on the evolving landscape of AI image generation, highlighting the collaborative efforts of the team to push the boundaries of technology and improve user experience. Listeners are encouraged to tune in for future updates and developments from the Google AI team.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today we're talking about native image generation with the team behind the new model that we're releasing. It's a giant quality leap. The model's state of the art, and we're really excited about both the generation and editing capabilities. You can ask to, for example, render the character from different angles, and it will look like the exact same character. When users interact with this, not only they're impressed by the quality of images, but they feel like, wow, this is smart. And can kind of have a fun conversation with the model over multiple turns. So I think this iterative process of creating is kind of like the magic behind it.

0:30And I think we're just scratching the surface on what these models can do. Hey, everyone. Welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today, we're joined by Kaushik, Robert, Nicole, Mostafa. These are the folks who are doing research and product for our Gemini Native Image Generation model, which we're here to talk about today, which I'm super excited about. So, Nicole, you want to kick us off? Why? What's the good news? I'm excited to hear about the release. Yeah, we're releasing an update to our image generation and editing capabilities in Gemini and Tutor 5 Flash, and it's a giant quality leap the model state of the art and we're really excited about both the generation and editing capabilities um and why don't i just show you what the model does because i think that's the best way to kind of get that across i'm excited and i played around with it like once but i i have not done as much playing around as y 'all have so i'm excited to see some some examples um great i'm i'm gonna take a picture of you okay um and let's just start with let's say zoom out and show him wearing a giant banana costume and keep his face visible because we want to make sure you know it looks like you all right it's going to take a couple of seconds to generate but it's still it's still pretty snappy which i think you remember from our last release like it was a pretty fast model um this was one of my favorite things i feel like this like pace of of editing uh makes these models a ton of fun to play with can you make it slightly bigger for me can you You can go full screen, I think.

2:01Click on this. Click on that. Let me just click on this. So there we go. This is Logan. This is still your face. And what's awesome about this model is that this still looks like you, right? Like this is you, but it's actually like you're wearing a giant banana costume. And now there's like a nice background of you walking through a city. That's so interesting because this picture is in Chicago. And that actually is pretty much what that street looks like. World knowledge coming on this model. and now let's keep going and let's say make it nano. What does that mean? What does make it nano mean?

2:39So let's see what the model does. When we first released it on LM Arena, we gave it the codename nano banana. Yeah. And people started speculating that it's an updated model from us and it is a model updated model from us. And there you go. now the model takes you and creates this like cute nano version of you wearing a giant banana costume i love that that's awesome and the awesome thing here is obviously like this was a very vague prompt right like you were like what does this mean i actually did not go with that mat um but then the model's creative enough to kind of interpret it and then like you know create a scene of this where like that that it fulfills your prompt and it still makes sense in the context and it keeps all the rest of the scene relevant um and this is really exciting because um it's the first time I think they were seeing kind of LLMs be really able to like keep the scene consistent across these multiple edits and have users use really natural language to interact with the model, right?

3:34I don't have to put in a super long prompt, like I'm just giving it very natural language instructions and can kind of have a fun conversation with the model over multiple turns. So that's super exciting. I love that. How good is it at like text rendering stuff, which is one of the use cases I care the most about? Do you want me to put something on this picture? Why Why don't you give me a prompt?

4:00Gemini Nano. That's the only Nano thing that comes to mind.

4:08I feel like this is the use case that I'm always trying to do is like announcement tweets with Bill Boris with text on them is what I love is my use case. All right, let's go. There you go. Nice. Cool. And so this is a relatively simple text, right? Yes. It's a pretty small number of letters, like easy words. And that worked really well. We do have some gaps in text rendering that we call out in the release. And we're working really hard on it. Folks on the team, Kaoshi maybe can talk about that, are working on making text rendering even better in our next model. I love it. Any other part of the, any other examples you want to show?

4:48Or like, is there any other like metric story around this launch? I know one of the challenges, and I'm curious actually how you all think about this is like the eval story is like a lot of human preference stuff is like what you're measuring. It's like hard to have like a, I think there's probably some things that you could have a source of truth on, but I'm curious how, yeah, how you all think about that for this release, but also just in general as we're training these models. I think generally with like, you know, multimodal stuff like image and video, it's like very hard to kind of like hill climb and you know um the kind of like historic approach has been to use like a bunch of human preference and kind of like hill climb bats um obviously like images are like super subjective so you're kind of like um getting like signal from a large group of people and it takes time right it's not necessarily like the fastest metric and it takes like um like real like hours to kind of get anything back from it so generally like we've been working really hard to come up with like other metrics that we can like hill climb as like me train um and I think like text rendering has been a really interesting story because like I think uh Kaushik has been you know talking about it for a long time one of the like biggest advocates of it uh and we were kind of like brushing him off for a long time about how like you know this guy's a little crazy like he's really obsessed with text rendering um but eventually it kind of became like one of the staple things we looked at and you can kind of think about it like um when the model learns how to do this like structure for text it's also kind of able to learn other structure in an image as well and like in an image you have these like um different like frequencies and you can have like structure which you can think of and but you can also have like texture and stuff like this so it really gives you signal into like how good the model is at generating the structure of the scene um and i'll let kashik talk a bit more about it because he's like the the main guy yeah i'm also curious like what the the initial conviction was is it just like as you were doing like a bunch of research experiments like it became clear that this was the case or yeah i'm curious to double clear on it yeah i think it started from a place of figuring out what these models were bad at so we in order to kind of improve any model you need a signal for uh what is not working well and then you try a bunch of ideas whether it's related to the model architecture data or other things and once you have that clear signal you you can definitely make good progress on it and i think if we go back a few years there were pretty much no models that were doing a decent job and even prompts that were on the order of uh short short lines like this gemini nano prompt here for example so as we uh spent more time looking into this metric and always tracking it right whatever experiment we run now if we track this metric we can make sure that we don't regress on it and just by virtue of having that as a signal we might even find that changes that we didn't expect to make a difference here actually do make a difference and then we can make sure we continue improving that metric over time yeah and like Robert said it's a great way to just measure overall image quality in in the absence of other other metrics for image quality that don't saturate very quickly right i think humans i was actually a little bit skeptical of the human rater approach to doing evals for image generation but i think what at least i've realized over time is when you have enough humans looking at enough prompts across a variety of categories you actually do get quite a bit of good signal but obviously this is expensive you don't want to always be asking a bunch of humans to grade images so looking at this text rendering metric for example while a model is training gives you great signal as to whether it's performing like you expect that's super interesting i'm curious about this interplay between the native image generation capability native image understanding capability we did an episode with ani and that team has obviously been pushing super hard on like Gemini has state-of-the-art image understanding is it like a reasonable mental model as our models get better at understanding images uh there's like some of that capability is actually transferable to generation as well and vice versa like is that i feel like yeah and so basically to hope is that we end up with uh with native image generation or native uh native multimodal understanding and generation and learning all these modalities and different capabilities at the same model within the same training run is that you want to end up having positive transfer across these different axes, right?

9:30And it's not only for understanding and generation for a single modality, but also it's about can we learn something about the word that is like from images or videos or audio that is going to be helping us from on the text understanding or text generation. So for sure, image understanding and image generation are like sisters. So like we definitely see still they're like going hand in hand in interleaved generation, for example. But also the ultimate goal is to see, like, let me just give you one example. So for example, language, we have this like phenomenon that we call it like reporting biases.

10:09And what it means that you go to your friend's place and when you come back, you never talk about their normal sofa in a conversation, right? But if you show someone an image of that room, it's there. So if you want to learn about a lot of things in the whole world, like images and videos, they have that information there without explicit requests for those information. So what I want to say is that eventually with text, or with other modalities, you can learn a lot about different things, but it might take more tokens. So visual signals are definitely a good shortcut for learning about the board.

10:47And back to the understanding and generation question, as I said, these two hands-in-hands and coming into the interleaved generation, you can see that there's actually a huge help from understanding to better generation and the other way around. So mid-generation can help, like you draw something on the board to solve a problem. So maybe you can better understand a problem that is given to you as a visual image. So maybe we can actually show some interleaved generation that is kind of related to understanding and generation going hands-in-hand with text as well. Let me transform this subject into a 1980s American glamour mall shot in five different ways.

11:38All right, fingers crossed, this works.

11:43Okay, this looks promising. And this takes obviously a little bit longer, right? Because we're trying to generate multiple images and then we're also trying to generate the text that would describe what's in those images. And one of the things that you'll notice about native image generation is that it's generating these images one after another. So the model may choose to look at a previous image and either try to generate something very different from it or try to generate a minor modification of it. It at least has that context of what is already generated. So that's what we mean by native image generation models.

12:14They have access to multimodal context and then they generate an image. Yeah, that's interesting. And my mental model had always been that it was like just, I guess maybe that doesn't even make sense, but it would have just been like four independent forward passes or something like that. But this is actually like all in a single, it's all in the context of the model, all in the context of the model. That's super interesting. And what's nice is then the style is kind of similar, right? It's also the models doing this funny thing where it has you twice in every single one. Interesting. Could we make some of these full screen?

12:44I'm going to make some of these. So this is Arcade King Logan.

12:50If we scroll, this is Rad Dude. And see, like, none of these descriptions that go with the images were something that we came up with. The prompt was just like, you as a 1980s American glamour mall shop. Mall Rat. You should consider some of these outfits. And the fourth option, chill bro. See, and like you have a different outfit in all of them. They all look like you. The fact that you're there twice is probably a little bit of a failure mode. But it's really cool to be able to see the model kind of come up with these five separate ideas. Give them different names, give you different outfits, right?

13:27And like keep the character consistent. And this is not just useful for character building, but this is also useful if you have a picture of your room. Yeah. And you can say like, hey, help me decorate this in five different ways, right? And maybe you can go from like really creative to maybe something more conservative that's a little bit more incremental to what you're doing. And we've seen a lot of people on the team are using it to like redesign their gardens and homes. And it's been really cool to see that and more kind of like practical application, which is us making fun of 80s Logan. I vibe coded for my girlfriend in AI studio, actually, an app to visualize her office with every different color of blinds or of curtains.

14:05and she was like, I don't know what curtain color is going to fit this vibe. So it literally just, this is with 2.0, and I'll have to retry it with 2.5 to check all the different vibes. It actually worked really well. It was like a very helpful and like doesn't, sometimes with 2.0, and actually this will be a good thing to retest, sometimes with 2.0 it would like change the bed or like change, like other artifacts would change, not just the curtain. So it was interesting to see that use case, one of my favorites. You should give it a try. The model does a pretty good job keeping the rest of the scene consistent.

14:36And we call this kind of pixel perfect editing. And that's really important, right? Because sometimes you want to just edit that one thing in your image, but you actually want everything else to stay the same. Again, if you're doing character building, you just want to turn the character's head. But like everything that they're wearing is to be the same across the scenes. And the model is really good at that. It will not always 100 % work, but we're really excited about how far it's come. Robert, you're inside something. Yeah, I was going to say, like, I think one really cool thing is like just how fast it is still, right?

15:02like you know well how long was this whole generation let's give this a uh this is 13 seconds wow so i think i think each image was each each image was 13 seconds right ish ish ish and so okay this is the cumulative now yeah yeah yeah me not the ai studio yeah yeah so so i think like the pool thing is like even when you know 2.0 came out i was using it for very similar things like i had a bookshelf i had all the stuff on the ground i'm like decorate this like what configuration of these items should be placed on my bookshelf and you know my girlfriend might not have agreed with the output so like sometimes we want to like iterate on that and so like rerunning it really quickly and iterating so even if sometimes like it kind of like fails you just tweak the prompt rerun it and you get something like really good afterwards so i think this like iterative process of creating is kind of like um the magic behind it and any difference in how folks who had tried 2.0 as an example um and like one of the examples for me using 2.0 was like wanting to be like do um only single edits like one at a time like if you had said like if you had asked it to like change six different things like the model would sometimes not do a great job of that any of those like is that still something that we you should still do like those like type of targeted edits with this model or any other just like general like usability or like things that folks should know as they're as they're playing around with the model this is something that i wanted to mention basically so uh one of the magics of interleave generation is that it offers you to do a new paradigm for image generation, right?

16:29Like, so if you have a very complex prompt, you know, you're talking about six different edits. What if I go with like 50 different edits, right? So now that the model has a really good mechanism to grab information from the context, like pixel perfect and use it in the next turn, what you can do is you can ask the model to break down the complex prompt, either it is editing or for image generation into multiple steps and do edits like one by one over different steps. So for the first one, you do this, like edits, like these five different things. And then for the next one, the next five, and so on and so forth.

17:05So it's like very similar to the test on compute that we have on the language side, right? So you spend more flops and you let the model to bring basically this thinking in the kind of like the pixel space, plus breaking it down into smaller pieces that you can like really nail down that specific stage, but like accumulate it, you can do whatever complex task you want, right? So I think like, again, this is the magic of interleave generation that we can think about, you know, incremental generation of like really complex images as opposed to traditional way of doing it, which was like really pushing hard for getting the best image in one shot, right?

17:42Like at the end of the day, there's a capacity that it can push the middle. You know, at some point you realize that, okay, you know, 100 details, we cannot do that. But when you have this one interleaved generation breaking the steps, you can always go for any capacity and any complexity that you want to generate. One of the things that's always top of mind for me, especially as like you're, Nicole, you're also the PM for our Imagine models. How should people think about developers or just like people who have knowledge of all the models, like Imagine versus this like native capability that we have.

18:18Yeah. And you know this, but our goal is to always like build one model with Gemini, right? So ultimately our goal is to always like bring all the modalities into Gemini so that we can benefit from all the knowledge transfer that Mustafa was talking about and ultimately build towards AGI, right? On the way there, there's a lot of usefulness out of having specialized models that are just very, very good at a specific thing that you need them to do. And Imagine is an amazing model for text-to-image generation, right? And we have a lot of different Imagine variants that also do image editing, and those are available in Vertex.

18:49And they're just optimized for that specific task, right? So if you just want text-to-image, and you want just one image out of that model, and you want really amazing visual quality, and you also want that to be really cost-effective and kind of snappy in generation time, Imagine is the place to go, right? If you want some of these kind of more complex workflows where you want to generate with the model, but then you also want to edit in that same workflow and you want to do it across multiple turns or you want to do some of this like ideation like we were doing with the model of like you know what what design ideas could you help me come up with for you know my room or this library um then Gemini is the place to go right so it really is kind of that more multimodal like creative partner where it can output images it can output text um you can be kind of less precise with the instructions that you give to Gemini um because like when we you know at the beginning we said, like make a nano, because it has that kind of world understanding and will just more creatively interpret your instructions.

19:45But imagine it's still a great family of models for developers to go to if they want like a super optimized model for that specific task. Yeah, one of the examples I was trying today, and I'm curious what your take is on which model or like if the native image generation model fixes this problem was I was saying like, generate this image and like make the this is my my dumb billboard use case. I was like, make the billboard use case. I need billboards. make the billboard the style of some company that I mentioned. Is that something that like Native Image Generation benefits from because it's like a little bit better at this world knowledge piece relative to like imagine being like really good at if you give it a good prompt, but like less good at the like in understanding my implying my prompts.

20:30You're actually intent behind the problem. Yeah. So I think that's part of it. The other part is with Native Image Generation, if you just want to grab that style reference that you have from that you know other company that you were trying to emulate the style of you can also insert that into the model and use that as a reference right so the fact that you can then also input an image as a reference like helps with that prompt and that is just easier to do in Gemini natively than it is in Imagine um so I do you should try it yeah you should let us know we should add this to our evals I'll let you know whether around the billboard use i'll make a billboard eval we'll have a logan email i love that one um back to this thread of like the progress from 2.0 one of the most fun things was when that model launched people were sending us tons of feedback about the experience in ai studio and then ultimately the gemini app um just like general failure modes for the model and all that stuff uh i made my only contribution to the original launch which was adding that hot tag in ai studio We're bringing the hot tag back for this model, actually, and it's going to go away on the other model.

21:34How, like, can we talk about that story of just, like, the progress and, like, the failure modes that we did get a ton of feedback on of, like, things that didn't work well for 2.0 that now hopefully work well for 2.5? Yeah, I mean, we literally sat on like X or Twitter and like went through a bunch of feedback. And literally, I remember like Kaushik and I and some of the other team, like gathering all the failure cases and making evals out of that. So we have like a benchmark that we take from like real user feedback just from Twitter. And it's just people adding us and saying like, hey, this didn't work.

22:10And like for every model we make in the future, we kind of just like append to that. so that we know for example like when we released 2.0 one of the failure cases sometimes we would see is like if you edit it would add your edit but it wouldn't necessarily be consistent with like the rest of the image right so that was like one of the things that was like in that and then we hope climbed and then there's plenty more so kind of um we're always just like gathering that feedback yeah send us send us the examples that uh that don't work well any any ones for you all that like particularly stand out of things that like just did not work before that now is like a slam dunk i don't know if there's anything top of mind you you all play with i think that like the team plays with this model i think so much in the i assume in the process as we're actually building in and bringing into life i don't know if there's any like go-to use cases for you all to test and like is this actually a good model yeah i think one thing i've noticed specifically while playing with the 2.5 model is that in the 2.0 model actually one of the things that we thought was going to be hard was consistency from image to image, but specifically the cases where you have an object or say like a character that you're building and you want that character to remain consistent across images.

23:23And if you actually leave the character in the same place that it was in the input image, it turns out that this is actually quite easy. And the 2.0 model could do this really well. It could, for example, add a hat, change the expression and stuff like that while kind of keeping the pose and overall structure of the scene the same what the 2.5 model adds on top of what these capabilities look like in 2.0 is that you can ask to for example render the character from different angles and it will look like the exact same character but from for example the side or you could take a piece of furniture and place it into a completely different context, reorient it and create a whole scene.

24:06But that piece of furniture, it would remain faithful to the original that you uploaded while transforming it in very substantial ways, not just taking the input image and pasting kind of those pixels into the output image. I love that. One of the reactions that I had about some of the 2.0 stuff was sometimes the images would almost look like as you would do, like add something, like I picture my face and add a goofy mustache or a hat or something it almost looked like it was like superimposed or was like kind of like photoshopped onto it is that something that is also like similar to this like character it's like it seems tangential to this character consistency but it feels like it's a similar-ish problem where it's just like taking pixels from memory and like putting them into the image almost versus the pixel transfer i'm curious if that's like a capability that's improved yeah and i actually I think that comes down a lot to the actual teams working on this model.

25:01The previous model, actually, we were kind of of the mindset that, okay, it did the edit, that's it, like it was successful. But when we started working more and more closely with the Imagine teams, they would look at the same exact edit that we were looking at from the Gemini side, and they'd say, this is terrible. Why would you ever want the model to do something like this, right? So this is one example we're blending the perspectives from both teams uh so on the gemini side the instruction following world knowledge all of these things and then on the imaginative side making the images actually look natural aesthetically pleasing and genuinely useful uh so i think it takes both of these and having these teams work together on this that led to 2.5 being much better at the stuff you're describing i love it um yeah and just on that point we actually have folks on the team who mostly come from the Imagine team who have like a really honed aesthetic taste.

25:52And so a lot of the times when we do evals, they will actually just look at like hundreds and thousands of images and be like, no, this model is better than this other model. And a lot of other people on the team will kind of look at it and be like, okay, you know, like we like, like you kind of have to hone that, I think, like sensibility over a couple of years. And I've gotten a lot better at it over the years. But there's definitely people on the team who are like amazing at it. And we always code to them and we try to pick between models can you train uh auto raters on people's like personal we haven't been able to do it yet and that's a fun side project i'm very excited for for as germany gets better kind of an understanding to have like an aesthetic operator based on you know one of the folks on the team who's really amazing at this just put that person to to provide training signal for us yes yes this is we'll we'll take that as a side project after this.

26:41I love that. Lots of progress on 2.5. And obviously, I think folks are going to be super excited to try out the model and all that stuff. What comes next? We've made a great model. I'm sure we have more stuff cooking in the pipeline, but I don't know how much we want to say about the future direction and what other capabilities hopefully will land in the future. So when it comes to image generation, I think we do care about the visual quality. but I think one thing that is again like new and and we want this with like unified omni model is smartness you know like you want your image generation model to feel smart you know when users interact with this not only they're impressed by the quality of images but they feel like wow this is smart you know like one example that I have in mind and I'm looking forward to to to see this happening and it's a bit controversial because I cannot even define it well is when I ask the model to do something, it doesn't follow my instruction, but it does something that at the end of the generation, I say, I'm glad that it didn't follow my instruction.

27:43It's even better than what I actually described. So it has this kind of edge to it that... Is that like the... You think the model is intentionally doing this or it's like it's kind of an unintended accident? Is that what you're trying to say? No, no, it's not just that, but it's basically sometimes underspecified or sometimes you think wrong about like some like something that is a reality but you know outside world with like the knowledge of Gemini um it's different from your perspective right and uh I think um again like it's not intentional or or what just like happens organically and uh I think again like you just feel that I'm interacting with a system that is like smarter than me right and when I'm asking for some images, I don't mind if it goes off the rail with my prompt and generates something that is different from what I asked because it's most of the time better than what I had in mind.

28:39So I think definitely smartness in high level is the direction that we are pushing forward while maintaining the visual quality or improving it. But there are so many specifics and capabilities and use cases, especially for developers, that I think this release has some, but next release is going to be also like, and we have these coming releases in the pipeline. I cannot share about the timeline, but it's just so exciting. Yeah, it should. Maybe it should, yeah. But I'm so excited. I'm happy, and the momentum is unmatched here, on the image generation side. I love that. Any other capabilities folks are excited about?

29:19I'm really excited about factuality. And so that kind of goes back to the point that sometimes maybe you need to make a little diagram or an infographic for a work presentation. And it's amazing if it looks nice, but that's not enough for that. It actually has to be accurate. You can't have any extraneous text. It just kind of has to both look good and also be functional for that purpose. and I think we're just scratching the surface on what these model kits can do with that and I'm really excited about some of these upcoming releases like us getting better at that type of use case so that my dream one day is that these models can actually make a slide deck for me for work that looks nice.

30:00This is every PM's dream. Every PM's dream. I'm trying to outsource that part of my job to Gemini and I think we play a really big part in it. Awesome, I love it. Well, I think folks are going to be super excited to try these models. Thank you all four of you and for the rest of the team for making this happen. So I appreciate all the hard work. I'm excited for this. And thanks everyone for watching Release Notes. We'll see you in the next episode.

From the publisher

Join host Logan Kilpatrick in discussion with some of the minds behind Google's new state-of-the-art image model, Gemini 2.5 Flash. Product and research leads from the Gemini team break down the technology behind its key capabilities, including interleaved generation for complex edits and new approaches to achieving character consistency and pixel-perfect control. With Nicole Brichtova, Kaushik Shivakumar, Mostafa Dehghani and Robert Riachi. 

Watch on YouTube: 

Chapters:
0:37 - New model introduction
1:21 -Demo - Image Editing
3:44 - Text rendering capabilities
4:44 Beyond human preference evals
6:44 - Text rendering as a proxy for quality
8:38 - Positive transfer between modalities
11:25 - Demo - Multi-turn, context aware image generation
13:54 - Pixel-perfect editing and character consistency
15:51 - Interleaved image generation
17:59 - Specialized vs. native models
19:52 - Understanding nuanced prompts
20:59 - User feedback shaping model development
22:37 - Improvements in character consistency
24:17 - More natural looking images from team collaboration
26:41 - What’s next for image generation models

More from Google AI: Release Notes

All 30 episodes
Behind the scenes of Google's state-of-the-art "nano-banana" image modelGoogle AI: Release Notes · 31 min
Listen in VO