Nano Banana Pro: Hands-on with the World’s Most Powerful Image Model

26 Nov 2025 · 36 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Google AI: Release Notes - Episode: Nano Banana Pro: Hands-on with the World’s Most Powerful Image Model

Episode Overview

In this episode of "Google AI

Release Notes," host Logan Kilpatrick discusses the newly launched Nano Banana Pro model, built on the Gemini 3 Pro framework. The episode features insights from team members Robin, Yilan, Abhishek, and Sherry, as they demonstrate the model's advanced capabilities, including visual reasoning, text rendering, and multi-turn generation.

Key Topics Discussed

  • Introduction to Nano Banana Pro
  • Built on the Gemini 3 Pro model.
  • Enhancements in text rendering, infographics, and structured content generation.
  • Capabilities of Nano Banana Pro
  • Text Rendering
  • Improved accuracy and quality for text rendered in images.
  • Side-by-side comparisons with the original Nano Banana model highlighted significant improvements.
  • Visual Reasoning
  • Demonstrations included generating accurate images based on complex prompts (e.g., wine glass full, clock showing specific times).
  • Multi-turn capability allows users to engage in extended conversations with the model without losing context.
  • Infographics and Complex Content Generation
  • Ability to create detailed infographics based on user prompts, grounded with real-time search data.
  • Enhanced support for multilingual text rendering, showing improvements in various languages without explicit additional training.
  • User Feedback and Benchmarking
  • Continuous user feedback drives iterative model improvements.
  • New benchmarks established for assessing image quality, especially in text rendering.

Key Examples & Demonstrations

  • Wine Glass and Clock Example
  • Demonstrates the model's improved understanding of common visual prompts and expectations.
  • Infographic Creation
  • Example of generating an infographic explaining photosynthesis.
  • Emphasis on the model's ability to ground information with Google Search for accuracy in real-time queries.
  • Multi-Turn Conversation
  • Comparison of previous models to demonstrate enhanced performance in handling complex user interactions.
  • Code Explanation Infographics
  • The model was able to generate infographics from code segments provided by users, showcasing high fidelity in image generation.

Feature Highlights

  • Advanced Editing Capabilities
  • Improvements in editing images to reflect user specifications, with a focus on maintaining consistency and accuracy.
  • Character Consistency
  • Significant advancements in maintaining character consistency across generated images.
  • Resolution Support
  • Ability to generate images in 2K and 4K resolutions, emphasizing the necessity for high-quality output in generated images.
  • Future Outlook
  • Discussions surrounding the potential for future model enhancements and additional features based on user requests and technological advancements.

Conclusion The episode concludes with excitement about the capabilities of Nano Banana Pro and its integration into various Google products, including the Gemini app and Notebook LM. The team expresses gratitude for user feedback and anticipation for further developments in AI image generation.

Key Takeaways

  • Nano Banana Pro represents a significant leap in AI-powered image generation and rendering.
  • Continuous improvements driven by user engagement and benchmarking are crucial for product evolution.
  • The integration of real-time search data helps enhance the model's output relevance and accuracy.
  • The episode highlights the collaborative environment within the Google AI team, emphasizing the importance of expert input and feedback in developing advanced AI capabilities.

Episode Details

  • Host: Logan Kilpatrick
  • Guests: Robin, Yilan, Abhishek, Sherry
  • Watch on YouTube: [Nano Banana Pro: Hands-on with the World’s Most Powerful Image Model](https://www.youtube.com/watch?v=hk6gwiZmSWA)

Chapters

  • 00:00 - Introducing Nano Banana Pro
  • 02:00 - Enhanced world understanding
  • 04:59 - Advanced text rendering
  • 05:49 - Gemini 3 Pro's influence
  • 09:30 - Multi-turn & infographics
  • 14:04 - Text rendering comparison
  • 16:26 - Multilingual text support
  • 18:22 - Infographics for learning
  • 24:00 - Multi-image input
  • 26:38 - Resolution & fidelity
  • 30:07 - Advanced editing & style
  • 32:09 - Practical use cases
  • 35:26 - Future outlook & thanks

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:07Hey everyone, welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today we're talking about Nano Banana Pro. We're here in Mountain View with Robin, Yilan, Abhishek, Sherry, folks from the Nano Banana team, the Gemini Image team. I don't know what our internal codename is for the team because there's 50 different names. But yeah, the sort of high-level news is we launched the Nano Banana Pro model built on top of Gemini 3 Pro, which has been super exciting. Obviously folks loved the original Nano Banana model. You've got the swag, which everyone pings me on Twitter and is like, how do I get the swag?

0:41I'm like, you can't get it. Exclusive. Hard to come by. But Robin, maybe you want to give us the sort of like high level for folks who maybe like miss the context of this model launch. And then we'll dive in and actually look at a bunch of demos and examples. I think so. Nano Banana Pro is particularly amazing at tasks like text rendering, infographics, actually like generating structured content. I think we have plenty of cool examples where we can do tool code and actually generate content while doing a search on the internet. And on visual reasoning also, it's super cool. We can do multi-turn generation.

1:17You can do edits where you take multiple people in and blend this into an image. And yeah, we're excited to showcase some examples today live of the model. I'm excited too. I want to make sure we also talk about this text rendering part because when we had we did the an episode about Nano banana one or the original Nano banana model and Kaushik and others were talking about how sort of like text rendering is actually like one of the main things from like an overall image quality a model quality perspective that like we can use to benchmark the model So I'm curious like if I don't know if we have like a public benchmark that we put out or that there's something we track But I'd be interested here But Abhishek, do you want to show us some examples to sort of kickstart how this model is better?

2:01I think is the top level question. Yeah, so we can drive into multiple examples. Like, let's start with one common example which people have been trying. And let me compare against Nano Banana, a side-by-side comparison. Let me just, so people have been trying this wine glass example where you want the wine glass to be full. It's a difficult example because, as you can imagine, most of the wine glasses, the pictures which people take or the pictures which are online, rarely you see full wine glass because it's served with something different. And also the time. Most of the time clock images, people see always 10-10.

2:45That's the standard time which is there in the clock. and so whenever you ask these imaginative models to come to show this wine glass or the clock example it always show either shows a wine glass which is not full or a clock which is like not showing the time which you want it's either 10 10 or some standard time which is mainly a data issue because that's what exists mainly in the internet uh but to a surprise like it's not like something we optimized our model for we trained our model and then we started seeing that people were actually trying out this prompt and it works like surprisingly well you can see in this like even on my first try like you click in on these so we can see them in full screen yeah uh try to make sure it's really full yeah it looks full yeah it's actually very full uh you can see the time on the clock it's also like perfectly depicting 5 30 uh whereas in the right hand side if you see like the nano banana output which is uh it's not showing the wine grass which is full it's also not showing the clock which is like accurate uh so this is like some example of like better world understanding I would say just because it's been just because it's trained on Gemini 3 which is a much better multimodal understanding that really helps us for image generation whereas all this better world knowledge comes up to this model so this is one particular example for the clock one when we were pressure testing our model not only we can generate the clock with the targeting time, but we can also editing that.

4:18So it's a, I think with this joint like text image and word knowledge understanding, and then also bring that knowledge into editing the target image. That's really cool. We have never seen that high successful rate for Nano Banana. Yeah, that's super interesting. I wonder, I've been seeing a bunch of these like new practical use cases enabled by the original Nano Banana model and VIO coming together. And I feel like with even more high fidelity editing, it'll be interesting to see how many more new use cases come up with that. But Abhijek, we have another example. Yeah, so I just took an example image.

4:57I asked the model to show all the consonants in yellow, all the vowels in red, and you can see it perfectly generates this image where all the vowels are described in red, all the consonants in yellow, I can keep like adding and it's kind of suggesting that the model is actually able to reason like it's not like an easy prompt, you have to really dive into each character. Yeah, and I think the fine grained capacity of like following the text prompt and the image content, it's both showcased on this example and also on the wine glass example where it doesn't output just a generic image but really reason about like what's the specific user input and trying to match exactly like the full wine glass, for example, instead of like just outputting an average wine glass.

5:43Yeah, I actually have a, before we look at another demo, I have a broad question, which is, and Ylan, maybe you can answer this one, how much of the sort of like capabilities that we're seeing here are just like coming out of the box because 3.0 Pro is a super smart model versus like, is it just that like that gives you all a better baseline and then it's just like much easier to hill climb on stuff or is it like actually it's you know some of these things were free versus some of them you really had to go deep on or actually some of them it was harder because 3.0 pro has its own biases or trying to do things like i'm curious um yeah i think um it's it's both i would say like first of all the mountain image model is trained based on the gemini 3.0 model as abyshek said it brings a lot stronger word knowledge and multi-modal understanding capabilities.

6:31We also put a lot of work into preparing the data and we have a much larger data set, data collection that we use to train this pro model and have much better synthetic captions that we generated for the images. I believe that was an important part of bringing a lot of the capabilities such as word knowledge and like description of the clock for example and also like a full wine glass and also maybe some of the infographic capabilities that we're also really proud about yeah this this story of the fusion between image understanding and like how as we continue to push the frontier from an image understanding standpoint it actually translates back to the image generation i think is really interesting and you sort of get this unique flywheel one of the um one of the use cases that I think this is independent of image generation that we have a ton of customers that I think are looking into this is around robotics and like if you can get high fidelity sort of synthetic captions of like the the spatial world it potentially is one of the unlocks for robotics which is a sidebar but also really interesting.

7:40The other side is the generation they hold with the understanding because right now we can see the visual reasoning of our model they can do a segmentation they can do bunion box action and I also can do a a little bit planning of robotics. We actually have robotics expert testing our models. So it's accurate to see the mixture of all the experts coming together. So first pressure testing our model and then figure out a way if we can further improve the model. That's really cool about like working at Google with all the experts. Yeah, so want to add, so because like we use Gemini itself to its leveraging its multi-model understanding ability to generate almost all the captions we have.

8:20So right now, we're also having a much closer collaboration between the generation team and also the understanding work stream so that we can really have, as you said, a flywheel about what our needs are and what feedback we have and how that can be used to improve multimodal understanding in Gemini. Yeah, we've had a long thread waiting to get JB on the podcast and a few of the other folks. So he's on our list of folks to talk about. Yeah. And specifically about understanding, I think it's also interesting, like we were talking about robotics and you have plenty of tasks where actually you would rather segment an object doing image generation rather than like use text to describe like exactly where this object is in the image.

9:03So I think you can really think about like generation and looking like key understanding capabilities of the world that you wouldn't have through text, for example. Yeah, that is super interesting. I don't know if we have a robotics example, but I'd love to see. I don't know. Abhishek, maybe not to a robotics example, but if you want to go to, we can look at our next example. I don't have a robotics example, but just adding on to things which we improved over Nano Banana, I think one thing which is like we saw a lot of user complaints where the model was not good. The Nano Banana was not perfect for multi-turn, so we have tried to improve the model a lot for multi-turn.

9:39So in this case, you can see like I just asked another reasoning example, like arrange all the words alphabetically and it then rearranges everything into alphabetical. So I think the model would feel much better for such multi-turn conversation. It can, you can go up to like five, ten conversations and you won't, hopefully you won't see much issues which we, which people saw before. So that is like one particular example. then some other examples like this thing I'm really excited about which is like so I had a github I took my github repo I wanted to understand my previous code which I wrote using tensorflow so and this is something which I've been using exclusively at google which is I paste someone's code I don't want to really read all the code that is written or ask Gemini to summarize it, I just instead just ask my model to say, generate some infographic poster explaining this image.

10:47And since our model can support like 2K or 4K resolution, so that's another positive. Let me try generating something at 2K. And yeah, you can see it's a code. It's like, uh we can uh i've been able to play with uh codes which are even like 500 lines of code and i just feed everything to model it just simplifies my daily life so much uh uh oh i have once i it started explaining and let me just see to explain uh generate an image corresponding to it i wonder does the ordering of the prompt change this like is the i don't know if we have i assume we have best practices somewhere, but like do prompt first, context first, or what is there a preference on ordering?

11:34Just instructify the prompt. Be explicit and say generate an image. Yeah. Abhishek didn't say please, that's why. I did not say to generate an image, I just said to generate an infograph and started writing text. Now I've explicitly said infographic image, so now it's generating the image. to your key question also like how do you organize the information the answer uh do you send an image back or text and i think uh gemini has been like text first and there is three that know that we have like really good like text rendering and like detailed generation we can really think about like how can we use gemini to actually like generate image where it's actually like when it's best than uh text yeah i feel like this is an interesting and i think the the image output is an example of this the genui stuff which a bunch of the products with now with gemini 3 actually has this as well which is like when do you output text when you actually write code to make an interface for users so it feels like as gemini gets smarter it's also learning like not not just text like what are the other modalities that when is tool calls as another example so um this looks awesome though yeah yeah it's like like please zoom in a bunch i'm trying to yeah so it's started look right yeah it looks I haven't looked at TensorFlow code in a long time.

12:48It looks convincing to me. Yeah, it looks quite convincing. It analyzed from my code that I was reading in some Amnesty data. First, I trained a teacher network. It wrote all the parameters of the teacher network, like what were the con layer sizes, the different layers which were used. And then once teacher is trained, we train the student network, which is whose architecture is perfect. Just to clarify, this is not the architecture for Nano Banana Pro. this is a source project so yeah don't don't get your hopes up and yeah this is the distillation loss you get teacher outputs and then you come uh you you can get uh student uh distillation loss normal loss uh you have all the hyper parameters and everything described uh how many steps you train the teacher for what optimizer was used uh the same for student and yeah it's like perfect I feel like it's perfect.

13:44Like I don't see any I mean, I think I don't have context to make sure that the like the It is like factually accurate from an architecture standpoint, but the text rendering and overall setup looks awesome Can we do the original Nano banana model as a side-by-side of this too? I'm just curious I think this is one of the places where There's lots of domains where like single turn editing I feel like the previous model was actually really good at that and continues to be good But I feel like this is one of the use cases where it's like clearly night and day difference from a text rendering capability standpoint.

14:14Yeah. And you can see the whole like flow diagram, like, Abhishek actually didn't ask for it and it came up with this through thinking and like organized in a very like clear way the information. It's crazy. And have you noticed that this has poster aspect ratio? So, yeah. Yeah. Yeah. So, Nano of Nano always generates a square image. Our model like is adaptive, it kind of understands our tool. This seems like a prompt for which I have to generate a different aspect ratio image. So it's more intelligent in that way also. I love that. Can we zoom in on this? Yeah, so this is Nano Banana. I think firstly, like you can see the text, for instance, it's not perfect.

14:55It made some mistakes in the text for Nano Banana. Yeah, the text rendering is not perfect. But if you compare against Nano Banana Pro, like we were hardly able to find any mistake in the text. Yeah, this was better than I remember though. Honestly for Nano Benet 1, I like the pro model But it is interesting that yeah Yeah, I want to say that for text rendering this is quite a hard test to measure in terms of metrics so During the last launch phase we actually have a small team with very proactive engineers like for actually working on like defining the prompt set with very simple short text prompt just measure that you can Character by character verify the metrics verify the successful rate and also whether like long under short under specified text prompt and long prompt and The prompt with like different language space and they come up they come up with like alterator human eval, super reader, and then eventually we also assembled a team within our work stream, people on different languages like French, Chinese, Japanese, yada yada.

16:07So we actually came up with a wide range of rich metrics in order to measure the success of Pro model. Yeah, that's quite impressive. This was one of the threads that was most surprising, and I think we got lots of feedback from the original model about sort of like I-18N non-English languages and i feel like i saw a bunch of i forgot what the metrics were specifically but this model is like now soda across a bunch of the other languages that are not english specific which is actually great because i think the original model was like super popular like all these like uh country uh geography specific trends in like indonesia and other places where like people were trying to do really cool things with the model so being able to render i assume rendering text in those languages as well is state-of-the-art.

16:51So we tested other models. We don't have any external benchmarks, but we tested different models as well as Nano Banana on a variety of different languages. And we are almost better in almost all the languages, I would say. And this is a behavior which we did not explicitly added in the model. We did not explicitly add training data for each language. We just targeted a general improvement and then in the end when we evaluated all different languages, it was very surprising for us to see that everything improved, which was super good too. And I think to piggyback on what Ilene was saying earlier, it's really a matter of having a very wide variety of data and improving the amount of concepts and data with a lot of world knowledge we train on compared to Nano Banana.

17:44Yeah, we should have just gone and benchmark max the clock and the wine stuff. We didn't do that, which I think folks will feel as they use the model. And I do think it would be easy to go and try to hill climb a bunch of the narrow stuff, especially because we got lots of feedback from what folks wanted. So it's cool that we have made general progress. Do we have more examples? Yeah, I have some. So on the side of text rendering, we are excited for people to try out infographics where people can use the model to explain different concepts in an infographic fashion. Like, for instance, let me show me an infographic explaining photosynthesis.

18:28And is this grounded with search as well when you do these infographics? Yes. We have an option here to ground with search. So for a lot of queries, if you're asking for some real-time queries, like show me an infographic about the weather for the next seven days, it will use search to be grounded with search. That's awesome. And can we talk, as you're typing this prompt, Abhishek, and as it's generating this image, the reasoning piece is new for this model. And maybe the model was doing some reasoning sort of behind the scenes before. but is there any like interesting part about this? Like, is that part of the story of how we're hill climbing quality?

19:08And again, is it reasoning in, like, I assume it's not actually like generating an image. It's just like generating like a really similar to how thinking works on the text models. It's just generating text and then it uses the like reasoning text to then make the image. Yeah, exactly. Yeah. And I think that's where, again, like the fact that the image model is so good at following the text that's given as input, really plays nicely with thinking because like the thinking trace is really long and still like the model is able to reason based on it and like generate an accurate image. That's interesting.

19:41My intuition would have been that would actually like degrade model quality because you have like way more like actually again, one of the challenges for Nano Banana originally was like sort of you had to have like more short terse requests and like not give as you gave it more context like it sort of overloaded the model at least a bunch of tests that I was doing personally. The other way around. Really? Yeah. Okay. See? My intuition doesn't match the way that the model is trained, which is why. Yeah. And also the thinking, the model also does self-critic. So it generates results and then try to compare, like see the response, compare it with the content and see if actually achieved the user intent.

20:18If not, the model will redo it and try to improve on itself. So that might also be the reason why for some of the challenging prompts we can actually see a higher successful rate That is interesting and Robin just to clarify it is what you're saying is true for The pro model and the original nano banana model or just for the pro model? It's true in general like It's good to basically give a lot of context and define clearly what you want and the longer the prompt the more detailed you can specify what you actually want to generate. Okay, cool. This looks impressive. So I just asked you to explain photosynthesis in detail.

20:57I'm not a biology expert, but it's like talking everything in detail. It talks about all the photosynthesis equations. It talks about all these things in chloroplast, granum, everything. And I think you'll have to really zoom in to find any text mistake. Like I don't see anything on top of my mind, but I feel it's going to be really useful for people to really Understand deeper topics much easily in an informative fashion instead of going through all the texts We like visualize all this information via an infographic image This this passes my like seventh grade science vibe test. Yeah, we can also like This 1k or 2k generation.

21:41This is like 1k, but I can ask you to make this image simpler Because we had the same impression with Abhishek, it felt amazing, but we're not biology experts, so it wasn't clear to us was it actually accurate or not, and it was a bit too complex. But we sent it to biology experts and it's correct. It was the seventh graders who were looking at this. We have biology professors and researchers helping us to grade those responses. nice yeah this is like much simpler uh how how well does this work for stuff like like obviously like uh photosynthesis is like very in distribution i would assume for like the knowledge of the model is it able to sort of grok the the nuance of like topics that are like not like i don't i don't know what a good recent example of this would be but like a net new research paper about like something that's like far afield that like maybe isn't super like well represented in training data?

22:40Yeah, like, so I tried, like, once the Google earnings were out, like, I just asked it, like, show me the latest Google earnings in infographic fashion. It went through Google search. Since it's Google search grounded, it found all the Google earnings and then generated the infographic. So I think since it's Google search grounded, like, you can really tested with very recent things which are for sure out of distribution also and the model should be able to do it quite well that's awesome i love it i feel like folks are going to get a lot of value out of this uh infographic use case i feel like i'm guessing uh i don't know if it's going to be available in notebook lm but this feels like a very notebook lm uh bring to life experience yeah it's available in notebook lm like and then the notebook lm team have integrated like lots of new things so it's really awesome to play it in notebook lm also yeah i feel like audio overviews plus this plus i'm like it's yeah i feel like it's gonna be awesome um i want to take can you do is it just uh text and uh and image inputs or can i also put in like audio and video and stuff like that as well uh right now it's only image and text we uh we hope to support video also but yeah currently it's image and text input.

24:00What about like native audio in? Is that possible? Not right now. Not right now. It's great that we launched everything simultaneously. Like yeah, no, no, no, no, no, no simultaneously launched with notebook. There are other few exciting changes which people will see all like simultaneously getting shipped. So it's great to be in Google and see like all of these things all simultaneously being launched. I have right now all the four presenters and Logan, so let me just say show them celebrating. Tell it to put us in a Google office and see it. Oh yeah, yeah, I will put that in a follow-up prompt then.

24:38And so from a user perspective, would I expect like, is there any difference in from like a model inference time perspective, like putting in all these images and fusing them, Does that like generally on average like take longer for the model to be able to bring the references into a single image? Or is it like basically the similar time to image generation across all of them? It's a bit longer, right? But like the difference is not that much. I think like mainly the different the bottleneck is image generation. Yeah, yeah. This looks good. Could we zoom in and... Yeah, but I think we can make it...

25:13it did not generate a photo realistic one. We can generate a photo realistic one. I wish this is what was happening where bananas and balloons and yellow champagne. Shall we try again? Yeah. Yeah. Yeah. And in terms of latency, I don't think there will be much difference between inputting one image versus inputting multiple images. It should be roughly the latency should be close to like 30 seconds or something, which is a bit higher than Nano Banana, but yeah, quality comes at a price of... Yeah. Yeah. I feel like, and so that's one of the, I think the like world knowledge grounding stuff. If you have more references, I'm trying to, I'm trying to accumulate my list of like when I should be going to the pro model versus the original Nano Banana model, other things.

26:04It likes, it likes doing us in this comic book. What is the prompt? I can't see. It's just show them celebrating in a Nano Banana pro launch party. Let me just say show a photo realistic image and you you're not saying please that's why i'm also excited about um is there any like interesting technical detail about like what the in my mind i think of like 1k and 2k and 4k as like i don't sort of think deeply about this and i just imagine there's some magic happening behind the scenes but like from a model perspective is there like is this actually like technically difficult if we can make 1k work like the what's the difference is it a model understanding piece that's different between 1k and 4k or something like that or is it just mainly about infra and actually if you have like serving cost in mind of course like generating 2k or 4k is going to be quite more expensive compared to 1k so you really want to make sure that if you start generating 2k it's actually going to be worth it and the image quality is going to be much better.

27:08So in a sense, you could just up sample a 1K image and say it's 2K, but that's not going to bring something new. Whereas when you generate 2K, you really want small text to be perfect. And that's the kind of thing we think about when we propose higher resolution, for example. Yeah. Also, in addition to training, for example, 2K or 4K model, there is also a requirement on the training data's resolution. Got it. It has to be higher. Yeah. We need 8K now. now that we zoom in on this this looks this looks very good yeah uh i think yeah it like perfectly almost perfectly captured all of our faces like i'm looking at the hands i'm looking at i thought abhishek i thought our hands we were holding the same yeah i also we're not we're not we have we have separate hands and we have our own glasses yeah yeah the character consistency is a big thing for there is also a label oh yeah Where?

28:05Nanobanana. Yeah, the text rendering in natural image setting. Yeah. So great. What are these little things that are sitting on the table? I'm not sure. I think the model figured that we are launching some new product, which looks a banana looking thing. We're here. Electronic. Banana razor. Electronic gadgets. I love it. That's awesome. Abhishek, do you want to talk about how we were fighting to make sure that character consistency is also good in the pro model? Because it was so good in Nanobanana. Yeah, I feel like it's a high bar. Yeah, it was a very high bar for us. And I think this took us the maximal amount of time to hill climb on.

28:42Because when we trained the model, we were super happy with almost all other capabilities. But consistency was something which we found that people really loved about Nano Banana. We were initially struggling to come to parity with Nano Banana. But now if you play, you'll find it's even better than Nano Banana for consistency. but it took a great deal of time both from data curating better evals and changing our training strategies a bit but it did take us quite a bit to hill climb on this character consistency but we are now really excited for it it should work much better and it should support multiple people input and yeah it's a new capability I think I love playing with, I love imagining myself in times 30 under 30 poster.

29:33Hey, if Nano Benet Pro goes well, who knows? Who knows? Maybe they'll make it happen. So character consistency, taking a major leap, the sort of like real world knowledge, taking a major leap, sort of some of these vibe tests start to work, the class. The text rendering, infographics. Text rendering. Any other things top of mind from a model capability standpoint? I would say these are the main things. We definitely improved also on style transfer compared to Nano Banana. I think because of this work knowledge and better grounding, we're better at a bunch of editing capabilities like chart editing.

30:09Imagine you have a very complicated graph with pie chart, a different format of the layout. I want this because I can't make, I'm incapable of making slides. My slides are horrible, but then I see beautiful visual slides and I'm always jealous. You can just talk to the model to improve the style, the arrangement of the text and the bar chart. Yeah, so part chart into pie chart and some other things. It can also do some mass computations directly from the numbers in your image and then make a process result into the edited images. So I see example like you presented with a Confucian matrix. That's a common thing in statistics.

30:49And then you ask the model, okay, tell me the percentage in each entry of the computation matrix. And that's pretty accurate. Can I do, maybe this is a two on the nose question. Can I do transparent backgrounds? Okay, no. I got a ton of feature requests. Everyone was like, when is transparent backgrounds? What stops the model from being able to do? I don't have a good understanding of how transparent backgrounds actually work from a technical perspective. but is it possible to like it's possible but it's a matter of like getting the right data because uh you don't have much data out there that has like the transparency channels and it's also like you need to actually like uh change a bit the way you train your model with while making sure you're not going to regress on like all of the tasks you're currently good at so it's a bit tricky to get right i we're definitely going to get there but like not for this actual version okay people want it so we need yeah i feel like getting transparent background images is the most painful thing anytime you're trying to do something so we can nail that and land it i feel like folks will love it yeah uh but in general like i i think uh even though like these are the new capabilities that you're talking about but even in simple like text image and editing the model should feel much superior like because that's the main metric we climbed on these all capabilities kind of emerged but simple text image queries simple edit image editing uh the model should feel much better than what it was before.

32:15We can even try some examples of like, I was trying to play with it, like let me just simply show me a funny pie chart around some random topic, and then we can ask it to even edit the image from pie chart to 3D bar plot or something. Yeah. And I think it really relates to this capability of like the model is really good at multi-turn generation. Yeah. I feel like multi-turn generation has started, I'm just remembering back to the original, I forgot what we even called the Gemini 2 flash image model. And you'd sort of like actually in some contexts and that model was great. And we were the first ones to ship that sort of like native image generation adding capability, but you'd like visibly be able to see some of the regression over time.

Read the full transcript

33:03And I feel like now it's sort of like cross the uncanny valley of like, you can just like continue to do the editing and generation that doesn't, uh, it like seemingly gets better over time, which is really cool. Yeah. Uh, like for instance, for this image, like you can see, like this is a normal text image query, but you can see that all the labels are like perfectly accurate, like, uh, 40%, 30%, all this like percentage labels on a pie chart as pretty accurately depicted here. And they actually add up to be 100%. Yeah, yeah. I gotta measure. We gotta get the true 100 % overlay. We should run code execution and use a visualization.

33:43I'm curious how close the actual gap is. But my vibe check is I feel like it looks right. Yeah, yeah. That's awesome. It looks quite good, no? Yeah, yeah. It looks quite good. and now let me see and yeah it did some reasoning to like make it funny going through memes I've already seen true

34:03staring into the fridge and hoping for a miracle is a real one I open up the fridge all the time like maybe there's something I really want in here and every time I'm disappointed or these MKs in the office I also remember it's just in general I think there was a use case you know like on our team, a lot of us that were in the US and some of our team and the rest of them, they're like in Europe, like for example in London. And every year when the daylight time, daylight saving time changes, it's mismatching in two countries. So there's a week when our meetings gets misaligned. So on that week, our helpful TPM send out a message to remind everybody there's a change in the meeting time and someone else on our team use that just chat message and sent that into Nano Banana Pro, and it made a very, very cute, like, reminder poster, like, with all the information, and, like, even has London bridges and stuff.

34:59And so it's really nice. Like, it's just nice to see high visual quality and, like, for just our daily tasks. Now, maybe, like, now we don't have to send chat messages. We can just generate a good-looking, cute image for that and to communicate what we want to tell others. Yeah, there's something about images too that like it's just so easy to like grok Information and I feel like we're so text-based and working and all this other stuff that like you just sort of Resonate so deeply with image, which I think is why this is so powerful and also just like letting anyone I the the pro models Leveling everyone up and I feel like nano banana did this where like anyone can go and edit images and start to generate and I feel like now the Now this model taking that even further I think is a cool capability This was a ton of fun.

35:46I'm super excited for folks to get their hands on this model. I feel like across the Gemini app, Notebook LM, AI Studio, the APIs, and I'm sure a bunch of other products, folks will be able to experience this model. So thank you all for the hard work. Thanks for sitting down and chatting about this stuff and excited for Gemini Nano Banana Pro 2 or whatever we call it in the future. I want a gig of banana, but that did not happen. So Nano Banana Pro. Again, thank you all. This was a ton of fun. Thanks for watching release notes. Hope you enjoyed and we'll see you in the next episode

From the publisher

Introducing Nano Banana Pro, a powerful model built on Gemini 3 Pro, designed to enhance text rendering, infographics, and structured content generation. Tune in to learn about Nano Banana Pro’s advanced visual reasoning and multi-turn generation capabilities, and how this next-gen tool enables complex image edits and real-world applications. In this episode, we discuss how user feedback and continuous benchmarking drive model improvements, ensuring a superior experience for developers.

Watch on YouTube: https://www.youtube.com/watch?v=hk6gwiZmSWA

Chapters:
00:00 - Introducing Nano Banana Pro
02:00 - Enhanced world understanding
04:59 - Advanced text rendering
05:49 - Gemini 3 Pro's influence
09:30 - Multi-turn & infographics
14:04 - Text rendering comparison
16:26 - Multilingual text support
18:22 - Infographics for learning
24:00 - Multi-image input
26:38 - Resolution & fidelity
30:07 - Advanced editing & style
32:09 - Practical use cases
35:26 - Future outlook & thanks

More from Google AI: Release Notes

All 30 episodes
Nano Banana Pro: Hands-on with the World’s Most Powerful Image ModelGoogle AI: Release Notes · 36 min
Listen in VO