In short
Podcast Episode Summary: How Google’s Nano Banana Achieved Breakthrough Character Consistency
Podcast Overview Podcast Title: Training Data Episode Title: How Google’s Nano Banana Achieved Breakthrough Character Consistency Hosts: Sonya Huang, Pat Grady Guests: Nicole Brichtova and Hansa Srinivasan (Product and Engineering Leads, Google) Description: This episode explores Google’s Nano Banana, an image model that revolutionizes how users can see themselves in AI-generated worlds. The hosts discuss the development, technical achievements, and implications of this model for visual AI.
---
Key Concepts
- Introduction of Nano Banana
- Cultural Impact: Launched as a global phenomenon, enabling users to create images that genuinely reflect their identities.
- Character Consistency: Focus on maintaining character consistency in generated images, which is a significant development in AI.
- Technical Achievements
- Breakthrough Techniques:
- Utilization of high-quality data and disciplined human evaluations to enhance character consistency.
- Innovations in multimodal design and long context windows that help maintain contextual coherence.
- Human Evaluation Importance: Continuous human feedback is crucial for achieving qualitative improvements in models.
- User Creativity and Applications
- Creative Use Cases:
- Examples of users employing Nano Banana for various purposes, such as creating visually engaging sketchnotes and personal storytelling.
- Users leveraging the model for educational purposes, such as understanding complex topics in a more digestible visual format.
- Future Directions in Visual AI
- Accessibility and Imagination: Emphasis on evolving tools to not only capture reality but also explore imaginative possibilities.
- Personalized Learning: Future applications may include personalized learning experiences through AI.
---
Insights from Nicole Brichtova and Hansa Srinivasan
Personal Experiences with Nano Banana
- Creative Outcomes: Both guests shared their personal experiences using Nano Banana, highlighting unexpected creative uses like video models for consistent character portrayal.
- Community Engagement: The unexpected ways the community has adapted and utilized the model were particularly eye-opening for the developers.
Achieving Character Consistency
- Technical Innovations: Discussion on the combination of model architecture and quality data that enabled improved character consistency.
- Evaluation Strategies: The role of qualitative user input and community testing in developing a reliable model that feels emotionally resonant.
The Role of Fun in Utility
- Engagement Through Fun: The guests noted that the playful nature of the Nano Banana model serves as a gateway for users to engage with more serious uses of AI.
- Emotional Connection: Users connect emotionally with the outputs, which fosters a deeper relationship with the technology.
---
Discussions on Model Development and Future Outlook
Multi-Modal Approach
- Gemini Model Integration: The Nano Banana model operates within a larger framework (Gemini) designed for multi-modal understanding.
- Future Capabilities: Potential for more integrated experiences across different media formats (text, image, and video).
User Experience and Interfaces
- Product Evolution: Ongoing focus on making the model user-friendly, while also considering professional user needs for precision and control.
- Interface Design: The need for new UI designs that accommodate complex user interactions without overwhelming them.
Ethical Considerations and Safeguards
- Addressing Deepfake Concerns: Discussion on balancing creative freedom with ethical considerations, particularly regarding misinformation.
- Watermarking Technology: Implementation of visible and invisible watermarks to indicate AI-generated content, ensuring transparency.
Conclusion The episode reveals how Nano Banana embodies a significant step in visual AI through its ability to maintain character consistency and engage users in imaginative ways. The discussion also highlights the importance of user feedback in refining such technologies and the need for future innovations that will further personalize AI interactions and learning experiences.
---
Key Takeaways
- Character Consistency is a major advancement achieved through disciplined data quality and human evaluations.
- User Creativity is driving unexpected applications and stories, showcasing the model's versatility.
- Future Visions involve personalized learning experiences and integrated applications across various platforms.
- Ethical Safeguards are essential to ensure responsible use of AI-generated content while promoting creativity.
This episode serves as a fascinating look into both the technical and human elements behind AI developments, encouraging ongoing conversations about the future of technology in our lives.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There's something about like visual media that really excites people. that it's like the fun thing, but it's not just fun. It's exciting. It's intuitive. The visual space is so much of how we as humans experience life that I think I've loved how much it's moved people. I think we're really now making it possible to like tell stories that you never could. And in a way where like the camera allowed anyone to capture reality when it became very accessible, you're kind of capturing people's imagination. Like you're giving them the tools to be able to like get the stuff that's in their brain, out on paper, visually, in a way that they just couldn't before because they didn't have the tools or they didn't have the knowledge of the tools.
0:41Like, that's been really awesome.
0:59Today, we're talking with Nicole Brichtova and Hansa Srinivasan. the team behind Google's nano-banana image model, which started as a 2AM codename and has become a cultural phenomenon since. They walk us through the technical leaps that made single-image character consistency possible, how high-quality data, long multimodal context windows, and disciplined human evals enabled reliable character consistency from a single photo, and why craft and infrastructure matter as much as scale. We discussed the trade-offs between pushing the frontier versus broad accessibility and where this technology is headed.
1:35Multimodal creation, personalized learning, and specialized UIs that marry fine-grained control with hands-off automation. Finally, we'll touch on what's still missing for true AGI and white space where startups should be building now. Enjoy the show. Nicole and Hansa, thank you so much for joining us today. We're so excited to be here to chat a little bit more about Nano Banana, which has taken the world by storm. We thought we'd start off with a fun question. What have been some of your own personal creations using Nano Banana or some of the most creative things you've seen from the community?
2:10Yeah, so I think for me, one of the most exciting things I've been seeing is like the... it didn't occur to me, but this is very obvious in hindsight, is the use with video models to get actually consistent cross scene character and scene preservation. How fluid is that workflow today? How hard is it to do that? So what I've been seeing is people are really mixing the tools and using different video models from different sources. And so I think it's probably not very fluid. I know there's some products out there that are trying to integrate with multiple models to make this more fluid but i think the the difference in the the videos i've been seeing from before and after the nano banana launch has been pretty pretty remarkable and it's like much much smoother and much more like what you'd want in the video creating process with scene cuts that feel natural yeah so that's been cool um and i don't know why it didn't totally occur to me that people would immediately do that but yeah one of my favorite ways that i didn't expect is how people have hacked around the model to use it for learning new things or digesting information.
3:19I met somebody last week who has been using it to create sketchnotes of these like varied topics and it's surprising because text rendering is not something that it's not where we want it to be but this person has hacked around like these massive problems that like get the model to output something that's coherent and he's used it to try to understand the work that his father's doing He was like a chemist at a university and it's like a super technical topic. And so he's been feeding his lectures to Gemini with Nano Banana and then getting these sketchnotes that are like very coherent and like visually digestible.
3:56And for the first time, I think in like decades, they've been able to have a conversation with each other about his dad's work. And that was really fun and something that I didn't see coming. I think people are really working around, you know, like this model is amazing, but obviously it's not perfect. we have a lot of things we want to improve and I think I've been astounded by the ways people have found to to work with the model in ways we didn't anticipate and give inputs to the models in ways we didn't anticipate to bring out the best performance and unlock these things that are kind of mind-blowing did you guys in the building of it was there a moment like an aha moment where you kind of felt wow this thing's gonna be pretty good we just talked about it yeah I think Nicole had the aha moment.
4:40I had one where, so we always have an internal demo where we play with the models as we're developing them. And I had one where I just took an image of myself. And then I said like, Hey, put me on the red carpet and like full glass, just total vanity prompt. Right. And then it came out and it looked like me. And then I compared it to like all the models that we had before. Um, and no other model actually looked like me. And I was like, so excited. Wow. Um, And then people looked and they were like, OK, yeah, we get it. Like you're on the red carpet. And then I think it took a couple of weeks of other people being able to take their own photos and play with it and just kind of realize how magical that is when you get it to work.
5:21And that's kind of the main thing that people have been actually doing with the model, right? Turning yourself into a 3D figurine where it's like you want a computer, you want a toy box and then you as the figurine. So like you three times like that way to be able to kind of like express yourself and see yourself. in new ways and almost kind of like enhance your own identity has just been really fun. And that for me was like, oh, man, this is awesome. What was it about what Nano Banana did with you on the red carpet that was miles better than what everyone else has? It looked like me. And it's very difficult for you to be able to judge character consistency on people's faces you don't know.
5:58Yeah. And so if I saw, you know, a version of you that's like an AI version of you, I might be OK with it. But you would say like, oh, no, you know, like parts of my face are not quite right. And you can really only do it on yourself, which is why we now have evals on many team members where it's like their own faces and they're looking at the model's output with their own faces on it. Because it's really the only way that you can judge whether or not someone looks like you. Yourself and like faces you're familiar with. I think like when we started doing it on ourselves and it's like I see Nicole a lot.
6:29So like Nicole versus like random person we might eval on. Right. It's just a very big difference in terms of judging the model capabilities. And yeah, I think it's one of those things that it's like so fun, that preservation of the identity is so fundamental to these models actually being useful and exciting, but is surprisingly tricky. And that's why we see a lot of other models not quite hitting it. Well, I was going to ask you, I would imagine that character consistency is not just an emergent property of scale. And so maybe two questions. One, I'm sure there's stuff you can't tell us, but what can you tell us about how you achieved it?
7:08And then two, was that an explicit goal heading into the development of this model? Yeah. So I would say, I mean, yeah, I think there's definitely things that are tricky to say here, but I would say there's like sort of different genres of ways to do image generation. And so that definitely plays a part in how good it is. And I think it was definitely a goal from the beginning. It was definitely a goal because we knew it was a gap with the models that we released in the past. And generally consistency for us was a goal because every time you're editing images, right, like you want to preserve some parts of it and then you want to change something.
7:47And prior models just weren't very good at that. And that makes it not very useful in professional workflows, but it also doesn't make it useful for things like character consistency. And we've heard this for years from even advertisers who are trying to advertise their products and putting them in lifestyle shots. It has to look like your product 100%. Otherwise, you can't put it in an ad. So we knew there was demand for it. We knew the models had a gap. And we felt like we had the right recipe, both in terms of the model architecture and the data, to finally make it happen. I think what surprised us was just how good it was when we actually finally built the model.
8:21Yeah, right. Because I think we felt like we had the recipe exactly as Nicole said, but there's still always until you're seeing the model, you finish training, you're actually using it. You don't know how close you're going to get to that goal. And I think we were all surprised by that. yeah and I think the other thing is if we think about like what people expect out of editing when that you edit on your phone apps or like photoshop you expect a high degree of preservation of things you're not touching yeah and depending on how the models are made and how how the design decisions behind them that's very tricky to do but it's something people really like it's it's one of those things were like, it's shockingly technically difficult, even though it's something I think a lay person who's using the models would be expect to be like the basic thing about editing is like, you don't mess with the things you don't want to be messed with.
9:13Yeah. Back to that moment where you saw yourself on the red carpet and wow, that's actually me. And it took some of your colleagues a couple of weeks to have the same experience because they tried it with their own photos. The question is beyond, Hey, that's actually me that, you know, the qualitative test. Is there some sort of an eval that you can put against that to make it quantitative that, you know, we have achieved the thing that we set out to achieve here? Yeah. So I actually think, I think face consistency exactly for the reason Nicole said is, is quite hard. It's quite hard for other people to do.
9:46Yeah. Um, I will say in general, I think what we found with image generation in particular, that's unlocked a lot for us is like human evals are important. And so I think they're foundational. We have a team that works on helping us build sort of good tooling and good practices for evals and having humans actually eval these things that are very subtle. Like if you think about image generation, like faces, aesthetic quality, these are things that are very hard to quantify. And so I think human evals have been a big game changer for us. I think it's definitely, I think it's a combination of there's human evals, there is a very technical term, eyeballing, of the model results by different people.
10:30And there's also just community testing. And when we do community testing, we start internally. And we have artists at Google and at Google Define to play with these models. Our execs will play with these models. And that really helps, I think, kind of build that qualitative narrative around like, why is this model actually awesome? Because if you just look at the quantitative benchmarks, you could say like, oh, it's 10 % better than this model that we had before. And that doesn't quite grok that emotional aspect of like, oh, I can now see myself in new ways, or I can now finally edit this family photo that I cut up when I was five years old.
11:07And I probably shouldn't have. People have done that. When then I'm able to restore it. I think you really need that qualitative user feedback in order to be able to tell that emotional story. I think this is probably true of many of the Gen.A.I. and AI capabilities, but I think it's especially true of visual media where it's very subjective versus if you think about something like math reasoning, logic reasoning, where you can really ground it in an answer. And so it's more easy to have these very objective, automated, quantitative evals. To get to that level of character consistency from just one 2D image of someone is really, really hard.
11:50Can you walk us through maybe a little bit what are the technical breakthroughs that helped you drive to that level of character consistency that we actually haven't seen anywhere else? I mean, I think a key thing is like having good data that teaches the models to generalize. Right. And the fact that this is a base, it's a Gemini model. Yeah. It's a multimodal foundational model that's had seen a lot of data and has good generalization capabilities. And I think that like that's kind of the secret sauce. Right. It's like you really need models that generalize well to be able to take advantage of that for this.
12:27Right. Yeah, and I think the other nice part about doing this in a model like Gemini is that you also get this really long context window. So, like, yes, you can provide one image of yourself, but you can also provide multiple. And then on the output side, you can also iterate across multiple turns and actually have a conversation with the model, which wasn't possible before, right? One, two years ago, we were fine-tuning on 10 images of you, and it took 20 minutes to actually get something that looked like you. And that's why it never took off in the mainstream, right? Because it's just too hard.
13:00And you don't have that many images of yourself. It's too much work. And so I think it's both kind of the general, like Gemini gets better. You benefit from that multimodal context window. And you benefit from the long output and ability to maintain context over a long conversation. And then you also benefit from actually paying attention to the data, focusing on the problem. a lot of the things we get better at come down to there's a person on the team who's like obsessed with making them work like we have people on the team who are obsessed with text rendering and so our text rendering just keeps getting better because that person just like is obsessed with the problem yeah it's like it's not just about throwing high quantities of data in right i think that's one thing that's really important is it's there's this like attention to detail um and quality of, you know, all the things you're doing with the model.
13:50There's a lot of, there's a lot of small design decisions and decision points at every point. And I think that like detail orientedness of high quality is data and selections are really important. Yeah. It's the craft part of, I think, the AI, which we don't talk about a lot, but I think it's super important. Yeah. How big was the team that worked on it? To ship it, it took a village. Yeah. Especially because we ship across, many products so i think like there's like sort of the core sort of modeling team and then there's you know our close collaborators across like all the surfaces yeah when you put them all together you easily get into like dozens and hundreds but um the team who works on the model is much smaller and then the people who actually make all the magic happen and we had a lot of infrastructure teams like optimizing every part of the stack to be able to serve the demand that we were seeing which was really awesome um but really like to ship it we were joking that it takes like a small country.
14:44When you build something like this, do you build it with particular personas or particular use cases in mind? Or do you build it more with a capability first mindset? And then once the capabilities emerge, you can map it to personas? It's a little bit of both, I would say. Before we start training any new model, we kind of have an idea of what we want the capabilities to be. and some design decisions like how fast is it at inference time right they also impact which persona you're going after yeah so this model because it's kind of a conversational editor we wanted it to be really snappy because you can't really have a conversation with a model if it takes like a minute or two to generate that's really nice about image models versus video models like you just have to wait that long and so to us from the beginning it felt like a very consumer-centric model.
15:37But obviously, we also have developer products and enterprise products, and all of these capabilities end up being useful to them. But really, we've seen a ton of excitement on the consumer side in a way that I think we haven't before with our image models, because it was very snappy, and it kind of made these like pro level capabilities just really easily accessible through a text prompt. And so that's kind of how we started it out. But then obviously, it ends up being useful in other domains as well. Yeah. And I think one of the differences in philosophy, so previously we'd worked on the Imagine line of models, which were straight image generation.
16:15And I think one of the big philosophical goal changes in these Gemini image generation models is generalization, is a more foundational capability. So I think there is also a lot of... There's things where we want this model to be able to be good at this. like representing people and letting them edit their images and have it look like themselves. But I think there's also a lot of things like that are emergent from the goal of just having a baseline capable model that like reasons about visual information. Like I think one thing that's surprised me, I guess, as a callback to your earlier conversation is people can put in math problems, like a drawing of a math problem and like ask it to like render the solution, right?
17:01So you can put in a geometry problem and say, what is this angle? And that's an emergent thing of a foundationally capable model that has both reasoning, mathematical understanding, and visual understanding. So I think it's both. Can you maybe share, just out of curiosity, what's a good way to understand maybe the family mapping and the relationship between Gemini powering Nano Banana, VO, you know, all these other adjacent products and models that are all driven and benefit from the generalization and the scale of Gemini itself, how you co-develop and then where you want to take it from here?
17:43Our goal has always been to build the single most powerful model that can do all these things, right? You can take in any modality and you can transform it into any modality. And that's the North Star. We're obviously not quite there yet. And so on the way there, we had a lot of sort of specialized models that just got you great results in a specific domain. So Imagine was an example of that for image generation. VO is an example of that for video generation and editing. And so I think we're both kind of developing these models to push the frontier of that modality. And you get really useful outputs out of that, right?
18:19A lot of filmmakers are using VO in their creative process. but you're also learning a lot that you can then bring back into Gemini and then make it good at that modality. Image is always a little bit, I think, ahead of the curve because you just have one frame, right? It's cheaper both to train and at inference time. So I think kind of a lot of the developments you see in image, I expect you to see in video like six, 12 months down the line. And so that's always kind of been the goal. And so we have separate teams kind of developing these And then I think with image, we're now moving closer to Gemini and to that vision of that single most powerful model.
18:59And you will see that, I think, with some of the other modalities. And along the way, we'll release these experiences that are just like really powerful and like really exciting in that modality. So like VO3 was really awesome because it brought audio into video generation, right, in a way that we haven't seen before. Genie 3 was really awesome because it lets you in real time kind of navigate a world. And so in order to push that frontier, it's very hard to do all of that at the same time right now in one model. And so to some extent, these specialized models are kind of a testing ground. But I would expect that over time, Gemini should be able to do all these things.
19:34Oh, that's so interesting. Okay. We got to ask you about the name. I suspect that the name was a bit of a... It's an amazing product. I suspect that the name gave it a little bit of a boost. because it's so easy to remember and so distinct. So was it a happy accident or is there some creative genius who knew that this is going to be just the right name? It was a happy accident. So I think as many people know, the model went out on an ILL Marina where many models do. And part of that is you give a code name. If anyone hasn't used ILL Marina, you get to put in your prompt. You'll get back two responses from two models.
20:14they have code names until they're publicly released um and i think it was like we had to someone we were going out at like 2 a.m and uh nicole nicole's our wonderful pm there's another pm you have nina and someone messaged her being like what do we name it and she was really tired and exhausted and she was like this was the name of stroke of genius that came to her at 2 a.m this is you it was not me it was somebody on my team okay who named who named the model um i can't take credit for this another one of our pms um but what was really awesome was like a it was really fun i think that really helps yeah it's easy to pronounce it has an emoji which is critical for branding and she didn't overthink it in this era but she didn't overthink it and what was awesome is everybody just went with it once it went live and and i think it just like felt very googly and very organic and ended up looking like the stroke of marketing genius but no it was it was a happy accident and it just sort of worked out and people loved it and so we leaned into it and now there's you know bananas everywhere when you go into the gemini app um which we did because people were complaining that they were having a really hard time finding the model when they came into the app yeah and so we just made it easier yeah and yeah exactly i think there's like publicly people were like, nano banana, nano banana.
21:34How do I use nano banana? I had someone at Google I work with be like, how do I use nano banana? And I was like, it's Gemini. It's right there. Just ask for an image. Yeah. But I think that's the thing is like, I think Google's always had this really fun brand, right? Like it's like, it's not like, it's been a consumer oriented company at its inception and like i think it was really nice to to play on that that image people have of like google as a fun fun place fun company um and have this fun name it's also just like a really nice path to fun being kind of a gateway to utility right i i i think nano banana and just the model in general and what you can do with it like put yourself on the red carpet do all the childhood dream professions you had it's like a really fun entry point but what's been awesome to see is that like once people are in the app and they are using gemini they start to use it for other things yeah that then become useful in their day-to-day life like you use it to study and solve math problems or you use it to learn about something else and so i think it's maybe a little bit undervalued sometimes to like have a little fun um not just with the naming but also just like with the products that we built because it kind of gets people in gets them excited and then it helps them discover other things that you know the models are awesome at yeah i think other users like my my um like my parents and their friends are using it i think it's because it like had this reputation it was really easy it was really fun it felt unintimidating to try when you try it and you're like actually this this is very easy to work this works very easily it's very easy to interact with there's no like technology there's no like you know technology i think can sometimes be intimidating to people especially ai right now yeah um and i think the chat bot naturalness has broken a lot of that barriers but maybe more so with younger people yeah um and i think this like fun like yeah my mom like made was like making these images and having a great time and and and then realized she can use it to like remove people from the background of her images like these very practical things right it started very silly turned very practical then people can use it to realize like actually they can give you them diagrams or help them understand stuff.
23:48So I think there's also like a big accessibility component. Yeah. Where do you want to take from here? Maybe both from a model side and from a product side? On the product side, I think there's kind of a couple areas. Like on the consumer side, I still think we have a long way to go to just like make these things easier to use, right? You will notice that a lot of the nano banana prompts are like 100 words long and people actually go in and copy paste them into the Gemini app and like go through the work to make it work because the payoff is worth it. But I think we have to get past this prompt engineering phase for consumers and just like make things really easy for them to use.
24:25I think on the professional side, we need to get into like much more precise control, kind of robustness, like reproducibility to make it useful in actual professional workflows, right? So like, yes, we, you know, we're very good at editing consistency and not changing pixels, but we're not 100 % there. And when you're a professional, you need to be 100 % there, right? Like you really need kind of these precise, maybe even like gesture-based controls, like over every single pixel in the frame. So we definitely need to go in that direction. And then I think there's like a general direction that I'm really excited about, which is just about visualizing information.
25:01So the example I had about sketchnotes at the beginning and somebody kind of hacking their way around using Nano Banana for that use case, you could just imagine being able to do that for anything, right? And a lot of people are visual learners. I think we haven't really exhausted the potential of LLMs to be able to help you digest and visualize information in whatever way is most natural for you to consume. So sometimes it's a diagram, sometimes it's an image, and sometimes maybe it's a short video that you want to learn about some concept that you're learning in a biology class or something like that.
25:35So I think that's like a completely new domain that I'm really excited about, just these models getting better and getting past the point where 95 % of the outputs that you get out of these models are just text, which is useful, but it's not how we consume information in the real world right now. So it's really interesting. So on the product side, then, are you alluding to the fact that you might want to vertically integrate and build a little bit more product around it? And also, are you alluding to the fact that maybe the way you interact with some of these models isn't just through pure language and prompting over time, but more UI?
26:07Yeah, yeah. I definitely think the chatbots, I think, are an easy entry point for people because you don't have to learn a new UI. You just talk to it and then you say whatever you want to do, right? I think it starts to become a little bit limiting for the visual modalities. And I think there's a ton ahead room to think about what is the new visual creation canvas for the future? And how do you build that in a way that doesn't become overwhelming? Because as these models can do more and more things, it's very hard to explain to the user in something that's very open-ended what the constraints are and how do you work around that?
26:45How do you actually use it in a productive way? So I'm really excited about people kind of building products in those directions. And for us, you know, we have a team called Labs at Google that's led by Josh Woodward. And they do a lot of this kind of like frontier thinking experimentation. They work with us really closely where they take our frontier models and they think about like, what's the future of entertainment? What's the future of creation? What's the future of productivity? And so they've built products like Notebook LM and Flow on the video side. And I'm excited that maybe flow could kind of become this place where you could do, you know, some of this creation and think about what that looks like in the future.
27:21I think in the short term, it's very clear that, you know, this model has things that it's not perfect at. And so in the short term, it's obviously, it should work the way you expect it to every time, not just a lot of the time. And really make it so seamless and fix all these like small things where it's just like a little bit inconsistent in its performance. I think long term, it's, I think, Nicole covered that, which is, to me, it's in order to have that reality of really rich multimodal generation. So like right now, if you ask Gemini to explain something, it'll usually just explain in text unless you ask it for images.
28:03But if you think about like the platforms that have really taken off in the last like 10, 20 years for learning, right? Like we think of like Khan Academy started on YouTube. We think about like Wikipedia has a lot of images. Like it's very image focused. If you look up any math thing, you like diagrams. And so like that should become more like an actual part of the flow and a part of the way you use these models. And to enable that from a modeling point of view, it goes back to like we were talking about this multimodal understanding. and seamless generalization between modalities. Maybe the other interesting area, as we think about kind of these models being more proactive at pulling in, whether it's code or images or video, when it's appropriate for the user intent.
Read the full transcript
28:49I think this other exciting, I started out as a consultant in my career. And so obviously I made a lot of slide decks in my time. I still do. And I think there are some of these use cases where you don't actually really want to be in the weeds of creation. Like what you really want is, let's say you're updating your stakeholders on how a project is going, right? You want to pull in some context. Maybe it's meeting notes. Maybe it's a couple of bullet points. Maybe it's, you know, some other deck that you've created in the past. And then you maybe just want Gemini to go off and like do all the work for you, right?
29:20Like pull that deck together, format it, create appropriate visuals to make it really easy to digest. And that's something that you probably don't want to be involved in. and it gets more into these agentic behaviors versus I think for some of these creative workflows, like you actually want to be creating, you want to be in the weeds, you want to think about what the UI looks like that makes it easy for a user to accomplish the goal. And so like if I'm designing my house and I'm actually into designing my house, then I probably actually want to play with it and like play with textures and different colors and like what would happen if I remove this wall.
29:54And so I think there's kind of this spectrum of like very hands-off, like just let the model go off and like pull in relevant visuals, materials for a test that makes sense, all the way to like, how do you actually make a creative process like more fun and remove the tedious parts and remove the technical barriers that exist today with tools that we have? It's like this mix of giving the user fine grained control, like the precision control they want, but also at the other extreme, having the model be able to understand the user quest and anticipate, right, like the need and the outcome that it should be and do all the intervening work in between.
30:30It's almost like when you actually hire a professional for something today, right? Like when you hire a designer, you give them a spec and then they go off and then they do all that awesome work that they do because they have all this expertise. And so these models should be able to do that. And they can't really do that in many domains today. What do you think the next competitive battleground is in this world? I think there's still work to be done on making these models more capable. And so this idea of having a single model that can take anything and transform it into anything else, I think nobody has really figured that out.
31:04But I do think in order to actually drive adoption, there's probably two things. One is user interfaces. Like we still rely very heavily on the chatbots. And we talked about this. Like it's useful for some things and it's a great entry point. But it maybe isn't useful for all the things. And so I think starting to think about much more deeply about who are the users? What are they trying to do? How can the technology be helpful? And then what product do you build around it to make that happen is probably one. Do you think five or 10 years from now, the frontier will be advancing as quickly as it has advanced over these last few years?
31:42Five to 10 years from now feels like 20 years from now. Just the space and you guys probably see this too. The space is moving really quickly. Yeah. And, you know, if you asked me two years ago, I would have told you the space is moving really quickly. If you ask me today, I will tell you it's moving faster than it was two years ago. Okay, I'm going to ask you a very different question.
32:05So I know Google's very, very sort of careful and very concerned about deep fakes and that sort of thing. And I have to imagine when you saw how capable this model was, there's a big conversation about, okay, well, how are we going to make sure people don't use it in the wrong sorts of ways? How does that sort of a conversation go inside of Google? And are you guys sort of like happy with where it ended up? I think it's an ever evolving frontier also because it's this mix of you want to give people the creative freedom to be able to use these tools. Right. And you want to give users control to be able to use these tools in a way that don't feel overly restrictive.
32:45And you want to prevent the worst harm. Right. I think that's always the balance that we spend a lot of time talking about. And so obviously when you look at the outputs of the model, there's a visible watermark that says it's been generated with Gemini. So that immediately indicates that it's AI content. And then we also, in every output that we produce with our models, image, video, audio, there's SynthID embedded, which is invisible watermarking. And so those are kind of the visible ways or invisible ways in which we verify that our content is AI generated. We're very invested in it. And, you know, we believe that it is really important to give users those tools to be able to understand that when they're seeing something, it's not a real video or it's not a real image.
33:30And then obviously, when we develop these models, we do a ton of testing internally and also with external partners to kind of find as the models get more capable, you find new attack vectors, right? And like new ways that you have to mitigate for. and so that is like a very important part of model development for us and we continue to invest in and as as the models get better and there's new new things that you can do with them we also have to develop kind of new mitigations for you know making sure that we don't create harm but also still give users the creativity and the control in order to make these models usable in a product i mean i think it's a very very hard balance to strike right um because you will always have people using a tool in good faith you'll also always have people using it in bad faith um and i think i think it's hard it's like is it a is it a tool is it something that has responsibility so i think we we take this very seriously um users obviously are also responsible for what they do with the model but synth id really is an important technology that lets us like release these capabilities to people and have have some faith in that we can still verify right and and have a tool to to combat this missing the the risk of misinformation um but it's a it's it's a super tricky conversation and i think it's one that i've seen everyone take very seriously um there's a lot of a lot of conversations about how to balance both is that the standard now across the industry synth id yeah it's a google standard it's the google standard i believe there's Like every Google, like Imagine, the Imagine line, Veo, they all have SynthID when you use them in any product surface.
35:16All right. You told us we can't go five to 10 years down the road because things are moving too fast. We'll go one to three years down the road. Thank you. Two questions. One, what will be possible that we can only dream about today? and two what will the resulting change be to the way that we all live our lives i really hope that a year or two from now you could really get like personalized tutors personalized textbooks in a way right love it yeah i there's no reason why you and i should be learning from the same textbook if we have different learning styles and different starting points but that's what we do now, right?
35:59That's how our learning environment is set up. And I think across all these breakthroughs, that should be very possible where you have an LLM tutor that just figures out your learning style. What are the things you like? Maybe you're into basketball. And so I need to explain physics to you with basketball analogies, right? And so I'm really excited about learning just becoming way more personalized. And that feels very achievable. And we obviously have to make sure that we don't hallucinate and there's like a high bar for factuality. And so we need to ground in sort of real world content, but that I'm really excited about.
36:34And that really, I think just, it removes a lot of barriers for people, right? To your question on like what the impact is going to be. I think it just becomes much more, it becomes much easier to learn basically anything in a way that's very tailored to you that you just can't do right now. Could that be a Google product surface? somebody should look into it yeah and i think for the way it'll change how we live and work i think i think we i think working on these technologies i've already seen how it changes the way we work right because we we obviously use them um a lot uh i'm getting married we made our save the dates with our model.
37:18And so what I really think we'll see is, and just work, the amount, part of, I think the reason that the innovation has accelerated is we have these models, you have like code assistance, you have just like, you can use models to like filter things, to analyze huge amounts of data, like it's drastically increased our own workflows. Like what I can do this year versus two years ago It's just like an order of magnitude more work. And I think that's true of the tech industry. It's not true of a lot of other industries just because that integration into their workflows or into their tooling hasn't happened.
38:01So I think some people are like, oh, it's going to replace me. But at least what I've seen is it really just actually changes the amount of work an individual can get done. What that means for businesses or economically, I'm not sure. but I think it means we will just see people be more empowered to hopefully do more in the same amount of time. Like maybe you don't have to, you know, I have friends who are in consulting and spent a lot of time. They're like, I should spend a lot of time, like two hours making slides, tweaking, moving logos around. And like, hopefully they won't have to do that.
38:33They can actually spend time thinking about what the content of the slides like should be thinking, working with clients. And I think that that's hopefully what we will see in like one to two years. Given the trajectory that you see in these capabilities, are there interesting areas that you think startups should go do that Google itself might not get into? I think there's a ton of spaces, even just in the creative tools. I think there's a ton of room for people to figure out what do these UIs of the future look like? What is the creative control? How do you bring everything together? We see a lot of people in the creative field work across LLM's image, video, and music.
39:14in a way where they have to go to four separate tools to be able to do that. So like a lot of people ideate with LLMs, right? Like give me some concepts, like use an idea that I have. Once you're happy with that, you take it to an image model. You start to think about where are the key frames that I want to have in my video. You'd spend a lot of time iterating there. Then you take it to a video model, which is yet another surface. And then at some point you want to have sound and music and mix it all together. And then you actually want to do maybe some heavy handed editing and you go to some of the traditional software tools, that feels like these kind of workflow-based tools are probably going to spin up for a lot of different verticals.
39:51So creative activity is just one example of it, but maybe there might be one for consultants so that you can more efficiently make slide decks and presentation and pitch decks to clients. And so I think there's a lot of opportunity there that some of the big companies may not go into. yeah there's a lot of like how do we make this technology useful for x workflow right like sales finance like i'm saying a lot of things i don't know about in companies like financial workflows but i imagine there's like a lot of tasks that could be automated could be made much more efficient yeah um and i think startups are in a good position to really like go understand the specific client use case need that niche need and and do that application layer right versus what we really focus on is the fundamental technology um i think i'm just really excited by the number of people who've been excited by this model yeah if that makes sense like a lot of people in my life like a lot of aunts uncles my parents like friends like they've used chatbots they ask it things they get information my mom loves to ask chatbots health about health information but there's something about like visual media that really excites people that it's like the fun thing but it's not just fun it's exciting it's intuitive the visual space is so much of how we as humans experience uh experience life that i think i've loved how much it's moved people like emotionally in excitement wise like i think that's been the most exciting part of this for me my kids love it yeah he uh my my three-year-old son tied our dog leash which is this like fraying you know brown rope like over himself so he looked like a warrior i took a picture of him and turned him into this warrior super yeah exactly it makes him feel superhuman yeah and And my husband will read.
41:52So he uses Google Storybook to read him these stories about lessons that he learned in school. You know, if he if there was like an incident on the playground with another kid or adjusting to a new school. And it's made I mean, it's made these characters that look like him and my husband and me and our dog and our daughter. And these fun stories and lessons that we're trying to teach him to the personalization that you talked about. So I really, really love this feature. It's going to be totally different from growing up. And it's awesome, right? Because this is a story for, you know, one or five people that you would have never had made.
42:29Right. Like and other people probably don't want to read it. I would love to if you want. Yeah. But I think we're really now making it possible to like tell stories that you never could. And in a way where like the camera allowed anyone to capture reality when it became very accessible, you're kind of capturing people's imagination. Like you're giving them the tools to be able to like get the stuff that's in their brain out on paper visually in a way that they just couldn't before because they didn't have the tools or they didn't have the knowledge of the tools. Like that's been really awesome.
43:02That's a nice way to put it. Thank you so much. Thank you for having us. It was awesome to have you.
43:20Thank you.
From the publisher
When Google launched Nano Banana, it instantly became a global phenomenon, introducing an image model that finally made it possible for people to see themselves in AI-generated worlds. In this episode, Nicole Brichtova and Hansa Srinivasan, the product and engineering leads behind Nano Banana, share the story behind the model’s creation and what it means for the future of visual AI.
Nicole and Hansa discuss how they achieved breakthrough character consistency, why human evaluation remains critical for models that aim to feel right, and how “fun” became a gateway to utility. They explain the craft behind Gemini’s multimodal design, the obsession with data quality that powered Nano Banana’s realism, and how user creativity continues to push the technology in unexpected directions—from personal storytelling to education and professional design. The conversation explores what comes next in visual AI, why accessibility and imagination must evolve together, and how the tools we build can help people capture not just reality but possibility.
Hosted by: Stephanie Zhan and Pat Grady, Sequoia Capital




