In short
Podcast Notes: Google AI - Release Notes
Episode Title
Gemini 3: Launch Day Reactions
Episode Overview In this episode, host Logan Kilpatrick discusses the launch of Gemini 3 with key team members Tulsi Doshi and Josh Woodward. They explore the capabilities of the new AI model, its applications, feedback from users, and the iterative process behind its development.
---
Key Topics Discussed
- Introduction to Gemini 3
- Launch Excitement: The team expresses enthusiasm about the release of Gemini 3, highlighting it as a significant advancement in AI capabilities.
- Key Features:
- Enhanced multimodal understanding (text, image, video).
- Improved reasoning and output richness.
- High performance in agentic tasks and coding.
- Multimodal Understanding
- Capabilities: Gemini 3 excels in various modalities, enabling users to generate complex outputs from simple prompts.
- Real-World Applications:
- Video and image understanding for enhanced user interaction.
- Potential for innovative applications like coding and creative tasks.
- User Feedback and Iterative Development
- Collaboration with Users: Extensive feedback loops have been established with users, allowing the team to refine the model based on real-world usage.
- Challenges: Balancing user needs across different products (e.g., search, AI mode) while maintaining model performance.
- Balancing Speed and Quality
- Development Philosophy: An emphasis on "relentless shipping" to get user feedback as quickly as possible, leading to iterative improvements.
- Quality Benchmarks: The model aims to not only perform well on technical benchmarks but also provide tangible benefits to users in real-world scenarios.
- Generative Interfaces
- Innovative UI Features: Gemini 3 introduces generative interfaces that allow for dynamic, interactive content creation based on user prompts.
- Examples: Interactive widgets for educational tools (e.g., demonstrating sorting algorithms) and dynamic content generation for creative tasks.
- Agentic Capabilities
- Personal Assistant Features: The model can act as an agent, helping manage tasks, generate to-do lists from emails, and facilitate organization.
- Proactive AI: Future features aim to enhance proactivity, helping users manage tasks more effectively.
- Compute Demand and Scalability
- Challenges with Demand: The team discusses the substantial demand for computing resources required to support Gemini 3 and other models.
- Strategies for Management: Creative solutions are being explored to optimize resource allocation across various applications.
- Future Developments
- Upcoming Models: Plans for the Gemini 3 family, including smaller models (e.g., Gemini 3 Flash), are in the works to cater to different user needs.
- Continuous Improvement: The iterative approach allows for learning and adaptation based on user interactions, paving the way for future advancements.
---
Key Takeaways
- Gemini 3 represents a significant step forward in AI capability, combining multimodal understanding with user-centric design.
- User feedback is crucial to the development process, helping shape practical applications of the model.
- Proactive features and generative interfaces are set to transform how users interact with the AI, making it more personalized and functional.
- Balancing compute demand remains a challenge that the team is actively addressing through innovative strategies and continuous optimization.
---
Conclusion This episode provided an insightful look at the launch of Gemini 3, detailing its capabilities, user feedback mechanisms, and the team’s vision for future developments. The discussions highlight the transformative potential of AI in various applications, emphasizing a user-centric approach to development.
For further details, you can watch the episode on [YouTube](https://www.youtube.com/watch?v=mci0f2dy7G0).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:07Logan Kilpatrick I'm on the Google DeepMind team. Welcome back to Release Notes. Today, we're joined by Tulsi Doshi, who's the product lead for actually more than just Gemini models, all models now inside of... Yeah, those are Gen AI. Gen AI models inside of DeepMind. And Josh Woodward, who leads product for AI Studio, Gemini app, and Google Lab to the whole suite of DeepMind's and actually more than DeepMind's product services. Tulsi, Josh, we're sitting here because of Gemini 3.0, and we're sort of entering the Gemini 3 era. do you want to give us the sort of high level of the exciting headlines for this model launch?
0:42Yeah, we're super excited about this model launch. So we're releasing Gemini 3 Preview. We're releasing it in a wide range of surfaces, and Josh can talk more about what that looks like. But I think the reason we're doing that is because we're seeing this model really be this step change in terms of what you can do with it and what you can build. And so a couple of things that we're really excited about. I mean, one, the model is just state of the art when you think about reasoning and its depth and its nuance, its ability to kind of take simple prompts and actually turn that into extremely rich outputs.
1:19So that's, I think, one starting point. It's multimodal understanding is also incredible. So I think one of the things I'm really excited about is its video understanding, its image understanding, its ability to actually go deeper between the context of your text and the images, and then combine that with its abilities in coding, for example, to actually create just really cool outcomes, I think is awesome. right that ties then to like it's an amazing vibe coding model it's also our best model yet for agentic use cases and tools and agentic coding so there's like a whole step change we're bringing in that regard and then I think lastly one of the things we always look at with Gemini is we're trying to build well-rounded models so we want to build models that are actually really smart which this model is but also models that we think are going to be really good for users in our products.
2:13And so this model crosses the 1500 ELO on LM Marina, which I think is it's 1501, which is awesome. And that one matters. That one matters. Matters. And but I think like we use that as kind of a measure of well-roundedness of like the the usability, the accessibility of the model from like a stylistic standpoint. And I think that's something we're also seeing in some of our testing in the products as well. And so I think the combination of all of those things means this model is going to be really awesome for this tagline we're using, bringing anything to life, because you actually can now take these different modalities and formats and use them to build really cool things.
2:57Yeah, I mean, I think I just add, there's one thing for a model to be awesome on the leaderboards, which this one is, but when you get your hands on it, it's going to be, you're just going to love it. Yeah, it's really good. We're putting it in the most surfaces we've ever done on day one. So we're going to basically write crackadon in the morning, get up, put it in the Gemini app. It's going to go into AI mode. That's for all the sort of consumer use cases. Developers will be able to get it at the same time. AI studio, it's going into Vertex. There's a new sort of coding product we're putting out called Google anti-gravity.
3:31It'll be there too. So it's just going to be everywhere and we can't wait for people to try it out, give us feedback. Yeah. And I think actually like on the, where it's going to be everywhere. One of the cool things with Gemini is this time we've actually partnered really closely with a lot of customers externally. So I think this is one of those times where we actually have a lot of partners outside of Google also who are going to be shipping the model tomorrow and who have been working with us to give us feedback and help us make the model better. And so that's also exciting. It's not just that our products are going to be scaling this model.
4:01It's hopefully broadly we're going to be scaling this model, which is really cool. Yeah, I love the meet the world, meet developers in my world, but like meet the world where they are and like all these products and like people don't just use Google products. So putting the models out there and getting them in their hands is super important. There's an interesting thread actually on this product story. And we were just talking to Cori before this about sort of like the partnership between DeepMind and like the PAs and all this stuff. And I'm curious, Tulsi and actually Josh as well from like a DeepMind collaboration, like how much as we simship across the services, like, you know, the perspective of those products matters more and sort of influences the models.
4:40And then also like, is that like a net, is it like just harder? Is it more constrained? So like, does that actually help us in like, Hey, we need this model to be really great for search. And that's pushing us in this dimension, which then, you know, skews the model capability in some positive way that we don't, or is it really just like more juggling of all the constraints? I think it's both things. I mean, I will say, Josh, tell me I'm wrong. I think it's been the strongest partnership we've had in terms of getting a model into the hands of the product, in terms of the feedback loops that we've had with the product teams.
5:11Because I think now, you know, we have actually a user base who is using our models regularly that we can get feedback from, that we can run live experiments, that we can actually, you know, solicit feedback from. And that is actually really huge, right? Because then we can bring that feedback back into the model. And I think you're right, like there are constraints, right? So the kinds of things that developers are looking for may not be the same things that are important for the app, which are also not necessarily the same things that are important for AI mode. And so it is a bit of a tension game we're playing between all of these moving pieces.
5:46And I think that's hard from a model side. But actually, like, as a PM and as someone who's like trying to think about this holistically, it's awesome because we're now getting the feedback to be able to actually have real conversations about those tradeoffs. Right. To be able to say, OK, the model is now going to do this really well, but at this cost are we okay with that and and what are the trade-offs um and i think that's been really good and i think this time around we've had a bunch of hard conversations actually i don't know if we can share any of them but like is there a specific like i was in a code oh i've got one okay so um some of the viewers may have seen we've been testing this model kind of in the wild silently and um sometimes not so silent uh yeah go to canvas and it's just there one morning.
6:30Oh, that was a mistake. But you kind of in an early rev of the model, what we were really interested in kind of trying to optimize some aspects of the persona. And I remember there was one rev that we really liked on that one dimension, but it regressed on a whole bunch of other dimensions. And so that'd be an example where we've got really positive feedback, but like Tulsi said, we're trying to make this a well-rounded model. And so if it's great personality, but it's like forgets how to do tool calls, for example, Not so good. Yeah, yeah, so these are kind of the kind of revs to revs that we look at together.
7:04And I would say, I think one of the things we're really excited for people to try it out is because I think we're now getting the product model feedback loops working really well. And so when people are giving feedback in any of these products, it's now just like seamlessly starting to like flow back in. And that's really exciting too, because you can see the quality gains happen when we do these kind of blind tests side by side. And it's kind of awesome because even something like we say tool use as like a broad phrase, but even the way that tools are being used in the Gemini app is a very different construct than like how a developer might use tools in a very different context or product or like how anti-gravity expects tool use to work.
7:44And so I think now that we actually have these different product surfaces that we're getting feedback from, we're actually able to see those differences and those tradeoffs. And I think that also allows us to actually have a much more holistic understanding of what do we mean when we say tool use or what do we mean when we say code performance or what do we mean when we say, you know, creative writing abilities or something like that. Yeah, I feel like the product scaffolding sort of grounds a bunch of those capability conversations in a way that is like really, you have to like proxy with some of the academic benchmarks, which are oftentimes like actually not representative of how users actually use the models in certain cases, which is tough.
8:22And you both alluded to this. I'm curious when, and this was sort of, you know, I get all the tweets of why are we not shipping Gemini 3 earlier, which I'm going to start just sending them to you too. Please. and sort of and like I think it gets out a broader conversation of like when uh obviously Gemini 3 is out the door now at the time of this conversation uh being released but like when was the moment where we knew like hey this is the model we want to get out the door and obviously we've been doing tons of this iteration but um yeah is there a way we can ship faster is there is there why not ship earlier like in the past we've done all these like crazy experimental checkpoints which we've gotten in terms of feedback, both positive and negative in a lot of cases.
9:05So I'm curious what the balance - The calculus was there? Yeah. I mean, we do want to ship fast. And I think that's actually super important to keep alive. Like I think this whole idea of relentless shipping, put things in the hands of users as soon as possible, get feedback as soon as possible. I think that we want that to continue to be the theme. And I think we want that to be the theme across all of our models, not even just like Gemini Pro, for example. I think in terms of this model, in particular, one of the things we really wanted to focus on was it's kind of real world usability. So real world usability, both in our products, but also with developers.
9:42And that means actually getting a lot of feedback and actually iterating on that feedback. And so we spent, you know, a period of time putting the product, putting the model in the product to Josh's point about the persona discussion we had. We put the model in the hands of customers. And then we basically used that to also help us calibrate, okay, where is the model meeting our expectations where like it matches what we are seeing internally? Where is the model not necessarily meeting those expectations and how might we want to adapt it? And so we had kind of quality goals and bars for ourselves that we wanted to meet on our benchmarks, but also on like how it would work in the products, what kinds of, you know, experiences we wanted to create with the model.
10:24And so it was all kind of working towards that. I still think the goal is like ship it as soon as possible within the constraints of of that realm of things but it's about like making sure that it hits those those goals and i think now from the previous models we've set we're actually setting more complex goals for ourselves um so every time we release one of these models we're we're setting the bar higher for ourselves that also means that meeting the that bar is harder i think yeah yeah yeah i mean the other thing i would add i feel like the speed is is great because when a model comes we measure it in like days and hours until it's actually ready.
10:59So I guess we could get to minutes. Maybe that's the goal. But I mean, I think some of the turnaround on this is just amazing. And the teams working across either the developer side or the consumer app side, the modeling side, obviously, it really is. You're seeing, I think, out in the world, the momentum, like the compounding. And so I think it'll get faster. But at this point, we're kind of like days, hours, and maybe we go to hours, minutes. That's the next jump. For both of you, what was that moment, and Josh, maybe you start with this one, that you tried one of the Gemini 3 checkpoints and sort of the use case or the demo that clicked, was there any of them that you sort of saw this and you're like, this is clearly a step change, and I'm happy to give mine as well.
11:43Yeah, we should go around. I mean, I've had a few of these. I think this one's really good at vibe coding, it's very good at web dev, and you can just describe something that's in your head, hit a button and it's there. And that is kind of a crazy sort of compression of time and skill and expertise. So that was one. I'm really excited about the multimodal capabilities too. We're doing some stuff where kind of both on the Gemini side and the Google lab side, we're really caught up in this idea of how you can transform content. And this model is very good at that, better than I've ever seen. So those would be two off the top of my head.
12:18What are some examples of the transforming content piece? Like is this? Yeah. Yeah. I mean say you're a student You've got a whole bunch of like video lectures handwritten notes any of that you just load it in hit a button It goes anything around like I'm very excited about also kind of visual formats We're starting to play around a lot with that There's a Experimental feature we're putting in the Gemini app where it'll literally you can sort of describe something It'll make an interactive website for you on the fly. That's crazy. And that's not canvas though. That's it's bit Yeah, it's kind of built its own little thing.
12:50And I think to me, that's a whole new frontier with these models. We're calling it generative interfaces on the team. But this idea that just pixels are streamed in as they're ready based on your prompt. And I think that's a whole leap. We've never been able to do that before. So that's now possible. Still a little slow right now. That's why we have it as a labs feature. But I hope people play around with it a lot because we're doing experiments where it's like, give like the history of van Gogh's life by period with example artwork and it'll just build a whole timeline and just let you go through it and it's all just generated um pretty amazing that is that is crazy yeah we'll see three examples from my side one is I do think uh the vibe coding piece really does hit um and I think it does because the visuals are so rich and the interactivity is so rich um and so I was like I was using it literally to create like a fun game to play with my niece with like a single prompt, what you can get, the like her ability to actually engage with it.
13:47I think that kind of fidelity of the experience, we haven't been able to see before. And that I think is really hits home. Are you using AI Studio? I'm using AI Studio. Of course. Of course. Fear not. The second one though, that I think is really interesting is actually one of the things we don't talk as much about with this model is its multilinguality is also really good um so its ability to write in like hindi or actually in gujarati which is like not as common a language but is like from the western part of india which is what my family speaks uh that is also really strong and i think that was another wow moment for me which is like oh this model is really good yeah um and it's actually able to do something that i personally uh struggle with you know in terms of daily life that i would actually really love to see kind of play out.
14:38And then the third is my husband is a very avid pickleball player. But the like video understanding of the model and its ability to kind of break down a video and like break down a set of moves and actually like give you critical feedback. I think it's like the like all three of these examples are basically combinations of the model being able to understand something that maybe you by yourself either struggle to do or it's harder to do or just requires a lot more focus and attention and the model is able to pull out some of those nuances and actually like help you with them in a much like richer kind of way um and i yeah i think it's just been awesome also i will say like the vibes of the model um even in terms of like the uh the style in which it responds um it feels different than other models it feels smart yeah um and i think that also stands out when you when you engage with it i love i think the one other thing I would add to, you said it, is like to me we're getting in a point in Gemini 3 where it's like the combination of a lot of these things is really starting to show and differentiate.
15:45So like when you go back, Gemini 1, long context, Gemini 2, more modalities, some of the coding started to come in, 2.5, even more coding. Now we're kind of in like agent tool use. It's all starting to kind of almost like ladder and sort of build on itself. And I think that's what we're starting to feel, at least I'm feeling on the team, is like there aren't these little isolated model features anymore, it's actually like the combination that starts to get really interesting. And I think we saw this even with Nano Banana a few months ago where you kind of take something as like put it together.
16:16So to me this is like, Gemini 3 gives us a huge foundation of sort of things we can combine in new ways. And that's where you get like multimodal and vibe coding, suddenly you've got some really interesting interactive thing that just before wasn't really possible or that easy to do. Or there's this cool demo someone on the team built, and I think we're actually using it in the materials, where they took a picture of a handwritten recipe in Korean and then turned it into like a full vibe coded, like family recipe app in English with the measurements where you could adapt them for like different switch.
16:49you could switch um and like again that's a great example of like the combination of like it has strong multilingual capabilities it has strong multimodal understanding so you can literally take a picture of a handwritten recipe in korean you can then get a fully fledged like interactive app on the other side that you can actually use in your day-to-day life like that combination of things is really cool yeah i feel like this uh this gives so much credence to the bring anything to life narrative i'm like maybe that's the gemini long term like do we have a gemini model slogan like i know there's a gemini app like the three p's yeah yeah all the research to reality no you're talking to the person do we have one because maybe we should have right now bring anything to life is the new gem the new the new we are i mean for gemini three i think that is for all practical purposes the slogan and i think what's cool is you talk about that example with the recipes and handwritten in Korean and everything.
17:44In some ways it's mind blowing, but in other ways, it's like, yeah, this is how it should work. Think about kind of the boundaries and sort of human computer interaction. I think German i3 like blurs those even more in a way we've never done before, which is cool. Yeah. What was yours? Oh, yeah, that's a good question. Well, mine was, Josh, we were in Arizona and I spent like all day making demos with 2.5 Pro And I was like gosh 2.5 Pro is such a good model And I was like wait I should probably try these with the new checkpoints And then it was just like I was like okay this thing's incredible like especially like it was like very Front and center in my mind like this was the limitation of 2.5 There was these cases where I which is and then I could just like see automatically and then I had a Interaction where I was DMing with someone and they were like hey try to make this random game and there was like a RPG sort of like fishing game and it just like one-shotted it and then I went and uh you know Demis worked on some of the original sim roller coaster type no not roller coaster sim city 1984 or something like that sort of tried to remake a game like that and it just worked uh using one of the earlier checkpoints and then again with with some of the the most recent checkpoint and it's just it's it blows my mind I think the video game example feels so pertinent because there's so many people who like learned it feels top of mind for me because there's so many people who tried to learn computer science with the end goal of like building games because it's just fun interactive thing like building games actually sucks and it's not fun in a lot of cases um and i feel like gemini 3 has this power to like help people bring the thing that they want to life uh you know what's actually awesome is we use gemini 3 so youtube has this notion of playables yeah yeah right where you can go to youtube and play a game and we actually used gemini 3 well by we I mean D-Shin who's awesome and does incredible things um he actually used Gemini 3 to build a bunch of YouTube playables that we actually are are going live uh on Tuesday with the model um and they're awesome and they're like so fun to play and I think like just the idea that you can create this like vast set of creators who can actually build these types of interactive experiences is really awesome but yeah I think it was after your Arizona trip where you you pinged and you were like, we gotta ship this.
20:02We gotta ship this. It's awesome. But it felt like it. I think if you spend time with the models all the time, it's just so clear that 3.0 is, or 3 is just a step change. So it's exciting. Did the SVG art people influence our decision on social checkpoints to use? Actually, it's really funny because we did, after seeing some of that feedback come through, we legitimately like maybe a week or two ago like sat down and had a bunch of people try across our different checkpoints svg art just because we were like is it really that different checkpoint to checkpoint like are we really seeing one checkpoint be like so much better at this than another based on kind of the the twitter feedback yeah um is there a good benchmark for svg art no not really no um but i think the other thing is like it is a very narrow sliver of how the model operates right um and i think what's interesting is also one whether something is cool from an svg art perspective is like very much up to some interpretation right um two i think what we've found is that some are really good at certain types of things and then others are like other checkpoints are really good at other things but it's not like universal like all svg art one is really good and then the other one was really not good kind of thing um so it's very like mixed feedback in terms of what we were also getting from from people i i had to go do a bunch of research to learn more about SVG art because people were showing some of the demos and I was like I had no idea that SVGs could move.
21:33I was like they're like showing like fully animated sequences and I was like I thought SVG yeah the fan I thought SVGs were just static images but they're like the programmatic representation of them which is also the thing that is most surprising to me is like I get that it's like I assume it's like the underlying code capability that as it improves it sort of pushes sort of SVG ability but then you look at what is actually being outputted and it's like the SVG representation is actually just a bunch of like random numbers and letters which I'm like is crazy that that that the model is actually able to like pick up and learn like I get I can kind of intuitively understand how the model learns to code because it's English basically and a representation of this structured thinking but I mean this is kind of two in some ways right it's just more I think like in it also reflects the model's ability to like reason about how something will manifest in two dimensions yeah right um because it's able to actually take the concept right so like oh i want to generate a span and i want it to look realistic and things like that and then actually translate that into those numbers and letters right that actually will reflect the colors and the points in like an x and y axis buttons to turn it up and down right they get fully interactive it's cool it is cool are there any other examples of like things that we saw from early our internal tests our external testing that like sort of we all of a sudden we're like hey we actually maybe do want to track that because that would be an interesting use case that um i think svgr was one of those and i'm trying like the voxel stuff has been around for a while stuff yeah i'm trying to think we've been tracking a lot on the app side a lot and there's not an official benchmark for this either but like emoji usage, how concise our model is.
23:20So sometimes we've been a long-winded Gemini model. You'll notice with three that we're like much tighter. It tends to be, it picks up on the subtleties and sort of the way it writes. So there's a whole set of kind of persona style, kind of house style things that we've definitely been monitoring with the various checkpoints that viewers have probably been seeing out there. So that's definitely something we've been looking at. Yeah. Yeah. It's also been interesting because persona and like style is so subjective. And so that's one where I think feedback actually really does make a huge difference because my perspective on what good persona is is probably very different than both of yours in terms of like what I want to interact with, right?
23:59And so I think that's one that like really stands out. But that one we were also tracking even before like social feedback, right? It's one that we were kind of really trying to monitor ourselves and kind of keep a pulse on. uh games is something we knew was exciting to your point like we knew that like games is a rich thing that you can build but i think when you put these models out and you start to see what people do uh you start to really see the magic of what people in the rain create and the age of what people can create and then that i think inspires a set of like demos but it also inspires a set of like just people to go start exploring and building something and i think that also is like a really nice feedback loop too yeah i have a meta question which is i think the last time all three of us were sitting down for this type of we need to sit down more and do these more often but was io a few like yeah yeah six months ago which is crazy um and to see like it is very my like meta step back take is like it's just crazy how the model capabilities continue to improve and it just like dwarfs like it just makes like 2.5 when 2.5 launched like you know it had its challenges and hopefully we addressed a bunch of the things that people had feedback on and sort of built on on the things that people really liked about 2.5 but um 3.0 just makes it look so it's like it makes it look like a goofy toy example and it's just kind of weird and i'm curious like how both of you think about that as you take a step back but also like you know when we make the next like it's just crazy to me to project forward like we'll make another step change that will make this model now that we're sitting here so impressed by and hopefully everyone else is impressed by feel like kind of a goofy toy uh and you know six months or something like that yeah no pressure i mean i do i think about this a lot people talk about ai time and how like six months of normal time translated to ai time but it is crazy what we were out like outside shoreline amphitheater chatting like this yeah and we were it was a very good mod it still is very good it's like you know incredibly 2.5 pro is like incredible um and here we are it's like every week we're shipping you know so in some ways that's why it probably does feel like forever ago yeah i think it's also interesting because like we were talking about this too is it puts actually more pressure is not the right word because that like sounds negative i think but it it puts more intentionality on like the model product story together because like as you have these models that can do cooler and cooler and cooler things.
26:31Yeah, you can talk about that through the benchmarks. And, you know, we obviously want the developer community to explore things and they will come up with amazing use cases as always. But it also puts, I think, a lot of pressure on us to come up with like, how can we really take advantage of the strength of this model to really build amazing things for users? And it's interesting coming from like, you know, you typically talk about you want to start with where the user is and the user's needs and then work backwards to build a product. you obviously still want to do that. You still want to focus on where users can get a lot of value, but you also want to also start from what the model can do and what kind of unique capabilities it has.
27:09And then what can that empower for users? Yeah. Both ways. Both ways. And I think like that was actually an easier conversation to have when the models were a little less capable because the range of what they could do was smaller. And so they kind of fit into your like, okay, yeah, these are like clearly things we want to build. Yeah. And now I think it's also pushing for more creativity of like, okay, now that the model has this range, what can you build? Gosh, that's a great question. We talked about vibe coding stuff, but there's a bunch of new, it's not just Gemini 3 sticking it into the existing Gemini app.
27:42There's a bunch of new product stuff, so maybe we actually talk about that as this through line. Yeah, I mean, there's a few things coming I'm excited about. One, across Gemini and Search AI mode, there's going to be a whole wave of kind of generative interfaces. Generative UI. And just to clarify, how is this different for folks who have seen Vibe coding stuff? Is that similar? So the model is generating code and then there's function calling behind the scenes? Or can you give a little bit more detail of the user who hasn't seen this before, how does it work, what should they expect? Yeah, so the shift we're starting to see, and this is really the frontier, the beginning of it, is in the past, some engineer here in the building would have coded the UI.
28:21It would have been one way for all users until it got changed. And what we're seeing now with Gemini 3 is it can actually make a lot of those design and implementation choices on its own. It's a big shift. So you may do a query like, hey, I'm planning an upcoming trip. Give me a three-day itinerary. Up until Gemini 3, we would have responded with a wall of text with some pictures. Now the model is going to respond obviously more concise like we talked about, but it's also going to lay out the page. and so at a practical level sometimes we do that where it's actually saying here's a tool call give me this tool that's going to lay out a I don't know carousel I can scroll through other times what it's going to do is actually almost act like a design agent and be like I'm going to go through a design process and think about what's the best way to present a three-day trip to Rome and so it may lay out a table it may give you kind of something almost looks like a magazine style layout yeah and all of this we're kind of experimenting right now like how much can we sort of trust the model to lay this stuff out?
29:21So it's super cool because you're giving the model widgets, a style sheet, different tools it can use, and you're kind of saying, go wild model. And so I think for a lot of people, what they're going to feel is, oh, this is a much more visual, kind of immersive, interactive, but also kind of customized to me sort of view. And I think that's what's exciting is like, it's just the beginning of that. And so I think if we look ahead throughout the Gemini 3 series as well as 3.5 and beyond this idea of a models being able to kind of compose more and more things feels like an interesting theme to kind of explore it's also like you know Josh mentioned widgets like I have this like particularly fun example from for myself at least on AI mode which is you know if you ask in AI mode with Gemini 3 about bubble sort right which is like just a very sort of standard notion of how you can sort uh you know elements in a in a series um it actually creates this like interactive widget that like you can play with and see like how things actually get ordered and how bubble sort works the real question is why were you googling how to yeah you thought you'd do bubble so you're preparing for like a coding interview somewhere but like show us how bubble die no but but i think the reason it's fun is because I like have this like it's you know um credit to to Madhvi who is um you know one of our co-workers who was showing me this example but it's like it's kind of a blast in the past because if you remember like when you were learning about these things you're learning about them from like static like textbooks right and you're trying to like take the written notation and like map it and be like okay like I only try to visualize it like let me try to write it out like okay now it makes sense versus being able to actually see it yeah and play with it and reflect on that in the context of how you search like it really feels like we're bringing the google mission to life you know the like make information what is it universally accessible and useful yeah like it's that kind of gemini 3 brings the google mission to life not only not only but even the google mission is coming to life that's awesome um josh other things genui in AI mode and the Gemini app, other things in Gemini app?
31:35We're also trying to explore some new features around how Gemini can act as an agent for you. So you've seen probably over the last month or so, we've put out a new API for computer use, which is really cool. This Gemini 3 model takes it kind of one level, even beyond how we think about multi-step actions, being able to kind of take sort of different tool calls and just do stuff for you. So the Gemini app's gonna have an agent, kind of experimental feature. It's interesting, you can go in and say things like, create to-do lists out of everything in my inbox, if you're behind like me on my inbox, and it'll sort of give you this bird's eye view and be like, let you add stuff directly to Google Calendar.
32:14So yeah, what's been your personal feel? Like I have a very high bar for this inbox use case and what's your sense is like, it's working, it's hitting it out of the park, or like what's the, where should my expectation be for how good it is at doing some of these inbox tasks? I think you should give it a try. See what you think. I think that one is actually approaching, like it's starting to get pretty useful for tasks, I would say. The other one we're seeing a lot is helping people do research across a lot of disparate things. So you could do this through a deep research path and get a report, or you can try this agent and it'll actually start trying to do things for you.
32:47And if it's not sure, it'll actually pause and say, Hey Logan, before I send this message or before I hit click this, click this buy button, do you want me to do this? So I think that'll roll out to all of our Ultraplan members on launch day, which we're excited. And we'll kind of take it from there as we go. It builds on a lot of the stuff we've worked on together, going all the way back to Project Mariner almost a year ago. That was some of our first forays into this. And there's still ways to go, but we're really excited. This is one where we're really excited to see the feedback. Is that how the Inbox one is working as well?
33:18it's like actually like you it's like actuating a browser and it's doing or is that like using a bunch of like tool use to send requests to like the gmail api or whatever or that yeah so in this case if you go into gemini and connect gmail which a lot of people do yeah it'll start doing kind of tool use and api calls in the background and then gemini 3 kind of oversees that and kind of orchestrates it for you um so there'll be a lot there i think that's like exciting to keep exploring as we kind of push forward ahead too. Yeah, I love that. Any proactivity stuff. I know just we talked about personal and powerful implicitly so far, but is there anything like on the horizon from a proactivity standpoint, which is my, this is my number one feature request.
Read the full transcript
33:58If I can get the Gemini app to just look at my email for me and look at and see all the tasks I need to be doing, that would be my ideal case because I don't want, it's just too stressful to look at all my. Overwhelming. I'm like, I already have chats and text messages. Shared pain here. I guess I'd say watch this space. Okay. Yeah, we're very interested in this. Right now, we're starting to see a lot of people use kind of scheduled actions in Gemini as kind of like a proactive way to engage. But we think there's a lot to do here. And I think some of the things we've talked about combining different modalities, letting people hook up different Google apps, kind of bring some of that together to orchestrate.
34:38Stay tuned. I'm excited. I will be staying tuned. I love when people say watch this space because I'm like where where do I watch this thing gemini.google.com gemini.google.com perfect it feels like one of the big constraints of like actually bringing the model and letting the world experience how it can bring anything to life um is like we need crazy amounts of compute uh because the demand is so it's not even like we don't have lots of compute it's just the demand curve is yeah it's like this but you know you can't build physical compute that fast six months um so i'm curious like for the gemini app and maybe more broadly across other products in in your scope like what how are you thinking about the balance from a like also from a user perspective of like where we're allocating compute because we think there's value for different use cases yeah it's super hard i would say like uh we could share some of our stories like the friday before launch uh i mean i don't know what's harder solving agi or solving the puzzle of compute and try to like, but I think like what makes it hard and you hit on this Logan is like, there's just demand everywhere.
35:42And so that's one thing. Well, I think we're making products now that people want and they are coming back to them and using them. I mean, even just in the Gemini app, our daily requests has tripled in the last quarter. That's correct. And so, and I know we've seen similar stuff on AI studio and developers and just like, so when you're streaming that many tokens, you really have to try to think through it. I mean, one of the things we spent a lot of time up here on Friday was like trying to think about, okay, where are the places where a lot of people can experience kind of the great stuff in Gemini 3?
36:13So there's going to be some consumer apps. There's going to be some developer stuff. There's going to be some customers. And so we try to think about that. Then we try to get very creative. So like there's a great kind of tool that came out of Google Labs called Flow. Helps you make sort of videos. And that was one we were like, oh, there's just not enough. And we're like, no, but there has to be because this product is like people love this product. And so we were able to find some crazy deal over the weekend where they're going to convert some of their chips to a different kind of TPU to get another TPU free.
36:41So, you know, you go to all these ways to try to, like, make it work. And we'll see. It's launch day. So we'll see how well our estimates played out. But we have a crazy Google sheet where there's all these estimates, like moving things around. Yeah, I think it's also interesting because one of the conversations we tried to have as a part of this kind of planning exercise. It's like, okay, what is a P0 experience we really want to ship with the model? What is P1? What is P2? And what's hard is it really does feel like you're comparing apples and oranges because part of the goal is to create like a wide range of experiences, right?
37:12Like Notebook LM is a very different and very compelling product than the Gemini app, which is also a very different product than Vibe Coding on AI Studio. And so, you know, when you're now sitting down and looking at these products and saying, okay, where should we be bringing Gemini 3 to life? You're talking about very different versions of the products and you actually like Gemini 3 will manifest in very different ways in these products. And so actually if you really want to showcase the breadth of what Gemini 3 can do, you actually want to ship across that breadth, right? And so then you're like, okay, well, how do I think about like, how do I think about how the math works out across all these things?
37:49And the real problem is we keep making great models across all of these dimensions. and it's like i do think there's like something really interesting where it's um you know we had like vo was crushing it and then it was like where you know we couldn't get enough compute for vo and you know we made it happen and then nano banana happened and it was the same thing as like we couldn't get enough compute there's so much and then gemini 3 is happening it's gonna be the same story and the interesting thing and there's only more coming and there's only more coming and it sets this new floor and i'm like uh yeah it's crazy it's crazy to think about how much demand there is for this stuff.
38:21Yeah. I mean, this is where all the efficiencies matter. So there's like an unsung cast of heroes that are making the models, you know, two, three, four times sort of more efficient. I think the other thing you kind of have to just step back and realize probably the most consistent way to get good models is to sign up for one of our subscription plans. I know this is not a commercial. This is not a commercial. We do give the highest rate limits and we really do try to kind of people that are willing to to pay are going to get that too and if you're a student uh that's the other part if you're a student a actually launch day if you're a university student in the u.s yeah go sign up google ai pro for free for a year it's awesome that's 20 a month times 12 free uh it's totally worth it by the way not just for this model but because there's just a lot coming for yeah that's right you'll get this model you'll not only get the highest rate limits kind of in Gemini app, you'll get access to notebook, LM, flow, kind of all the good stuff coming.
39:17A lot of our developer products too. So it's exciting. It's crazy. It is a good deal. Storage. I mean, now this is just sponsored by Google AI Pro subscription. That's awesome. What any, any sort of, you know, obviously we're starting with, which is interesting to sort of, and I don't know if there's a story here, Tulsi, that we can tell, But last December we shipped Flash, Gemini 2.0 Flash, and it was awesome. And it was a frontier sort of state-of-the-art model. Now we're shipping Gemini 3.0 Pro or 3 Pro, depending who you ask. Gemini 3 Pro. Gemini 3 Pro, officially. Obviously people love, like Flash is actually what I think made Gemini popular in some sense.
40:03Like it was our workhorse model. We launched 1.5 Flash back at IO a year plus, a year and a half ago, which I think put Gemini on the map and has sort of pushed the Pareto frontier. When are we working on Gemini 3 Flash and all the other models? Do you have a sense of when we get smaller models? Yeah. This is my turn to say, watch this space. Which space? Including the app. but actually I think so we definitely want to build out the Gemini 3 family so I think pro is just the beginning of this and like we're already I think very excited about the direction we're going with flash and where we're going to be with the rest of the Gemini 3 models I think the reason for sort of shipping them in sequence instead of like bundling them all together and trying to ship them all as one package is one when you ship one of these models you get to learn a lot about how people are using them.
40:59And so you get to learn a lot about like, okay, where is 3Pro really resonating? Where did we think it was going to resonate versus where are people resonating with it? Where are we seeing people say, hey, actually, this is too expensive for my use case, or it's too slow for my use case, or I might actually need something different, which then influences what you do for Flash, right? Because Flash is meant to be the workhorse model. And we want to make sure we kind of account for those pieces when we're building that. And so So I think actually shipping them in sequence kind of allows us to like learn from one and build to the other.
41:33I think it's also part of the relentless shipping exercise. Right. If you want to ship quickly, we want to actually like put these models out there and see what people do with them and get feedback. So, yeah, but there there will be there will be some exciting stuff coming. Thank you both. This was an awesome conversation. Thank you both for the hard work of pushing to get the model and the model into all the products. I feel like folks are going to enjoy the moment. So hopefully we'll be sitting here for IO maybe sooner next year and launching other cool models. So thank you both. Thank you. Enjoy it.
42:04That's going to be great. And thanks everyone for watching Release Notes. We'll see you in the next episode.
From the publisher
Join us for a special episode of Release Notes as we unpack Gemini 3, Google’s latest AI model with key team members. Learn how Gemini 3 empowers developers with enhanced multimodal understanding, agentic capabilities for complex tasks, and generative interfaces that transform prompts into interactive applications. We discuss real-world use cases, the iterative development process driven by user feedback, and the strategic balance between model performance and broad accessibility across various Google platforms.
Watch on YouTube: https://www.youtube.com/watch?v=mci0f2dy7G0
Chapters:
00:00 - Introducing Gemini 3
03:08 - Gemini 3 everywhere
04:13 - The product-model partnership
08:20 - Balancing speed and quality
11:40 - Gemini 3 'wow' moments
27:47 - Generative interfaces and UI
31:44 - Gemini's agentic capabilities
33:55 - Proactive AI and future
34:55 - Managing compute demand
39:32 - The Gemini 3 family
41:45 - Conclusion

