In short
Google AI: Release Notes - Episode Summary
Episode Title
Launching Gemini 2.5
Episode Description In this episode, host Logan Kilpatrick sits down with Tulsee Doshi, Head of Product for Gemini Models, to discuss the launch of Gemini 2.5 Pro experimental. This multimodal thinking model is designed to handle increasingly complex tasks, showcasing improved reasoning capabilities and advanced coding abilities. The discussion explores the development process of Gemini 2.5, enhancements made, and future directions for the model.
Key Discussions
- Overview of Gemini 2.5 Launch
- Strengths of Gemini 2.5 Pro:
- Leading in reasoning capabilities and academic benchmarks.
- Excels in coding tasks, including web application and agentic code creation.
- Supports multimodal functions, including image and video understanding.
- Features a long context window, allowing for processing long documents or videos.
- The Transition to Version 2.5
- Significance of Versioning:
- The "0.5" indicates a significant advancement in model representation and performance.
- Introduction of 'thinking models' to enhance problem-solving capabilities.
- Development Process and Collaborations
- Cross-Stack Improvements:
- Integration of pre-training, post-training, and reasoning improvements.
- Collaboration among teams to ensure comprehensive enhancements across the model's functionalities.
- Safety and Review Processes
- Embedded Safety Protocols:
- Safety evaluations conducted throughout the development process.
- Teams perform 'red teaming' to identify issues and enhance model reliability.
- Multimodal Reasoning and Video Understanding
- Challenges and Innovations:
- Combining multimodal understanding, long context processing, and reasoning is crucial for video analysis.
- Example discussed: Analyzing cricket match videos for key events.
- Future Directions for Gemini
- Upcoming Features:
- Plans to scale access to the Gemini 2.5 Pro model for broader developer use.
- Enhancements in image generation and continued focus on improving model usability and dynamic thinking capabilities.
- Community Engagement and Feedback
- Emphasis on gathering developer feedback to shape future updates and capabilities.
- Exploration of how to balance performance with user engagement and experience.
Key Takeaways
- Gemini 2.5 Pro's Capabilities: The model represents a leap in reasoning and coding capabilities, setting a new standard for AI development.
- Importance of Collaboration: Successful model releases require coordinated efforts across various teams, highlighting the significance of teamwork in AI development.
- Safety as a Feature: Incorporating safety measures early in the development process enables quicker iterations and fosters innovation.
- Future Models: The continuous evolution of Gemini aims at building the most capable and user-friendly AI tools, with a focus on multimodal understanding and dynamic thinking.
Resources
- [Gemini Overview](https://goo.gle/41Yf72b)
- [Gemini 2.5 Blog Post](https://goo.gle/441SHiV)
- [Example of Game Design Skills](https://goo.gle/43vxkq1)
- [Demo: Gemini 2.5 Pro Experimental in Google AI Studio](https://goo.gle/4c5RbhE)
Episode Chapters
- 0:00 - Introduction
- 1:05 - Gemini 2.5 launch overview
- 3:19 - Academic evaluations vs. vibe checks
- 6:19 - The jump to 2.5
- 7:51 - Coordinating cross-stack improvements
- 11:48 - Role of pre/post-training vs. test-time compute
- 13:21 - Shipping Gemini 2.5
- 15:29 - Embedded safety process
- 17:28 - Multimodal reasoning with Gemini 2.5
- 18:55 - Benchmark deep dive
- 22:07 - What's next for Gemini
- 24:49 - Dynamic thinking in Gemini 2.5
- 25:37 - The team effort behind the launch
This episode provides valuable insights into the development of Gemini 2.5, the importance of safety in AI, and the collaborative efforts that drive innovation in AI technology.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:01We're here in Mountain View. Tulsi's back. We launched Gemini 2.5 Pro today. It is extremely strong in reasoning. It's an amazing coding model. It is the best model we've ever built. My head is somewhat spinning. We just shipped Gemini 2.0 originally a few months ago, and now we're already at 2.5. We're like, we really have to get this model out the door. We have to get it in the hands of developers. We have to see what people are doing with it. Shipping is a team sport. Long context, tool use, like so many of these things that we ship with our models are truly end-to-end innovations. What's coming next?
0:34I think a lot of what we're trying to do is just build the most capable models. So I'm actually really excited to see what people do with it.
1:05Hey everyone, welcome back to Release Notes. We're here in Mountain View. Tulsi's back. We launched Gemini 2.5 Pro today. It's been a ton of fun. Everyone's been sort of celebrating the state-of-the-art milestone for us. Can you sort of give us a rundown on what we launched today and why folks should be as excited as we are about Gemini 2.5 Pro? Yeah, so like you said, we launched Gemini 2.5 Pro today. Super excited. I think the really big thing about this model is that it is the best model we've ever built. And I think maybe even going further than that, I think it is one of the best models we have in the industry right now.
1:40It is extremely strong in reasoning. And so it is actually state of the art in a number of the common reasoning benchmarks. It's an amazing coding model. It's particularly good at creating really fun web applications. It's great at agentic kind of code applications. And it also actually is particularly good at code editing and transformation. So I I think it just actually makes for a really good coding partner. And so those both are really strong. And then I think also it builds on all of the great stuff we already have of Gemini Pro. It's multimodal, so it's great for video understanding and image understanding.
2:15It comes with long context. And so this one million long context window that we have, which allows you to process like long videos or documents. And so really overall, this is just an incredibly strong model. and we're just like really excited to actually put it in the hands of developers and customers and actually see what people are going to build with it. I think one other thing I'll say about Pro, which I think is really cool, is I think it's actually a really well-rounded model. So I think a lot of models, when you push for them to be really strong reasoning models, they're really smart at the benchmarks.
2:47But what's really cool about this model is it's also really good at, you know, style. And it's like a fun model to actually talk to. And I think that's also why we see it doing really well on leaderboards that are useful for like user preferences, right? So we actually have a 40-point jump in LM Arena ratings in ELO compared to the next model. And I think a lot of that comes from the kind of ability for the model not just to be really smart, but also to be, yeah, I guess well-rounded is maybe the right word. Yeah, I feel like that the balance of like passing the academic evals and the vibe test is like actually hard to do And it actually is kind of weirdly unintuitive that that's the case But I feel like it generally ends up being true from what we see like I think you're closer to this But like as you hill climb some random eval like it doesn't actually translate towards like something that users are interested in the model doing And it is this weird.
3:43Yeah, it seems like there's a weird disconnect Yeah, I think you use the right word which is like vibes of the model So, like, I was playing with the model we shipped today over the weekend, and I've been doing that, you know, with every model we train, just, like, trying some things out and seeing how it goes. What's your go-to personal benchmark for this? Yeah, so, I mean, I think I typically go through, like, three different things. One is I'll just, like, try the basic, like, hey, how are you? Like, how's it going? Kind of just, like, to see how the model, like, responds to, like, really simple, basic, like, what's going on.
4:14I'll do things like write a poem about walking on the beach where the sun is actually shining but also make sure that it references a sunset and also please make sure that it's about the month of March and like trying to see what the model does when you give it a set of very interesting instructions that kind of cover a different range of things and then what's been really fun more recently as our models have gotten better at coding is to also have the model try to make games, and so actually giving the model one-shot prompts and seeing what it can do. So I spent an unreasonable amount of time on Saturday playing Snake because I had made a web app using Gemini 2.5 Pro.
4:53I was like, it's awesome. I can play Snake. And to be clear, Snake itself is not a complex game, but what was awesome to see is being able to take a single prompt and then actually have the model build something visually aesthetic, So actually like colorful and with the right effects and actually being able to supply the JavaScript to really make the game more engaging. And I think the thing with the vibe check is you actually not only see the model doing the right thing and actually being able to like follow your instructions, but also when you look at the thoughts of the model and you actually look at the response itself, it actually feels more engaging.
5:31And I think that's important too. Yeah, I love that. And I have so many random threads about this, but let's come back to some of those examples. Because I think there's just like threads now all over the internet of like really, really cool one-shot use cases. We put out a bunch of stuff. Developers have already put out a bunch of stuff. So hopefully we'll put them in a show notes or something like that. Yeah, we should totally do that. I mean, like Jack today was talking about a great example that I think a lot of people use to vibe check models, which is, and Jack is our lead for thinking, but he talks a lot about like the ball bouncing around the square.
6:01And I love that as a vibe check prompt because it tests not only like the ability for the model to generate graphics But also like understanding physics and being able to actually like manage the physics of that And it's a very simple use case, but one that is actually like very evocative I think in what you can do Yeah, I put out a tweet exactly with the like ball bouncing use case and I used Like the prompt that I thought was the standard prompt and somebody was replying and being like hey it doesn't have gravity or something like that. And I was like, the prompt didn't mention gravity. And then I think Jack actually shared an example somewhere of with gravity and it's just like, it can do all types of crazy stuff.
6:37So yeah, hopefully we'll have a link somewhere in the show notes and folks can look at some of these examples. But we made the jump from Gemini 2.0 Pro to Gemini 2.5 Pro. And we were talking about this jump and what that means. Can you sort of talk us through like why this is so substantive, why we added that next version? I feel like my head is somewhat spinning. We just shipped Gemini 2.0 originally a few months ago, and now we're already at 2.5. I know if you believe we only shipped Gemini 2.0 three months ago. Yeah, it feels like a year ago. It feels like a year ago. Yeah, so why do we make the shift to 2.5?
7:11So I think for us, these 0.5 increments, so we shipped Gemini 1.0, then 1.5, then 2.0. I think 2.5 signifies two major things. One is a shift to what our models are going to represent going forward. So I think going forward, all Gemini models are going to be thinking models. And that's going to be a fundamental part of how the models approach problem solving. And I think that's already a big shift in thinking through the capabilities of the models and what they will represent. I think the second one is really the step change in performance. And that one is critical. If you look at the 2.0 series, that was already a step change from the 1.5 series.
7:48And here we see this significant jump. Tulsi, we were having this conversation earlier with Seb, who leads pre-training, with Melvin, who leads post-training, and with Jack, who leads the reasoning efforts, about sort of what you described as like this across-the-stack improvement, which it really does feel like this is true. and we sort of talked about how the pre-training benefits translates to better reasoning capabilities. Just from your perspective as the person orchestrating to make sure that we actually get a model that the world likes and developers can build with, is it just like this moment where all of those three paths converge together in a really natural way and it happens?
8:29Or was the intent like, let's make sure that this pre-training thing lands at the same time that all these other improvements land so that we can get the full breadth of the, how does that process actually happen? I think, you know, probably like any good system, there's a little bit of centralized organization and then individual ambition and execution, right? And so every part of this stack is doing innovations of their own, right? And they're driving progress to really think through research and technical breakthroughs. And so you have the pre-training side really thinking through what is their sort of scientific method of essentially testing pre-training improvements and growing.
9:08There's the post-training side that is thinking through specific capabilities and how they tune the model. There's thinking that is trying to drive new algorithmic innovations and kind of how we do reasoning. And so all three are driving those innovations. But I think what is really cool is, one, they are all trying to think through composability, right? So the pre-training side is trying to think through how do we best train this base model such that it is most composable and most, you know, client, I guess, to like downstream post-training. And so all of these pieces are actually trying to think through like how they be one puzzle and like how they fit together.
9:42I think the second part also is like intentionality of goal setting, right? So because we knew, for example, that code is an area we really do want to make progress on, that was something that we prioritized across all parts of this stack, right? So from a pre-training perspective, we thought about, okay, what is the data that, for example, would be required in pre-training to be awesome at code. From a post-training perspective, we thought about, okay, if we wanted to, for example, build better web apps, what would that look like? And then from a thinking perspective, we've also been thinking through how do we help the model reason about code?
10:13And so those three things then also come together well to push a single domain forward. And then I think that also translates across the board. But I think that's also been really critical as kind of driving towards a common force. Yeah, that makes a ton of sense. And actually to provide the counter example of this, like with 2.0 Flash thinking, Was that just like reasoning innovation or like we didn't do a bunch of post-training or pre-training work? Or like was it the full breadth at that point? With 2.0 Flash thinking, we of course benefited from the 2.0 Flash model. Yeah, yeah, yeah. Right?
10:43So the 2.0 Flash model was of course also had pre-training innovations and post-training innovations. And then we were thinking about, okay, how do we take that model and introduce reasoning into that and build on top of what 2.0 Flash is? So I think it still had all parts of the stack, But I think the difference between 2.0 Flash Thinking and now what we've introduced with 2.5 Pro is I think with 2.0 Flash Thinking, we proved that we could, even on a smaller model that is sort of more cost efficient and things like that, actually make it perform extremely well on reasoning and complex prompts with the introduction of thinking.
11:21But I think what we've done with 2.5 Pro is to the point about well-rounded is we've taken all those innovations and we've a heightened those themselves But we've also introduced this idea of like how do we also make sure that it's also just a really That it keeps the other aspects of what makes flash and pro great models, which is like the The strong multimodal performance the long context performance the style the the vibes if you will tool use Yeah, like all of those aspects. Yeah, I love that Something that is exciting about this is like, I think historically a lot of the reasoning narrative right now is like, test time compute is the only thing that matters.
11:58Like, you know, stop pre-training, stop post-training. There's no value. Yeah, just throw more, compute at the end of this process. And somehow magic is going to happen. And I think this is actually a great, like, verifiable example of like, well, that's not actually true. Like, there really is like the hard work on pre-training, like, is paying off and having a better model. Is there any like, yeah, how like? Yeah, I think actually Sub gave this example maybe earlier, I think, when we were talking with him about pre-training innovations. And I think like one of the examples he talks through is like you can have a model be really good at reasoning.
12:29But if it doesn't actually know the theorems, the underlying knowledge to then build out its reasoning, its reasoning will only go so far. And so I think the idea is that, for example, pre-training gives you a base foundation of knowledge. and that base foundation of knowledge can then be like better customized and tuned and tweaked in post-training for different use cases and you know built out and then you have uh inference time on top of that is but or as a part of that rather really um that can build that out even further and so I think about it as like test time is obviously important like inference time clearly is important and that's been proven both in 2.5 pro but in a series of models that have shown that extending the way that a model can think allows it to better produce outputs.
13:15But I think that is, you can do more if you have the strong foundation to build off of. Yeah. Tulsi, I think one of the most interesting threads of this launch is sort of back to this pace of how quickly we're able to make these models happen. I think we obviously, behind the scenes, we had a bunch of other model candidates. We were thinking about how we want to bring a new version of this model that thinks to the world. But like, what was the rundown in the series of events that led to today when we just shipped the model? Yeah, it's a great question. Like you said, I mean, this has obviously been something we've been working towards, which is to have a 2.5 pro candidate, really strong on reasoning, bringing thinking to our models, you know, actually having a Gemini series that is really built on that reasoning capability.
13:59And so that's what the team had been working towards across pre-post thinking. And I think one of the challenges we were encountering was to something I said earlier about the models being well-rounded. We were both trying to make models that were good at the vibes, if you will, right? Like good for users and our products and consumers who actually want to play with the models, whether that's like a developer or an enterprise customer or consumer. We also wanted the models to be really good at the benchmarks and at thinking and at reasoning. And so we were pushing on both of these fronts. And we were having trouble getting to a model that was really doing both of these things well.
14:33We were finding candidates where we felt really good about how it was engaging or how it was working on certain tasks, but then couldn't maybe push it as much on code or weren't seeing the gains on reasoning. Or we had another candidate which was like really great at code, but we weren't necessarily seeing the impact elsewhere. The team had been, you know, hill climbing on both of these goals and really pushing sort of innovations to build on both of these. And we got back some of our eval results. We were looking at the numbers. We were playing with the model and going through examples. And we're like, this is a really good model.
15:07And we're really excited about it. We did a bunch of our own vibe tests to kind of check the model on a variety of different prompts. I maybe built like 50 random web apps. I just tried to see what the model would do. And we were like, we're really excited about this. This is a really fun model to test. And so we're like, we really have to get this model out the door. We have to get it in the hands of developers. We have to see what people are doing with it. I think one thing that's also really interesting when you're trying to take a model through this is how do we also make sure that we're being intentional about what we're releasing, right?
15:38So we have a model that we're excited about. What are all the pieces that we need to work through? So for example, one area that I obviously think a lot about is safety. And how do we make sure that the model we're putting out is not just a good model, but it is also a safe model. And so one thing that's also I really like about how we've adapted our process to build quickly is safety is actually embedded in the development process itself. So every time we build a model checkpoint, we are evaluating that model checkpoint for safety when we're building it. And so when we're looking at the model's numbers and we're looking at how it's performing, we're also looking at safety numbers to say, hey, you know, what is it?
16:14How is it performing? And then we ask our teams to actually red team the model. The team has been bashing at the model, trying to identify issues. And it's funny, actually, because sometimes in doing so, they don't even uncover safety issues. They just uncover other random issues. And so we had the team reaching out and being like, hey, so we were doing some safety red teaming, and we found this other weird thing. What should we do about that? And so it actually also helps us patch kind of issues with the model overall, which is actually pretty cool. Yeah, I think the safety as a feature of the model development process is actually super important, because I think it's like a key note of how we're able to move quickly.
16:46Like, I think if you decouple safety from the model development process, you end up like, hey, we've got a great model. We want to put it out the door. Now we need to wait five weeks or however long it would actually take to go through the whole process. Yeah, it's also not fun for anybody involved, I think, right? Because then it actually makes safety like a blocker in the process. And that doesn't lend itself to innovation, right? Whereas I think when you actually enable safety to be a part of the process, what you're actually saying is how do we build a model that is also helpful, right? And ultimately you change the notion of safety from being like a wall to being like actually a way to make the model better and more useful to people.
17:25And I think that's actually a really nice change too. Oh, I love that. And Tulsi, so one of the big threads of this launch is the model's very good at multimodal understanding. Video understanding is something a bunch of folks have been talking about. Like why is there like something special we did to do this or like why like what's the highlight of why video understanding is so much better with with 2.5 Pro? Yeah, I think again, it goes back to the whole stack piece of things where with video understanding, video understanding is interesting because it's a combination of good multimodal understanding, right?
17:54So you need to understand vision. It also requires often long context because for example, if you want to look at a multiple hour match. So for example, my parents are big cricket fans. Some of these cricket matches are hours. We're talking hours of content, right? And so to actually be able to put that and process that, you need long context. And then the last part of it is also actually strong reasoning, right? So if you imagine, for example, taking a video of a cricket match and then being able to say, okay, I want to, can you help me identify all the points in the video where actually a wicket was taken, right?
18:29And so you actually then want the model to be able to analyze the video, be able to pull out critical timestamps of that video, reason about those timestamps, and give you explanations. And that's the kind of thing that brings together, I think, a lot of the magic of what Gemini models do particularly well, like the multimodal understanding, the long context, and the reasoning pieces. And so we're seeing that, I think, actually play out with this model. And so I'm actually really excited to see what people do with it. What I wanted to ask earlier, when you were mentioning sort of this trade-off between the vibes being right and academic evals, Like, what does it actually look like in practice for, like, you're closer to the model development process than I am, but, like, you have a model that's really performant on academic evals.
19:11Is it just that it's, like, harder to get the model to do, like, it's less, like, less good at instruction following? Or, like, what does it actually look like for, like, the vibes to be off? Yeah, I think it's less, I mean, instruction following is important, I think, even for academic evals, right? I think there's some certain foundations to a model, actually, that are important that build up everything else, right? So like instruction following and steerability, I think, are just kind of critical tenants to the model being able to do a lot of other things. But I think one way you can think about it, too, for example, is like maybe like another way to think about it is model behavior or like persona of the model.
19:45Right. So anytime as you're trying to improve a model, you can think of it as hill climbing some sort of goal. Right. And often that goal is set by an eval. So often you can hill climb academic benchmarks because you can you have a metric and you can sort of hill climb towards it. And just to clarify, we care about academic benchmarks for us matter a lot because it's just something that people universally agree upon? Yeah, it's a good question. And actually, we don't look at all academic benchmarks because we sometimes look at an academic benchmark and say, actually, we don't agree with what this benchmark is testing, or maybe it's a really leaked benchmark.
20:19And so it actually, like, being really good at it doesn't give us much value. And so there are academic benchmarks that we actively kind of are like, okay, that's nice, but we're not going to pay attention to it. But I think some of these academic benchmarks, so I'll give the example of Humanity's Last Exam, which actually 2.5 Pro is soda on, which is awesome. We got 19 % without tool use, or 18.6 % I think without tool use, which is awesome. And what does Humanity's Last Exam actually cover for folks who are in close? So Humanity's Last Exam is this example of, I think, 3 ,000 prompts that are supposed to represent really hard questions compiled by researchers and industry experts.
20:59And so Humanity's Last Exam is the kind of example of an academic benchmark that is really interesting to climb because it represents the kinds of questions we want Gemini to be really good at. And so then actually moving that metric is also moving towards a meaningful goal. Right. Another example of an evaluation like that is Sweebench. So Sweebench verified is an example of sort of like agentic code. And that's another case where like we really want the model to be good at these type of sweet agent tasks. Being able to, you know, push on Sweebench is a good way to to verify and validate that.
21:37So I see academic benchmarks as kind of a way to a motivate a specific destination of progress and then also help communicate to developers what the model is good at, right, and kind of where it has excelled, which I think is interesting. And I also think it's important to accompany any academic benchmark also with our own internal evals, right? So a big part of our efforts within Gemini are to building the right evals and making sure that we're actually measuring, you know, the goals that we have. Yeah, I love that. That's awesome. This launch has been a ton of fun. there's been lots of interest.
22:14I'm happy that we have a state-of-the-art model. What's coming next? I think we sort of have a sense of what the roadmap looks like, at least for the next few months. What can folks look forward to? Yeah, so much to look forward to. So first of all, 2.5 Pro released today experimentally, but one thing I know we've been talking a lot about is actually getting people access to the model in a way that they can build with it at scale. So one of the things we're really excited about is actually pricing the 2.5 Pro model and actually releasing it for use in production and use at large scale. So that's one thing I think we should all be really excited about.
22:50Very soon, hopefully. Very soon, hopefully. And I think this is hopefully us really listening to developer feedback and actually acting on it in a useful way. And so I'm excited to see what people do when we can give them more scaled access. So that's kind of more tactical. I think in the more kind of immediate term, we, of course, want to bring the 2.5 kind of series to more models. And so, you know, Flash is, of course, the next on our list to come forward. And then we're also thinking about how do we make these models more usable, right? So one of the challenges with thinking models is that they think often for a long time.
23:29And that's helpful to make the model more performant. But models don't need to think as much necessarily for simpler prompts, right? So going back to like my sort of different levels of vibe checking, for a web app that is really complicated, maybe the model needs to think longer. But for like, hi, how are you, it probably doesn't really need to think at all. And so how do we make sure that the model kind of like better learns, you know, how to modulate? And then also how do we provide developers more control? And what does that look like, right? especially when we're thinking about cost and latency and different types of applications.
Read the full transcript
24:04So that's like a lot of what we're thinking about in the near term is kind of like how to bring that to bear. We're also thinking about things like image generation and how to bring that into the fold of these models. And so there's a lot of really exciting pieces coming. I think if you then look like even farther out, I think a lot of what we're trying to do is just build like the most capable models to be able to help you, whether that's encoding, whether that's in building agents and awesome kind of agentic applications. You know, when we talked about 2.0 Flash in December, we were talking about Mariner and UI control.
24:39And I think there's a lot of really cool things that we can start to do as these models get more and more capable in terms of like building sort of end-to-end experiences that I think will be really fun. Yeah, and just one quick follow-up. So today's 2.5 Pro model does do a little bit of dynamic thinking. So it's maybe not the full version of how we want that experience to look like. But if I ask a simpler question, the model's been trained to know, hey, these are... Yeah, absolutely. I think for simpler prompts, the model will definitely think less than for more complex prompts. But I do think it is true today that the model probably overthinks quite a bit.
25:17And I think that's actually a good place for us to start. What we really wanted to do was kind of push how the model could use reasoning to solve awesome problems. And now I think we want to take a look and say, okay, how do we continue to make this more and more useful for developers and for customers who should be more cost conscious of like kind of how the model actually modulates that? Yeah, I love that. Tulsi, this last three days has been a ton of fun. I think the most fun I have in my job is when we're sort of sprinting through the chaos to get new models. Yeah, the little ship emoji. Yeah, the little ship emoji and getting stuff out the door.
25:49So this was a ton of fun. Huge thanks to your team for making all this possible and the rest of the Gemini. I mean, vice versa. I mean, shipping is a team sport. I think that's one of my favorite sayings these days. Like, it really does take so many people to make these launches possible. It's kind of amazing. Like, I feel like we have, we talk a lot about, like, the parts of the model training, but I think we don't always talk a lot about the parts that happen after. Yeah. Like, you know, we have Madhvi and the team building demos and actually, like, really testing the model. We have, like, our marketing and comms teams.
26:18We have our deployment team, like, actively figuring out how to serve the model efficiently and trying to figure out all the bugs and the kinks and getting the model stood up. Which is not easy. Which is not easy. We talked to Emma. Emma was the first person to come on this show or whatever we call it. And him and the serving folks and Evan and everyone, Joe and everyone else and Alvin in the trenches doing all this work. It's a beautiful orchestration of chaos to make this happen. Yeah, absolutely. And I feel like even in the last three days, but as we've been doing more of the shipping, I've been learning so much more about deployment and serving and just how complex that is to get a stable experience set up for an end developer.
26:59And yeah, I feel like it's been quite the team effort. Yeah, I need to trademark that long context is as much of a model innovation as it is an infrastructure innovation. I think we actually feel this in a lot of cases. There's challenges with making long context work because it's an infrastructure innovation and it takes effort to make it happen. context, tool use, like so many of these things that we ship with our models are truly end-to-end innovations. And yeah, long context is a great example of that. Getting it right is hard. And it's a huge team effort. Tulsi, this was a ton of fun. And you are now becoming the resident model launch co-host of this.
27:36So hopefully we'll have you on to do this again. I mean, we're going to ship a ton. So we're going to be doing this a lot. I'm excited. Hopefully we'll be in person in out of you again. Deal. I love it. All right.
From the publisher
Tulsee Doshi, Head of Product for Gemini Models joins host Logan Kilpatrick for an in-depth discussion on the latest Gemini 2.5 Pro experimental launch. Gemini 2.5 is a well-rounded, multimodal thinking model, designed to tackle increasingly complex problems. From enhanced reasoning to advanced coding, Gemini 2.5 can create impressive web applications and agentic code applications. Learn about the process of building Gemini 2.5 Pro experimental, the improvements made across the stack, and what’s next for Gemini 2.5.
Chapters:
0:00 - Introduction
1:05 - Gemini 2.5 launch overview
3:19 - Academic evals vs. vibe checks
6:19 - The jump to 2.5
7:51 - Coordinating cross-stack improvements
11:48 - Role of pre/post-training vs. test-time compute
13:21 - Shipping Gemini 2.5
15:29 - Embedded safety process
17:28 - Multimodal reasoning with Gemini 2.5
18:55 - Benchmark deep dive
22:07 - What’s next for Gemini
24:49 - Dynamic thinking in Gemini 2.5
25:37 - The team effort behind the launch
Resources:
- Gemini → https://goo.gle/41Yf72b
- Gemini 2.5 blog post → https://goo.gle/441SHiV
- Example of Gemini’s 2.5 Pro’s game design skills → https://goo.gle/43vxkq1
- Demo: Gemini 2.5 Pro Experimental in Google AI Studio → https://goo.gle/4c5RbhE
