In short
Google AI: Release Notes - Episode Summary
Episode Title
Sergey Brin on the Future of AI & Gemini
Host
Logan Kilpatrick
Guest
Sergey Brin, Co-founder of Google and Computer Scientist working on Gemini
Episode Description
In this episode, Logan talks with Sergey Brin about the advancements in the Gemini AI framework following a significant year of progress. The discussion encompasses the latest features, insights on AI model training, and reflections on the evolution of AI technology.
---
Key Topics Discussed
- Initial Reactions to Google I/O
- Sergey shares his excitement about the recent announcements at Google I/O.
- Noted the internal and external positive reception and the extensive developments across various products.
- Focus on Gemini’s Core Text Model
- Sergey emphasizes his primary focus on the main Gemini text model, which is seen as critical for advancements in AI development.
- The importance of self-improvement and enhanced coding abilities through the text model is highlighted.
- Native Audio in Gemini and Veo 3
- Discussion on the integration of audio capabilities into the Gemini framework.
- Sergey expresses surprise at how the incorporation of audio enhances user experience, making generated content feel more engaging.
- Insights from Model Training Runs
- Sergey provides insights on how training runs allow for observational checkpoints to evaluate model performance.
- The ability to assess models at various stages of training is essential for ongoing improvements.
- Surprises in Current AI Developments
- Sergey reflects on the unexpected advancements in AI, particularly the rise of language models.
- He discusses the interpretability of these models and how it aids in understanding their reasoning processes.
- Evolution of Model Training
- The conversation touches on the architectural similarities among AI models and the integration of new capabilities like tool use in the training processes.
- Emphasis on the shift from primarily pre-training to a more balanced approach that includes extensive post-training phases.
- The Future of Reasoning and DeepThink
- Sergey discusses the progress of the DeepThink model and its capabilities for extended reasoning.
- The potential for longer reasoning processes in AI is seen as a significant breakthrough.
- Google’s Startup Culture and Accelerating Innovation
- Sergey comments on the evolving culture within Google, likening it to a startup environment focused on rapid innovation.
- He notes the company's historical adaptability to technological shifts and the importance of maintaining momentum in AI advancement.
---
Key Takeaways
- Gemini’s Potential: The Gemini framework stands out as a foundational element for future AI developments, especially in text and audio generation.
- Revolution in AI Reasoning: The advancement of models capable of prolonged reasoning marks a new era in AI capabilities, pushing the boundaries of what these systems can achieve.
- Cultural Shift: A renewed sense of urgency and innovation within Google reflects a commitment to leading the AI industry, reminiscent of startup dynamics.
- Importance of Interpretability: The ability of AI systems to provide reasoning behind their outputs enhances user trust and safety in AI applications.
---
Closing Remarks
- Sergey expresses gratitude for the collaboration and acknowledges the hard work of the Google AI team.
- The episode concludes with a light-hearted moment as Sergey unboxes a TPU v4, emphasizing the ongoing dedication to computational excellence.
For those interested in the future of AI and Google’s role in it, this episode provides rich insights directly from one of the industry’s leading figures.
[Watch on YouTube](https://www.youtube.com/watch?v=o7U4DV9Fkc0)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:04Hey, everyone. Thanks for joining us. We've got an IO special. Sergey Brin. We're talking all things Google. Thanks for taking the time to chat. Thank you, Logan. And you know, you and I are all the time in chat spaces on all kinds of products, but it's nice to hang out in real life. Yeah, it is. My California experience, as always, it's incredibly fun. I spent a bunch of time with Cori yesterday and today, and like he's just, it's like you feel the warmth and the humanity of AI progress when you spend time in person with everyone. So it's been a ton of fun. But we're the general sentiment from the world and I think also from the team internally is like incredibly great day for Google.
0:46A ton of progress across models, across all of our products. What's your take? What's your reaction? Obviously, there's a lot of stuff we still need to do, but where's your head at? Yeah, I think it was definitely a phenomenal set of announcements. Honestly, I probably didn't even know about 30 % of them or so. You know, there's only so much time, and I'm kind of deep in the weeds on Gemini, and I didn't even know about the virtual fit, for example, for products in Google Search. I didn't realize we were shipping that. So there were many things that also even surprised me. That was great. I think the reception's been great.
1:30There are so many things, though. So I think it takes a while for people to explore them, wrap their heads around them. And obviously we're very busy delivering all of them right now. And it's a lot of energy across the way. Just make sure things are actually shipping, shipping smoothly, that people are able to sign up for Ultra, have all those new features and so forth. Yeah, I feel like IO is the start of lots of work for lots of other people. It's like at the finish line for some teams and then it's like the starting line for a bunch of other teams. How do you think about, obviously we launched more, there's a bunch of Gemini announcements, Gemini Diffusion we'll talk more about later, DeepThink, which is like continuing to push the frontier on reasoning models.
2:20I see you in the reasoning strike chat all the time, sort of pushing on people to keep pushing the frontier, which I love to see. um how do you think about sort of maybe maybe how you think about your focus but also just like generally the like deep mind team's focus around like vo and imagine and we have like a whole suite of generative media models we just announced a lyria our music model um and then also at the same time like the core gemini main model like how you know do you do you work on the gen media stuff at all or is it mostly like laser focus on on gemini at the moment mostly i'm on gemini the core text model primarily because i think that that's what will lead to for example self-improvement will help us code and develop the science behind ai even better so that's the number one focus that i have at the same time the generative media is so amazing yeah i mean that is just so superhuman you know I mean with a text model you know I mean there's some like math problems or whatever that I might be able to solve that it gets wrong or something like that or stumbles on a piece of code although that's less and less frequent and I actually end up now relying on Gemini to do some of my coding math and so forth but nevertheless it's sort of in the range of human given my artistic talent there would be zero chance of me ever coming up with either you know an image or a video i mean just the amount of work that that would i imagine would take if i was an expert you know videographer 3d renderer or whatever special effects person i mean i mean that's got to be you know a month of solid work yeah to get the thing i can get in a few minutes um and obviously it's visually so compelling that, you know, it just kind of sucks you in.
4:28You can't escape it. The audio piece with VO just makes it feel like I historically have, you know, personally, I think like generating videos is awesome and stuff, but it's always like kind of felt slightly gimmicky to me. And I think when I saw audio in VO3 on stage yesterday, I think it like made that was the moment for me where I was like, okay, this is actually so many people are going to be able to just cause also practically historically, like you could generate a video, but then like, you'd have to go and like, where does audio come from? You know, how do you, you know, sync everything up and now like you can make humans like talking and having a conversation.
5:06Uh, and like, it just does it all well. And it's, it, it blew my mind. Uh, yeah, you're right. I mean, I, I've been a huge fan of that. I I'm like a pretty visual person personally. I'm not like a very, I guess, audio person but uh over the years especially you know like with google glass i mean the moment we added some sound i mean that's just like it adds so much richness to have sound yeah i mean it's like you're better off adding audio than adding like 3d for example um although some of the 3d stuff's cool if you played with uh the big wearable thing yeah um but anyhow um yeah it's it's it's just an incredible change in perception uh when you get audio working and i knew i saw the model training the last month or two uh and uh you know i just saw it from checkpoint to checkpoint and i knew that wow this is just going to feel different yeah it'll be interesting to see how the like fusion of those capabilities because it does seem like there's a lot of like similarities to mainline Gemini like mainline Gemini model and obviously we landed native audio support um in the mainline Gemini model at IO um and in VO as well and I was having a conversation with Tulsi this morning about just like how you know are those similar breakthroughs are they different it sounds like it's actually technologically very different from a technology standpoint but it is cool that we have like other rails to do this innovation and like ideally it all upstreams back to Gemini in some way?
6:41Yeah, I mean, honestly, I think we've just taken a long time to ship the native audio in Gemini. It's been in there for over a year. Back in December, right? Oh, really? No, no, the base model has had audio that it's trained on for at least a year. And I don't know, there's always like, I think honestly it's just, there's just so much to do, so much to ship that nobody has, for whatever reason, gotten it out there. And I mean, like, native audio it in, native audio out. I think native audio in has been in there even longer. But to get through all the little hoops to make it really work well and so forth, I just think for whatever reason, it took a long time.
7:28But yeah, that's finally out there. I don't think that's done, as you say, the same way Vio does it, I believe does the audio also through diffusion just like it does the video yeah in fact if you watch the during the training run you can actually see it generating videos that are like you know um a couple percent into it it's these kind of you know the shapes aren't quite right and the words are kind of warbled and stuff like that but then it takes shape and it develops uh until you know at the end of the run you have what you have uh what you see today um so yeah I'm pretty sure that that's a diffusion-based audio yeah i mean diffusion is a really powerful technique you know as you know we shipped uh text diffusion for you know small early test run um i mean i i think that's one of the things that i'm grateful that we have you know the bench of machine learning researchers in there that we can pursue simultaneously different based techniques for across modalities.
8:34Yeah, the results of Gemini Diffusion look super, super promising so far. I'm hopeful the model progress is there and it all fully works because the demo works. We were talking off camera. The demo looks really good. So hopefully the capability translates well and everything works from that perspective. But you mentioned this before about watching the training run. I actually haven't seen what this looks like. So what does it actually mean to watch the training run? Oh, okay. Well, maybe you've seen for our text models, but, you know, we're able to test out the intermediate checkpoints, you know, 10 % of training, 20 % of training, and so forth.
9:12And it's, you know, the model is weak at those points in time, but it kind of, you can kind of get a sense for the trajectory. Yeah. And so, you know, usually, especially if you have a big training run that you have a lot sort of, you're using a lot of compute and you have high hopes for it, you're going to test it out in various ways through, you know, many times throughout the run. So you're going to have a pretty good sense of what it's on trend for. Yeah. So that's true for the text models. That's true for the Fusion video model for VO. All of these models kind of have these intermediate results that you can take a peek at.
9:58And if you're really deep in there, you're for sure checking them because you're nervous and excited about what exactly it will produce. Yeah. I need to come and hang out at the MKMORE and watch. Yeah, come over later. Watch some of them, yeah. Um, one of the, I was listening to Sunar's conversation with, uh, Dave Freeberg and he soon made the comment that, um, even 15 years ago, you and Larry and him were having conversations about, uh, and like the team at Google were having conversations about like what this future facing AI moment would look like. Um, and that there, it's like eerily close to what you all were talking about 10 or 15 years ago.
10:39So I'm curious, like how, like what, what are the things that have been most surprising to you about this, this moment? Like even, and we can ground it in products if you want to look at search or just the technology or like what, what's been surprising and what's been like almost what you would have expected to have happened? yeah you know i i think from an intellectual standpoint you can kind of reason your way through the singularity and uh you know famously ray cursewell did this i don't know however many decades ago um i mean it kind of um i don't remember what date he said it was 2037 or I can't remember.
11:20He put some date on it based on his extrapolation. Today looks like maybe that was kind of conservative. I don't know. But you can intellectually sort of reason through it. I think to see it happening is altogether different. Yeah. And I think when you're kind of talking, whatever, 15 years ago, i won't say you're like joking around you're like truly talking about it but you're kind of like you know imagining the science fiction future but it's almost like um like a game like a like you're just kind of uh you know chatting with the other folks who are interested in it like i think it's fun um but yeah as i said seeing it actually start to happen you know feels very different yeah and of course the way in which it happens is pretty surprising um and i can give you an example i mean the fact that language models seem to be right now the way that you know ai is developing I don't think you necessarily would have known that 15 years ago in fact DeepMind you know in the past and even now to a certain extent has bet a lot on this kind of physical grounding that it's important to have kind of a physical world to ground on and we're obviously doing experiments in that vein like Genie and whatnot but the fact that these language models have come as far as they have wasn't obvious.
13:09And an interesting side effect of that, especially with the thinking models, is that they are also surprisingly interpretable. Like you can look through the thoughts of one of these thinking models and how it's coming to the conclusions that it's coming to. You can't necessarily, without a huge number of tools, inspect the weights of the model and try to infer stuff from that. but you can understand a lot of its reasoning in very understandable terms. So I think that's kind of a, you wouldn't necessarily have guessed that 15 years ago. Yeah. So that's been an interesting surprise, which I think gives a lot of comfort, not infinite comfort, I'm not saying we should ignore it, But from a safety standpoint, the fact that these things do to some extent say what they're thinking, I think is a big plus.
14:14And yeah, there are papers out there about how they're sort of lying and stuff, but I think that's a relatively smaller effect. what's your sense just being close to the model training process today as far as like how different or how similar it's going to look as the models move from being like text in token in token out or text out to being like actual systems and i think we're actually like we've already taken it like gemini 2.0 like search is natively in there code execution is natively in there like the model is learning that in the process do you think the like training infrastructure or just like the way we think about models is going to be fundamentally different as they're not models anymore, they're really like full systems that we're baking for people?
15:05I think it's the confluence of a few things. One thing, it's kind of remarkable how similar architecturally all the different models are, including, for example, VIO, which you would think video diffusion's like very different than some text language model. But architecturally, there's a huge amount shared. It's kind of astonishing how much is shared. And a lot of it uses transformers at the core, which thanks to Nome and the crew, we've had now for approaching a decade.
15:40Now we are adding things like tool use and so forth. mostly those things come about in the during what we call post training and post training is an increasing fraction of the overall training you know it used to be that everything was like 99 % pre-training and it's sort of shifting now to now maybe it's 90 % will be 80 % and so forth and this post training which is sort of you know some people call fine tuning but it includes you know the rl kind of work that we do um you know used to be this just little bit of tiny bit of shaping you do in the end but now it's more and more material and the kind of things you're mentioning with tool use and so forth come in during that now much larger phase uh and um yeah i mean look that makes the model vastly more powerful.
16:40Yeah. I've got two more questions because I want to get you back in the office working so that we can keep making model progress. First one is just around the reasoning scaling. I think we announced, we showed the results of DeepThink, which is sort of continuing with ScaleUp 2.5 Pro and let it reason for longer and have sort of parallel thought processes. What's your sort of general reaction to that? I think it feels like we're so early in that scaling paradigm that there's like going to be a huge amount of additional unlock but uh you're obviously you're you're in the weeds on this one so i'm curious what your what your thoughts are um yeah we have a you know interesting we had about like five different approaches to doing that kind of thing and they all converged on this deep thing so it was uh um it was great to see all those people and those teams come together and that the You know, sometimes we fragment and it takes a long time.
17:39But in this case, we kind of took the best of ideas of all of them, combine it in one go. And yeah, it's definitely yielding stronger results, obviously. I think that the more that continues to happen, that is like a superpower. hour if you can have these models and i know lots of the top ai labs talk about this but if you can have these models instead of just thinking for a minute spinning out an answer if you can leave them to go for an hour for a day or for a month maybe and they actually get you a significantly better answer to a really important question that can be incredibly valuable yeah uh and that's that's kind of new and it's um and it's non-trivial it's a little bit like you know we cracked long context for input um we did that a while ago we've had million plus context for um a year and a half or something like that now we need infinite context so you got to keep pushing the team yeah i'm not saying millions enough but that generalization is sort of non-trivial.
18:54Yeah. You know, for a model to... It's as if you were to go through like Groundhog Day. You know, you just get over and over. You get to sample one day as a person. You try this, you try that. And now all of a sudden, you're meant to live your life where stuff happens from day to day and week to week and month to month. It's kind of a non-trivial generalization. But, you know, we figured out how to do that. On the output side, it's also, it's just kind of non-trivial if all you do is, say, short little math problems. Going from that, well, it's kind of like, you know, we interview people. We ask them 10 interview questions or whatever.
19:37And then we expect them to build these big systems over months. And it's not clear if that's actually the right way to test a person. But on the AI models, we do that, you know, a million times over. Like, we're only training them to do, you know, short, little clever math questions, coding things, whatever. Yeah. And then the expectation from there that, hey, they actually can spend a long time to develop something new that requires, you know, thinking over days, let's say. That's very non-trivial. But that is a gap that we're starting to overcome. And that's a huge leap forward. Yeah. This example that you gave around just like how we test and eval models is always my consistent reminder that so much of life is a, like this AI moment has taught me that so much of life is like actually an eval problem.
20:31And like the challenge of like even interviewing people and like trying to build a great team and all that stuff like is an eval problem at its heart. And like humans haven't solved that. Like it's not surprising to me that we haven't solved the like AI eval problem as well. It is non-trivial to do that. But my sort of final question for you, and this is just like, again, in reaction to everything that we're seeing with IO and the pace of innovation. I don't know how to slide on the screen. Actually, you know, Demis did where it was like everything we shipped in 2024 and then everything we shipped in 2025 so far.
21:02And it was like, I'm pretty sure the 2025 section was bigger than the 2024 section. So like clear acceleration happening. and it feels like at least on a personal level for me joining Google I've been here more almost a year a little more than a year now like Google truly feels like a startup experience to me and I'm curious to get your reaction to that but also just like your reaction in general having like seen Google grow and you know expand and everything that's happened in the last 20 years how you think about that um and the great question i mean first of all you know i think companies need to periodically reinvent themselves um and uh you know there are different important technology shifts i guess you know we started as a web company we had to make mobile work um we know we're never particularly good at social to be honest as an example but um you know now we're in the ai domain and i think from there like the it's exciting because in some ways google has always been an ai company we've always been about large-scale data and analysis um we are also the place that gave birth to a lot of the modern large-scale machine learning from Google Brain to the Transformer and so forth.
22:33I mean, it is in the company's DNA. So this is a shift we should be really well equipped to make. You know, any shift is probably hard on any company, but I'm feeling really good about it. And I think going from 24 to where we were honestly catching up on many levels to 25, and particularly with the launch of 2.5 Pro. I mean, that was just like a clear leap forward. And I know, whatever, on different benchmarks, maybe we were for a bit number one before or not. 2.5 Pro was a big step forward, kind of across the board. And even, you know, to date, it remains on most leaderboards, number one, you know, with style control, without style control, however you measure it.
23:29So that's been just a really exciting leap forward. And I think we've, it's both sort of cause and effect of the kind of science, the science engine we have going behind it. it's going to help propel us forward and it's thanks to what all of the science we were doing over the past year that we were able to finally produce that model and in quick succession after that a lot of other things have followed we've already gone through a few different iterations of the 2.5 Pro model I don't know if everybody noticed yesterday we launched the new 2.5 Flash I mean do you notice that's actually like on many measurements it's number two behind 2.5 Pro.
24:20Yeah. So we're like one and two on many different leaderboards now with that flash model, which I think with all the other announcements, I just think a lot of people might have glossed over. It got buried. But it's like a super fast model. It's really powerful. I think it's going to appeal to a lot of use cases. Yeah, but really with that kind of cornerstone of 2.5 Pro this year, I think for us to be able to build off of that and continue the momentum is really exciting. It's going to be a great year. Sergey, I appreciate you for taking the time. I appreciate you for pushing everyone hard. It's a ton of fun to see.
25:00And we have a special gift for you, which I'd love to see you unbox. And somebody will bring it over to us in one sec. Well, thank you, Logan. And while they bring it over, I'll just say thank you, Logan. I mean, I see you working hard all the time and, you know, making all of our customers and partners happy and tracking, you know, the millions of issues that might arise. I mean, it's not that easy, you know, having these models that so many people and businesses want and, you know, getting them deployed and not having the TPUs meltdown. And, you know, every little nuance from function calling to caching to all the million things.
25:47And I see you being really good about putting the customers first, communicating the needs back to the team, being just really on top of the ball. So thank you. The team's crushing it. No, thank you. The team's crushing it. Special gift for you. All right. Thank you. Shall I unbox it? Yeah, yeah. You got to unbox it right now. We got to capture. The thread of this was just one of the pieces, other than all the people inside of Google, that makes all this possible. Nice. It's our TPUs. This is the TPU v4, which internally, by the way, we call Pufferfish. It's probably not too secret. I think Pufferfish is v4, right?
26:27Yeah, yeah. I never know the external translation. We just call these. I mean, these were the hottest thing to come by a year or two ago. though we've moved on to newer generations now. But we still do a lot of our work on these. So this is great. Yeah, hopefully we'll get a bunch of these in the MK for the team. This is really cool. It's a real one. They had to take it out of a data center somewhere. It was not being used. We were not taking compute out. Really? OK. I'm going to come find us. We definitely need these. Yeah, we need the two views. Sometimes some of the early samples are a little bit faulty, and maybe it's one of those.
26:58Yeah, I'm sure it is. But I appreciate it. It's super nice. Of course. I'll give you a little zoom in on there. Thank you. Thank you. Thank you. Thanks, Logan. Thank you for tuning in for folks who are listening. This is Release Notes. Thanks for watching.
From the publisher
A conversation with Sergey Brin, co-founder of Google and computer scientist working on Gemini, in reaction to a year of progress with Gemini.
Watch on YouTube: https://www.youtube.com/watch?v=o7U4DV9Fkc0
Chapters
0:20 - Initial reactions to I/O
2:00 - Focus on Gemini’s core text model
4:29 - Native audio in Gemini and Veo 3
8:34 - Insights from model training runs
10:07 - Surprises in current AI developments vs. past expectations
14:20 - Evolution of model training
16:40 - The future of reasoning and Deep Think
20:19 - Google’s startup culture and accelerating AI innovation
24:51 - Closing
