Gemini co-leads on project origins and what's next

22 Jun 2026 · 42 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Origins and strategy behind Google DeepMind’s Gemini (the “single model” push), plus what’s next across Gemini 3.5/Flash, agentic coding, evaluation, and world modeling toward Omni.

Guests (backgrounds)

Jeff Hinton (DeepMind/Google AI pioneer; early distillation work; long-time codebase reviewer); Noam Shazeer (Google/DeepMind researcher; early Google Brain/LLM-era leadership); Koray Kavukcuoglu (DeepMind researcher focused on models and evaluation); Oriol Vinyals (DeepMind leader; previously led DeepMind efforts and steered Pathways/Palm; deep work on agents and world modeling).

Key claims

Gemini name reflects “map and reduce” and unifying fragmented teams/compute into one general model; product feedback is essential (like search) to improve models; Gemini 3.5 Flash targets agentic coding while preserving multimodal/tool use; distillation now “packs intelligence” more efficiently (teacher-student, not large ensembles); evaluation/generalization to “anything” is a major bottleneck.

Notable examples

distilling a 50-model ensemble trained on 300M images into a single model; “Flash next-gen” outperforming prior “Pro” generation; Omni framed as a “true world model” that can simulate dynamics for consistent future rollouts; tools are too slow for long autonomous agent runs.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Genesis of Gemini: A Unified Vision

0:00 to 0:45

Learn how the idea of a unified AI model led to the Gemini project.

“Even before we started the Gemini effort, there were a lot of people thinking about, you know, building incredibly general purpose models that could do things.”

Unpacking the Gemini 3.5 Release

1:05 to 2:35

Explore the key features and expectations of the Gemini 3.5 model.

“I think this is like a third and a half generation of Gemini models.”

The Importance of User Feedback

2:35 to 4:25

Understand how user experiences shape the development of AI models.

“It's always fun and exciting on a day-to-day basis.”

Collaboration Across Teams for Innovation

4:25 to 6:20

Discover the challenges and successes of team collaboration in AI development.

“You can't do it if you don't actually do it with the products.”

Building a Multimodal AI: The Path Forward

6:20 to 8:10

Learn about the evolution towards creating a multimodal AI model.

“Jeff, that's a great segue to sort of, again, going back to sort of the formation of the Gemini project.”

Understanding World Models and Simulations

8:10 to 9:20

Explore the concept of world models and how they enhance AI understanding.

“And I think everyone, all of us can see that the whole organization is very proud of what we have built together, right?”

The People Behind the Technology

9:20 to 12:30

Hear personal stories about the relationships that shaped the project.

“So you would activate different pieces of it for different kinds of things.”

Finding Common Ground in AI Development

12:30 to 14:00

Reflect on how shared experiences and common goals drove the team forward.

“When you say multimodal, you instinctively are drawn to human modalities like text and images and audio and video.”

Early Days and Mentorship

14:00 to 18:09

Learn about the initial interactions and mentorship experiences of the speakers at Google.

“And I said, hey, you know, I mean, just Chad, I want to introduce myself.”

DeepMind Acquisition Talks

18:10 to 21:38

Hear insights into the acquisition discussions of DeepMind and code reviews.

“Was it actually, Jeff, you were just like pointing at random directories and then Koraj just happened to know.”
Show all 17 chapters

Reflections on Progress

21:39 to 25:18

Discuss the unexpected advancements and areas where progress has been slower than anticipated.

“I mean, the good side, like, thinking back to it, it's all good.”

Challenges in Evaluation

25:19 to 27:19

Explore the difficulties in evaluating AI systems and generalization.

“Yeah, part of it, I would say, is we really need to come up with algorithmic things that just get much more out of every piece of data or example that a Model C's or every token.”

Disagreements in Research

27:20 to 28:07

Discover the perspectives on research disagreements and collaborative experimentation among the speakers.

“I mean, the whole dream of every AI researcher ever has been, how do we build systems that can generalize to things they've never been confronted with?”

Exploring Research Disagreements and Collaboration

28:07 to 31:00

The speakers discuss the importance of differing perspectives in research and the collaborative efforts in developing Gemini.

“And I want to preface this by, I think this is a positive thing.”

Predictions for AI by IO 2027

31:00 to 35:49

The participants make predictions about advancements in AI and the capabilities of models by the year 2027.

“We should do predictions just so that we have something to be wrong about a year when we reflect back on this conversation.”

The Future of AI Products and User Experience

35:49 to 39:57

The conversation shifts to the potential evolution of AI products and how users will interact with them in the future.

“Because like if it finished it in one day, you'd be way happier.”

Personal Projects and Reflections on AI

39:57 to 41:52

The speakers share their personal projects and reflections on how AI influences their daily lives and creativity.

“I'm curious if like sort of obviously all the AI coding stuff, like just maybe personal, less Gemini specific, anything interesting that y 'all are doing.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Even before we started the Gemini effort, there were a lot of people thinking about, you know, building incredibly general purpose models that could do things. Oriol was leading some efforts in DeepMind and I was sort of helping steer some efforts around the Pathways project and things like Palm and Palm 2 and so on. I actually said, this is silly. We are fragmenting our efforts and fragmenting our compute. And if we're going to build an incredibly powerful model, we need to all come together and work on building a single model. That's actually where the name Gemini comes from, the twins. We mapped and then we reduced.

0:32Yeah, exactly. Yeah, something like that. I thought it was because I had twins. That too, that too.

0:45Hey everyone, how's it going? My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today we're talking with Jeff, Korai, Noam, and Oriol about all things Gemini, the origin of the Gemini project, and so much more. So thank you all for sitting down and chatting. We're sitting here in Gradient Canopy. We just launched the Gemini 3.5 era of models, starting with Flash. Contextualize the moment. I think this is like a third and a half generation of Gemini models. We've shipped a lot of stuff, a lot of models in between. Maybe, Oriel, you want to sort of fill us in on the moment with Gemini 3.5?

1:15Yeah, we could do one each, maybe. Contextualizing. I guess we started 2023, I want to say. We've had a few releases, some half models or even 0.1, and been building on some foundation of multimodal tool use, agentic from the get-go, and just kept building the capabilities up. So today, it's exciting to be releasing the Flash version of 3.5. It's a very powerful series, and the focus of this one probably is on agentic coding, and then, of course, preserving and enhancing the rest of capabilities. I think everyone feels like this is really the times where coding capabilities and agentic experiences are defining what it means to experience AI.

2:03And 3.5 is a huge, huge step there. Yeah. Right. And I think everyone is actually experiencing that. And it is being recognized as a very strong model. These big release moments have in one way gotten less exciting because the thing that's like most at the top of everybody's mind is not even the big release to the public. It's like, what am I going to be using tomorrow to do my engineering and my research? And what are my friends around the office going to be using for their engineering and research? And will they be complaining at me? Will they be happy? It's always fun and exciting on a day-to-day basis.

2:41this. Reflecting back on sort of the initial like moment of and like coming together and forming the Gemini project and shipping those initial models. Was it obvious to you all that the like product story of like how we bring Gemini models to the world was going to be so important? Not from like obviously at Google we have lots of products and we bring AI to our customers through the products but actually to like improve the model. Was that like a we hope that this happens and it's something that we're going to intentionally work towards, or has it just become obvious over time that we need to do that because the use cases are much more complicated than the initial version of Gemini?

3:17I'm actually curious what you all think. Oh, yeah. For me, that's my job. I mean, I can weigh in. I think that was actually pretty obvious that if you have a model that is used by a lot of people, that you're going to get a lot of lessons and experience of what's working well, what's not working well. And we've seen this in search for many, many years is the use of search by our users really informs, you know, what are the things that's not doing well? What are the things we should be doing better? You know, aggregating lots of interesting usage statistics to understand, you know, that in a deeper way and then working on improving those things is important.

3:55And AI models should be no different, right? It seemed like that was pretty obvious from the beginning, but we had to have a thing out there that people were using. Yeah, that's the true test. like people using it and is useful to people. Because if you go in a box and you try to hill climb benchmarks, then you end up hill climbing benchmarks and maybe you leaked your benchmark. It doesn't end up well. You don't want to build intelligence in a black box. You want it to be useful. You want people to be using it. Therefore, understanding what is required, like scratching the frontier is both scratching the frontier of the research in terms of technical capability, but also scratching the surface of what is the next thing that you can enable users.

4:41You can't do it if you don't actually do it with the products. And those two hand in hand define what the frontier means. So at the time like Gemini starts, there's a lot of already machine learning models making it into products. And I think what seemed obvious is that if we create a single more powerful than the average of the other models, like powering everything, that had to be like a leap forward. Whether there was a single product that could be created around a single model, that maybe wasn't as clear at the time. But I think it was very clear that putting all the compute and intelligence into a single powerful model was going to leapfrog many things Google was using already machine learning for.

5:27And that was very exciting to be given that amount of compute and responsibility initially. But I think it's proven to be indeed kind of the core engine of Google intelligence. Even before we started the Gemini effort, there were a lot of people thinking about, you know, building incredibly general purpose models that could do things. Oriel was leading some efforts in DeepMind and I was sort of helping steer some efforts around the Pathways project and things like Palm and Palm 2 and so on. I actually said, this is silly. We are fragmenting our efforts and fragmenting our compute. And if we're going to build an incredibly powerful model, we need to all come together and work on building a single model.

6:10That's actually where the name Gemini comes from, the twins. We've mapped and then we reduced. Yeah, exactly. Yeah, something like that. I thought it was because I had twins.

6:20Jeff, that's a great segue to sort of, again, going back to sort of the formation of the Gemini project. I'm curious, how controversial was that? Obviously, like, as you say it now, and we sort of have, we've done three and a half iterations, the sort of all the organizational complexity of bringing teams together is now sort of behind us. Was it, like, blatantly obvious at the time that, like, we won't win and actually deliver on building the right product and models for our customers if we don't do this? or was it like sort of originated as like a more pie in the sky idea? Or like, I'm just curious, like what was your level of confidence?

6:52I was sure the right thing was to come together. I actually articulated in a half-page memo, like this is silly to fragment. We shared this half-page memo. We should put it somewhere or release it somewhere. But it felt like fragmenting our best ideas across different research teams that weren't really working together and also fragmenting our compute. Just both of those issues seem like things we should fix. And it was a little bit organizationally complicated and time zone wise. Like there were lots of people in London, lots of people here. Eight hours apart is never a recipe for easy collaboration.

7:24But I think we've done a really good job of navigating that and bringing people together. And now we have like a really good, amazing team all over the world. And we're cranking out good models. There were a bunch of teams building LLMs at the time that you just needed to match together basically. At some point, research in AI was a lot more academic. You go back 10 years, a lot more academic research. And at that point, how you organize it is not really the key element. It is more about the exploration and the speed of exploration is important. But as things get more and more focused, what you want is to really, to Jeff's point, this focused operation where rather than us trying to build things in parallel because these things require a lot more focus effort and each one of them is like a major operation in terms of many researchers coming together solving many problems at that point i think it was really really a good idea to okay this is the moment that we need to change i think like both organizations acted with like great urgency on that and enabled it i think that was an experience of course it's never easy bringing two organizations together, but I think everyone realized that this is the right moment and there's a huge value to be gained from this.

8:44And I think everyone, all of us can see that the whole organization is very proud of what we have built together, right? Gemini is really the fruit of that. It's a scale. It's the fact that you can, when you build one big beautiful giant LLM that it can do so many things. And so you really do need to put together, yeah, that many people, that much compute. Infrastructure teams, data teams, evals, and so on. Yeah, better to have one of those teams than five little ones. Yeah, yeah. I mean, one thing I would say is like from the start, we wanted Gemini to be, I mean, even pre-Gemini, one of the origins of the Pathways project was to explore, you know, a single model that could do many things, a multimodal model that could deal with all different modalities, a very large model that was sparse.

9:37So you would activate different pieces of it for different kinds of things. And all three of those things are sort of represented in the Gemini models that we have today. And I think now with Omni, we've gotten the multimodal kind of aspects of the now we can even generate video. You know, we used to be able to just generate images and audio Pretty awesome because you have the full capability of this amazing reasoning model that can deal with Lots of input modalities and can like edit the video. It's just produced I think Omni is a whole new capability, right? Like I mean we had Vio and Nanobanano of course like you can do text to video text to images But like what you want is really a model that understand all the modalities of the physical world so that it can understand the physics and everything, but together with text, because there's a lot of information about world in text as well at a very high level.

10:25Corey, really quick, I have a question about this. In the sort of IO keynote, we were sort of framing Omni in the sort of like world model section, and I'm curious like how much, like is it actually have a bunch of the genie world model stuff, or is this sort of just like positioning for the next stage where it takes in anything and puts out anything, and that's sort of our representation of world model? That wasn't abundantly clear to me. I hadn't thought about it that way. I'll give my opinion. Like, Oriole has worked on these things a lot. World modeling means that you really understand the dynamics, the physics, the visuals.

10:55And then, like, you have to be able to simulate that as well. Because that simulation aspect is critical for both us understanding if the model has it right. And also, when you want to rely on the model, you want a model that is going to be able to roll forward that simulation. And the decisions that are coming out of the model is based on those future simulations. That's why I think Gemini Omni is a different category that is really transforming what we had with Gemini that is mostly understanding and text output and video that is text input and doing the video modeling into turning to a really true world model.

11:31By training jointly, there's a hope that, of course, everything will transfer and making a better text understanding model will help the world modeling aspect. But I think we're seeing this every time we try. It's not easy, but as we get the recipes right, we see that, you know, back in the day, like rolling out a complex video scene forward, consistency, all these things, you kind of had to manually think about them and almost pre-specify how to get the visuals right over time. And when you turn the thing, the object disappeared. And just by training at scale and mixing all the data, more and more, we're seeing these capabilities emerge.

12:09And that's what's exciting and sort of the main premise, I guess, we were putting forward. And now, finally, we're going to outputting also amazing, consistent 3D worlds, sounds, all the things. I mean, it feels almost impossible. If you asked me a few years ago, this approach would work. Otherwise, probably we would have done it like 10 years ago. But it did happen. Maybe more GPUs. Yes, probably, probably. More data, yes. When you say multimodal, you instinctively are drawn to human modalities like text and images and audio and video. But I think really you want the model to understand a much richer set of modalities, like understand interesting scientific data that comes in genomic sequences or in, you know, chemical structure or robotic grasping data or LiDAR data.

12:57Exposing the model to a little bit of this kind of data makes it much better at understanding it when it does encounter more of it. I feel like part of the story of Google DeepMind being able to pull off this model and actually like again this like form the formation story is actually like people and the fact that you all actually like know each other and we were talking off camera before this about when did you all like meet each other and start working together and like hear hear of each other and I'm curious for all your all your versions of that story. Maybe I can go first since I think I know people the longest so.

13:32One way to put it. For many years, I did a lot of engineering, hiring, and recruiting in the very early days of Google. So I screened all the engineering resumes that came into Google for three years or something. It was amazing. They would just bring death like a giant stack of resumes. He'd be like, no, yes, yes, no, no, no, yes. It was like extremely fast. So I didn't actually interview Noam, I don't think, but he had interviewed and gotten an offer. And I think you were debating, should you take the offer? So I called you up on the phone in 2000. And I said, hey, you know, I mean, just Chad, I want to introduce myself.

14:07And, you know, I really like the kinds of things you're excited about and working on. I think you'd really enjoy it here. And I finished the phone call. Honest question. Were you just selling at that point? And did you like, was there something like he had an offer? So I'm selling because I want him to accept the offer. Yeah. Yes. I love it. So then I did. He did. He did. And then became my office mate for like three and a half years or something. So, yeah. I mean, I remember. I remember joining and everyone got a mentor to ask questions as a new hire because there are a million things you don't know.

14:38And I would ask my mentor and every time my mentor would know the answer. And it was like, wow, everyone knows everything here. And it turns out that Jeff was my mentor and it was just that Jeff knows everything. And had written half of the code base. Yeah, so then fast forward, I guess, maybe to 2012, I think. Yeah, yeah. So Oriel had interviewed with us. And I don't think I interviewed you, but I think you had an offer. And I was trying to convince you you were considering this in another company. So I called you up and I said, hey, you should really come here. We're really doing really interesting deep learning choice, deep learning models.

15:14And we're having an awesome time. We were all in the Google Brain team, wedged in probably a 30-person office just outside on the main patio of the main four buildings of Googleplex. Somehow I managed to convince you to come, which was awesome. and yeah that was I mean I remember lots of back and forth I had like the last maybe one year of my PhD so I was just writing the thesis no LLMs at the time so you actually had to you know write every single word lots of pondering but I joined and I maybe not exactly like mentor like Noam mentioned but we started two projects one of which was distillation and so I remember I mean the code base was complex, like C++, you're like in academia, so you don't know exactly how to implement things.

15:59But the idea was clear. And literally, I remember sitting by Jeff's desk, and he was just coding the classes, like distillation and KL diversions and so on and so forth. So we didn't have coding agents at the time. But I can say maybe for a little bit, Jeff was kind of acting as like the coding agent for the project. And I mean, a hard benchmark to Yeah, I think that project was good because Jeff Hinton did some of the very early exploration on MNIST, which is like a tiny, tiny data set that he could run on his laptop. And he had some good ideas about how to get a bigger model to transfer to a smaller model.

16:32And I'm like, we got to show this thing at scale. So we trained a 50 model ensemble for 300 million images, which was a lot at that time, and 50 distinct ones. So we grouped the categories. So this one was going to be good at cars, and this one was going to be good at wild animals. Then we transferred the knowledge with distillation into a single model, and it was much more accurate than the single model you could have trained on the raw data. And by the way, at the time, I remember compute was already constraining, but all you needed to do is ask Jeff, hey, we ran out of CPU, and he would just go to some website, change a number, and we doubled it.

17:07And we did that a few times for a few months. I had super user power in the quota saying. That was nice. I missed that. Yeah, sasway exponential growth sometimes stops happening. I remember first time we really sat together, talked was actually during the acquisition discussions of DeepMind that you flew to New York and to London. Sorry. There was this moment that there were all sorts of discussions and such. There's a bunch of people in the room. But then Jeff comes to me and says, let's look at the code. Okay. So I sat down on the keyboard. I'm like, okay, don't show me anything too sensitive, but I want to see that directory.

17:46Yes. So he hooks at that directory and we go inside. I'm like, okay, let's see this file. He's like, okay. And then I go and explain, okay, here's what we are doing here. Here's what we are doing there. This is this idea. This is that idea. I mean, at the time for me, it was a big deal, right? Like I'm sitting together with Jeff. I'm explaining to him, okay, these are the ideas and this is the code and like we are walking through it. Our first code review together. Yes, exactly. Looks good to me. Was it actually, Jeff, you were just like pointing at random directories and then Koraj just happened to know.

18:17We'd seen 15 talks, which was great. At the time, I would review pretty much all the code at DeepMind. So I would know pretty much everything going on. Yeah, and I think the company was like 55 or 60 people or something. So we all flew over to London, and then we not slept super well that night. And then we go in, and we see 13 consecutive 30-minute talks. Jeff Hinton had a bad back, so he's laying on the floor in the back of the conference room. We just flattened him. I heard that story. Yeah, towards the end of the day, I'm like, okay, this seems pretty promising. Let me just see the code because we'd seen some nice slide decks.

18:56That's crazy. We need a movie about this. I feel like this would be a good movie. Actually, another thread of this reflecting back to three and a half years, maybe even longer than that. Something now, as we sit here, that is both positive surprising and also negative surprising. Like something that maybe you wish we had made more progress and it's surprising we haven't. And also something maybe we've made way more progress than you. And obviously, so much of the stuff is so hard to have imagined five years ago. But anything that sticks out for all of you? Maybe I start with positive and very timely for today.

19:32I really didn't expect we could keep doing what we've been doing generation after generation, which is to pack the intelligence of Pro back into Flash. So it's kind of like that. happened in 1.0 and you could say, well, you know, it was the first run and it was fairly suboptimal in some way. So that makes sense. We improved the recipe. But in a way, even that seems to be sometimes even accelerating, depending on which version we look, that Flash next gen just outperforms pro previous gen. I mean, just even understanding distillation, how it works, I'm still like mesmerized. How can we pack so much intelligence per byte or per parameter?

20:10has distillation like fundamentally changed in a way like is it like it's sort of and i'm not super i know of the concept i don't know the details is there like architectural improvements to the way we do distillation which is part of how we can like keep packing more in or is it like the technique is relatively the same of what you y 'all came up with originally yeah i would say it's even simpler i mean we had some you know tweak with temperatures in the soft max and and and we had to take an ensemble of models. Don't tell what we had. No, no, I won't tell what we had. I'm going to spill the recipe.

20:43Just making sure. I'm going to spill the recipe. You have a really, really good teacher, and then you have a student. But you didn't need an ensemble of 50 teachers. You just have one really good teacher, and then one student. And you pretty much use the recipe described in the original paper with some modest tweaks. But the basic spirit of the idea is pretty much the same. Let me give you the most technical explanation. It's like squeezing the lemon. You squeeze the lemon, the juice comes out, it's the good bits. You put it in a glass, which is your small model. I like that. You should read the intro of the paper.

21:19It has some poetic intro as well about larvae and insects. The original paper, that was just like soft labels. Yeah, pretty much, yeah. Anything that's sort of, you're surprised we haven't been able to pull off, given how much progress across the board Gemini has made over the last three and a half generations? I mean, the good side, like, thinking back to it, it's all good. It's all about the beginning of Google, right? Like, we have this, what was it, this one-box philosophy, right? Like, Jeff, you must remember, like, the one box for everything. Like, the search box also, you could use it for...

21:54Type in something, it would show you sports scores. Type in something else, it would show you stock quotes. Right. And like on the back end, like, you know, these were all separate, very separate back ends and like our custom built, whatever, some of them were AI-ish and some of them weren't. Spelling. Did you mean? Did you mean is like largely GNOME's starter project, I think. Oh, yeah. The user would assume, oh, there must be some brilliant general purpose AI behind the whole thing and it knows it can do all these different things. and now we actually built it. We built the general purpose AI.

22:29It's one box. It is one box. It is one box and it's like one back end. We finally have the back end for the front end and we have the right interface because we built the one box. He wants a negative thing. No, no, no. But it's obviously like people want more. Is there something that you wish that... I think you should say it's hard for us, right? Because we have been in it and for us, especially for researchers researchers like like you don't operate with negativity that much like if something doesn't work it's a learning and like like you you put on top of it from your point of view what would you have expected to see and you are not seeing what's your disappointment that's a good i wouldn't frame it as disappointment um but he has it but he has clearly i have one i'm a part engineer part researcher engineers can be more negative all right okay um i mean i felt like we would make more progress on sort of continual learning and more kind of not so structured model architectures.

23:26Like right now we have MOEs, which are like lots of experts. They're, you know, all very similar structure. I felt like a much more organic style thing would be something we would make progress on. Yeah, you always imagine that kind of like bigger architecture. I still think that could be interesting, but we are not doing that yet. But, you know, what we're doing seems to work. I'm a little disappointed. that, okay, we haven't cured every disease yet. You can't just type in, invent me a cure for cancer or something, and it'll just do it. But we're moving along. Yeah. I think, and I'm curious to get y 'all's reaction to this.

23:59I think it's not a negative thing, but it's surprising to me, it seems like how much energy and effort it takes to merge the capabilities into a single model. Obviously, and that it's a really difficult juggling act. like you merge in a new capability it's not just like it works out of the box or something like that you it's like trade off against something and you have to make some change and try to like make up those gaps and i think it's not super intuitive to me as far as from from my point of view that is one thing that i'm amazed with the models that they're still like there's insane amount of capacity in the model and we keep packing stuff like imagine that the current models are not like that much bigger than what was happening like three, four years ago, right?

24:44But like we keep packing more and more and more capability and information. So like the fact that we can do that, like there is so much room in the model. Maybe this is the negative part, but to me, like we keep doing that and it's still there is room and there's so much more room in these models. And that's why it makes me actually excited because in terms of algorithmic AI development, there is a lot of room. I really believe in that. The models have much more capacity than what we are getting out of them right now. There's going to be big innovations that's going to enable us to do a lot more with the models.

Read the full transcript

25:21Yeah, part of it, I would say, is we really need to come up with algorithmic things that just get much more out of every piece of data or example that a Model C's or every token. Because I think if you look at the efficiency of say human learning, it's a thousand times better than what our sort of LLM learning can do. Like the LLM gets to see a thousand times as much data as a really capable human, and then gets to roughly the same capability, maybe slightly better in some things and not quite as good in others, but it needed a thousand times as much data. So if we could make it so that you could get a thousand times as much information out of every example, it'd be amazing.

26:03A human has heard what? Like a billion words in a lifetime. And then, yeah, a model has been trained on trillions and can remember them. To disagree a bit, though, right? We're pre-trained. I mean, it's not like you're the first human. So anyways, there's some arguments also about that. But the source code is so small. We got like gigabytes of source code. This is one of my questions. This is why you don't want this conversation happening. By the way, I have a hard one in terms of what's been difficult. I think evaluation is very difficult. Interesting. It's been somewhat underappreciated in the community, even from the academic era that Koray was mentioning, evaluating capabilities in isolation or what are the next big things that will happen and how to evaluate in a way that is not leaked into the data sets and that users will agree with the number.

26:58There's a lot of work and progress, but I feel like that's been maybe surprisingly hard, but perhaps because we came from a table of numbers in papers and now we have users and feedback, that's just been surprising and exciting because every time you find something difficult, you get motivated by trying to fix it. But evaluation is one that needs to keep getting better. I mean, the whole dream of every AI researcher ever has been, how do we build systems that can generalize to things they've never been confronted with? And that's really, you know, even when you're training specific models on particular tasks, you want to generalize to new examples of that task.

27:39But I think what we're trying to do now is generalize to anything anyone might ask. And that is kind of a hard problem. But by having a lot of users, you get a lot of feedback about, okay, well, we're generalizing pretty well in these kinds of problems, but in these kinds of problems, we're falling short. One of the questions, controversial questions that I had for you all was, you all have obviously worked together for a long time in different capacities. What are some research things that you all still don't all agree on? And I want to preface this by, I think this is a positive thing. The beauty of, I think, having people who have different perspectives is that there's disagreement and we try different things.

28:18I'm curious if there's anything that comes to mind. I'm trying to... Or you all agree. I don't think we would all agree, but I don't think there would be big, major disagreements. Because I think in the grand scheme of design of Gemini, I think this group has experimented with all sorts of things. I think we built a lot of ideas through experimentation. I know that Jeff always had this idea of building something a little bit more flexible and has more plasticity and more fluid. We didn't get there, but it's not like we disagreed on that. It's just that I think the current systems have sort of empirically showed us the way that this is the model that we are doing.

29:02But otherwise, I don't know if we had big disagreements. at any given time each of us is kind of spending more of our effort on say one particular thing or a few things and the others are not necessarily spending much time on that thing so you know like i'm spending a lot of time on you know what should future inference hardware look like because i think that's a super important capability for us to have um and you know you're not spending much time but i describe it to you in the kitchen you're like oh yeah that sounds good like when can we have it? Reality is a good way of getting people to agree.

29:36You see experimental results and see what works and what doesn't work. I mean, in general, Gemini is quite data-driven, I would say. Like, lots of people run experiments at small scale and then say, oh, yeah, here's the results. Like, oh, that looks promising. Have you tried combining it with this thing? And you need to use your pool of researchy compute in the most effective way possible. And being data-driven is I think something like Gemini, if you think about Gemini or AI in general, it pulls in so many things, like from hardware to model design to product to everything. So I think having this group to work together is actually one of the most important factors that actually make it work.

30:16Like Jeff, as he said, focusing on hardware, like Noam is focusing on models. Oreo has been focusing on models, now going very deep on agents and doing really deep work there. And I try to focus on, okay, like where are we going with Gemini? And like, I mean, like, are we working well with the products? Are we getting that experience? And like, are we running well? So I think all of us like work together in a way that are taking care of different important areas of like, because it's a whole technology transformation that is happening. And I think having people who are deeply thinking about different aspects of this technology transformation, I think that's what makes it work.

30:59I love it. We should do predictions just so that we have something to be wrong about a year when we reflect back on this conversation. Obviously, huge amount of progress, lots of exciting things from this year's I.O. if we were sort of sitting here 2027, which sounds like a made up year, is going to be around the corner of 2027. We're going to be sitting here at IO. Why is the made up year? I mean, just like, in my mind, I'm like, 2027 just seems not real. Like it's so far in the future and it's six months away or whatever. I'm going to be 50. Wow. Well, yeah, your 50. Happy early birthday. Your 50th birthday, we'll be celebrating IO 2027.

31:36Any sort of predictions of things that you would, you are hopeful that will actually land by then from a model capability perspective or anything like that. Let's try to predict, like, IO 2027 and what are we announcing. Yeah. Let's not do that. Let's do that. Anything, well, like, directionally, directionally. Just, like, given where we are now, it's, like, you know, coding, you know, obviously we've made a huge amount of progress on coding. How, you know, will we be saturated? Will we still be spending as much time focused on it? Same thing with agents. like just given like sort of the exponential that it feels like we're on for a bunch of these different capabilities.

32:11Maybe I'll jump in. I think one thing that might be happening in a year's time is like self-learning. Self-learning is the same as continual learning or different? I think they're related. Maybe like for some, it is the same, but like we are in an era where like models are a lot more agentic and like they're very good at writing code. We use them in our research. I think slowly we are going to use them more and more in our research and there will be a point where At least at some experimentation level we are going to rely on the models to improve different parts of Gemini And I think like next year we will definitely be on that path and like probably talking about it would be my prediction Let's see.

32:54Yeah, we'll probably be able to point to like some very significant thing in our models that was generated by the models and agents working to do self-improvement. Yeah. Under the guidance of. Right. And instead of suggesting to one of your team members, hey, why don't you experiment with this a bit and let me know how it's going next week, we'll be telling the model to do that. You have agency. I mean, hard to disagree with that one. But maybe to build on the continual learning, what more as a capability, I mean, the ability of a model to, through its experience, interactions to improve without the need to kind of update its weights.

33:37So some sort of like knowledge-based update that works really well, right? Like, I mean, we have examples of this working, but I don't think the capability has seen like a, you know, steep curve in terms of being so good that this would be an obvious thing that everyone would be using and turning on, you know in the model so that's one that i'm hopeful we'll see one year seems possible yeah a lot of interesting weird problems on that to solve i feel like i i see examples all the time where you like ask even in today's era of this you ask models the question it pulls in some like random personal context that is like completely unrelated about a friend's birthday party and somehow that's related to my question that has nothing to do with it so it does it does feel like it needs like another a year sort of in our own tech bubble right like i mean because we are in the research of this like from your point of view and you are much more plugged into real world than us i would say like like what do you want to see what do you expect that's a good question this is not a interview logan episode but the maybe we should have some of them no you don't want to hear what I have to say.

34:50The model is the product. That's all I have to say. I want the models to get better. No, I think the long running stuff, I think will be really interesting to see because it does, I feel like that's like a frontier that we can like really easily track. And like, even if coding models get 20 % better tomorrow and they're really good, like I still think you'll run into limitations for like how long you want the model to like run autonomously for. And it feels like that feels like it's, you know, IO 2027. If we're able to say this model's been running for like 30 days or something leading up to IO, like that would be like, I think really surprising to a lot of people.

35:24And maybe we won't say that, but like maybe something to shoot for. But that quantity of work that is being done independently by the model would be the good. Yeah, would be surprising. And I think actually, I think it takes the full stack to pull that off. It's like, you're going to need sort of like memory systems and you're going to need continual learning and you're going to need better hardware because it's going to cost a zillion tokens to let something run for 30 days. So. Well, and also you want your better hardware to have a late seat. Yeah, exactly. Because like if it finished it in one day, you'd be way happier.

35:52Way happier than reading 30 days. 30 days is a good marketing line, but I actually, I'd be happy. Oh, another prediction is I think, well, not a prediction for announcements, but I think one thing that these agents are going to stress is that all of our tools are too slow. Yeah, yeah. Right? Like a lot of the tools that these agents rely on, if you make the model infinitely fast, you are going to limit how much you can actually speed up real work because often they involve of interactions with tools that are designed for kind of human latency. Or the frequency of working, right? Yeah, exactly.

36:2129 and a half days of those 30 days are spent like waiting. List files, green files. Check everything. Another, actually maybe another sort of meta controversial question that I'm curious. And I think, Corey, I like the research take of this, which is why I'm interested. I asked Josh this the other day, which is like five years from now, Google either has like three products or we have like 10 ,000 products. What do you think? What seems more plausible? We have one. Only one product? Yeah, it's the model. Okay. I like it. Sure. Sure. I'll take that answer. What do the rest of y 'all think? I mean, I do think if you have an incredibly capable model, it can do many, many things.

37:07And I think you saw in the search demos today at I.O. that, you know, it can even create little apps inside search that are customized for you and the visualizations and write code. So in some sense, that's, I don't know if that counts as one product or 10 ,000 products or 10 million products. If there's a bunch of users. But like on a serious note, like I feel like people want to consume information in different ways. And I think something like search is fundamental. And I think like five years from now, we will definitely have search maybe with a much more magical box. But I think the idea that people want to reach information and consume that information for themselves, that learning activity, I think it's still fundamental.

37:56So I really think that it's going to be there and like probably we'll have many many more because it will be It will be easy to do products because they're all powered by the intelligence more and more I mean, I think there's many product outlets and there's you know Smaller number of things that make those products amazing So like if you think about the glasses that were demonstrated I know that's a product But it's gonna be made better because the models are better and they understand audio better and they can speak to you better but that's a distinct product from search. Exactly. I think it's clear to us there's definitely one model powering whatever it is.

38:34I'm not an expert, but as a user, sometimes I feel like I make an active choice of what I want to do with a digital device. I want to check my calendar, email, buy something, and having that division, that might be more of a human factor rather than the technology is incapable to present all this in a single product. But I feel like the choice of what I want to do, that focus, whether it goes away or we just evolve out of it, I'm not sure. But I find myself liking the separation of concerns sometimes. So betting on one product, I would not do it at this time, at least for myself. I guess we've been talking about the informational products, products that deliver information.

39:18And there you can just talk about how do humans want to consume that information? Is it visual? Is it text? Is it glasses? Is it some kind of brain computer interface where you get like the models internal embeddings like straight into your neurons? There's something weird like that. Actor processing. But powered by, you know, things like Omni, like maybe we will get into, you know, physical products in the future and start moving atoms and not just bits. But that is a prediction for the far future. I love it. Moving atoms and not bits is the future. One more fast round. Things that you all are building.

40:00I'm curious if like sort of obviously all the AI coding stuff, like just maybe personal, less Gemini specific, anything interesting that y 'all are doing. It could not be with code to anything like physical atoms in the real world, painting, carpentry, woodworking, whatever it is. Jeff, you go first. I mean, I think I'm just enjoying some of the consumer facing products that we're putting out that are now much more capable. I made a cute little Mother's Day card for my daughter, had her for savings. So that was kind of fun. I love it. Building Mother's Day cards. As you all know, we just like, I mean, we made the decision.

40:35We moved here. So that means a new house. New house comes with all sorts of things that you need to fix, learn, adapt. so like nowadays like house diy like it goes from like home automation to like um like fixing stuff with a nail and hammer like so that spectrum and like i enjoy that like i like being able to to hands-on things i love it i'm just trying to make the model smarter build some new model architectures yeah i've been trying to build a knowledge base of lots of research that I couldn't possibly process because we were too busy building and then create a brainstorming partner to just figure out what the next big things might be.

41:19I love it. That's awesome. Well, thank you to all four of you for taking the time to sit down. Lots of controversial answers, but it was wonderful. It's fun. I made this comment last year at IO in a conversation. I think I made this to you, Coria, but I feel like IO and bringing people together and launching this stuff, like you feel the warmth of like humanity as we build, you know, this technology together and sort of, I feel this conversation made me feel that. So I appreciate it. It was wonderful to sit down and talk. Um, and thank you all. Thank you all for, for listening and for watching this episode of Release Notes.

41:51We'll see you in the next one.

From the publisher

To mark the launch of Gemini 3.5 Flash, Logan Kilpatrick sat down at Gradient Canopy with four of the people who built it: Jeff Dean, Koray Kavukcuoglu, Noam Shazeer, and Oriol Vinyals. They talked about the origin of the Gemini project, the bet on a single unified model, why each Flash generation now outperforms the previous Pro, and what's coming next.

Watch on YouTube: https://www.youtube.com/watch?v=8hfpLa5wPGo

More from Google AI: Release Notes

All 30 episodes
Gemini co-leads on project origins and what's nextGoogle AI: Release Notes · 42 min
Listen in VO