In short
Google AI: Release Notes - Episode Summary
Episode Title
Koray Kavukcuoglu: “This Is How We Are Going to Build AGI” Description In this episode, host Logan Kilpatrick interviews Koray Kavukcuoglu, the CTO of Google DeepMind and Chief AI Architect of Google. The conversation revolves around the recent release of Gemini 3, its reception, ongoing advancements in AI research, and the collaborative approach of Google towards AI development.
Key Themes
- Gemini 3 Launch Reception
- Positive Feedback: The episode opens with a discussion on the positive reception of Gemini 3, highlighting initial excitement and satisfaction with the model's capabilities.
- User Engagement: Kavukcuoglu emphasizes the importance of user feedback in shaping AI development, stressing that the technology is co-developed with users.
- Continuous Progress in AI
- Innovation is Key: Kavukcuoglu expresses optimism about ongoing innovations in AI, acknowledging that the challenges ahead will lead to further breakthroughs.
- Benchmarking: The conversation includes insights into the role of benchmarks in AI research, focusing on how they evolve as technology progresses.
- Key Areas for Improvement
- Instruction Following: A significant focus is placed on enhancing the model's ability to understand and follow user instructions accurately.
- Internationalization: Making the model accessible in various languages is deemed crucial for reaching a global audience.
- Product Scaffolding for Model Improvement
- Integrating User Feedback: Kavukcuoglu discusses the importance of integrating user feedback into the AI development process, highlighting platforms like Antigravity and AI Studio as essential for understanding user needs.
- Collaboration Across Teams: The episode underscores a collaborative culture at DeepMind, where various teams contribute to product development.
- Future Growth Areas for Gemini
- Generative Media: The rise of generative media is highlighted, showcasing how models like Nano Banana Pro are designed to handle complex tasks like image generation and text synthesis effectively.
- Unified Model Checkpoints: Kavukcuoglu discusses plans for unifying model checkpoints to streamline AI capabilities across various applications.
- Engineering Mindset and Safety
- Engineering Focus: Kavukcuoglu advocates for an engineering mindset in AI development, stressing the need for robust, safe, and reliable models.
- Safety and Security: The discussion touches on the integration of safety measures throughout the development process, indicating a proactive approach to potential risks.
- DeepMind's Culture and Collaboration
- Team Dynamics: The importance of a collaborative and supportive team environment at DeepMind is emphasized, with Kavukcuoglu expressing pride in the team's efforts and achievements.
- Learning from the Past: Reflecting on the evolution of AI at DeepMind, Kavukcuoglu shares insights on how the culture has shifted towards a more integrated approach to innovation and product development.
Conclusion Kavukcuoglu concludes by expressing excitement for the future of AI, reiterating that the next six months will be as thrilling as the previous ones. Both he and Kilpatrick acknowledge the significant strides made in AI and the collaborative efforts that have made the release of Gemini 3 possible.
Key Takeaways
- Continuous innovation is essential for advancing AI technology.
- User feedback is a critical component of the development process.
- Collaboration across teams enhances the quality and capabilities of AI products.
- A proactive engineering mindset is necessary for building safe and effective AI models.
- The future of AI holds immense potential, and the journey is just beginning.
Additional Resources
- Watch the episode on YouTube: [Google AI: Release Notes - Episode with Koray Kavukcuoglu](https://www.youtube.com/watch?v=fXtna7UrL44)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Gemini 3, we're sitting here. Reception seems super positive. The vibes of the model are good. I'm very excited about the progress. I'm excited about the research. We actually pushed the frontier on a bunch of dimensions. This is how we are going to build AGI. We want to do it the right way. And that's where we are putting all our minds, all our innovations. It's not like it's this purely research effort that's off in a lab somewhere. Like it's a joint effort with us in the world. This is a new world, right? There's a new technology that is defining a lot of what users expect. We are in some sense like co-building AGI with our customers.
0:32So all of a sudden you enable a lot more people to be builders. Bring anything to life. Bring anything to life, right? Yeah, I feel like the next six months are going to be probably just as exciting as the last six months and the previous six months before that. We are lucky to be living in this age. It's happening right now. It's very exciting.
0:57Hey everyone, welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the DeepMind team. today. It's an honor to be joined by Koray Kavachulu, who is the CTO of DeepMind and the new Chief AI Architect of Google. Koray, thanks for being here. I'm excited to chat. Me too. Yeah, very excited. Thanks for inviting. Of course, Gemini 3. We're sitting here, we've launched the model. Reception seems super positive. I think we went out and we obviously had a hunch about how good the model was going to be. Leaderboards looked awesome, but I think putting in the hands of users and actually getting out is like...
1:27That's always the test, right? Like, I mean, we have been like benchmarking is the first step. And then we have been doing tests. We have been like with trusted testers with priorities and everything. So you get a feeling that, yes, it's a good model. It's capable. It's not perfect, right? But like, I think I'm quite pleased with the reception, really. People seem to like the model and the kinds of things that I think we found interesting. They also found interesting. So like, that's good so far. Like, this is good. Yeah, we were talking yesterday. and the thread of the conversation was just around appreciating this moment that the progress isn't slowing down, which I think resonates with me.
2:07And as I was reflecting back to the last time I sat next to you, we were at I.O. as we launched 2.5 and we were listening to Demis and Sergey talk about AI and all that stuff. I feel like the progress has not slowed down, which is really interesting. When we launched 2.5, it felt like a state-of-the-art model and it felt like we had actually pushed the frontier around a bunch of dimensions. And I feel like 3.0 delivers that again. And I'm curious what the scaling conversation of can it continue, continues to go. What's your sense right now? Yeah, I mean, look, I'm very excited about the progress.
2:39I'm excited about the research. Like when you are actually there in the research, there are a lot of excitement in terms of like in all areas of this, right? Like I mean, from data, pre-training, post-training, everywhere. We see a lot of excitement. We see a lot of progress, a lot of new ideas. at the end of the day like this whole like this this whole thing is really running on innovation running on ideas right the more we do something that is impactful that is in real world that people use you actually get more ideas because like your surface area increases the kinds of signals that you get increases and i think like the the problems will get harder the problems will get more varied right and with that like um i think like we will be challenged and these kinds of challenges are good.
3:27And I think that is the driver for going towards building intelligence as well. That's how it's going to happen. I feel like sometimes if you look at one or two benchmarks you can see squeeze, but I think that's normal because benchmarks are defined at a time when something was a challenge. You define that benchmark and then of course as the technology progresses that benchmark becomes not the frontier. It doesn't define the frontier. And then what happens is that you define a new benchmark. It's very normal in machine learning, right? Benchmarks and model development is always hand in hand. You need the benchmarks to guide the model development, but you only know what the next frontier is when you get close to it so that you can define with the new benchmark.
4:10Yeah, I feel this way and there was a couple of benchmarks like HLE originally, all the models were horrible on and doing like one or two percent. And I think now the newest with DeepThink is like 40-something percent, which is crazy. ArcGi2 was originally all the models could barely do any of that. It's now like 40 plus. So it is interesting. And then it's also interesting to see, and I don't have the context on why, the benchmarks that are static that do have a little bit of like the test of time, if you will. And like, I think they are probably close to saturating, but like GPQA Diamond, as an example, like continues to stick around, even though we're eking out 1 % or whatever.
4:47There are really hard questions there. Yeah, yeah. Like, I mean, and like those hard things we are still not able to do. Yeah. Right. And they still test something. But if you think about like where we are with GPQA, like it's not like, oh, there's a, there's like you are at 20s and you need to, you need to go to 90s. Right. Like, so you're getting close there. So the number of things that it defines as unsold is like, of course, like decreasing. So at some point, it's good to find new frontiers, new benchmarks. And defining benchmarks is really, really important, right? Because if you're going to think about, like, if we think about benchmarks as the definition of progress, which does not necessarily always align, right?
5:29Like, there's this thing between, like, there's progress and then there's the benchmarks. In an ideal case, it's 100 % aligned, but it's never 100 % aligned. like to me the most important measure of progress is like we have our models in real world and scientists use them students use them like lawyers use them engineers use them and then like like people use them to do like all sorts of things writing creative writing emails easy or hard right like that spectrum is important and different topics different domains if you can actually continue delivering larger value there. I think that's progress.
6:07And these benchmarks help you quantify that. Yeah. How do you think about, and maybe even there's a particular example from 2.5 to 3 or whatever, we could choose whichever model version change you want. Where are we hill climbing? And actually how in a world where there's a zillion benchmarks now actually of you could choose where you want to hill climb, how are you thinking about for just broadly Gemini, but also maybe the pro model specifically like where do we try to help climb for that i think like there are there are several important areas right like one of them is instruction following is important the instruction following is where like the model needs to be able to understand the request of the user and to be able to follow that right like um you don't want the model just like answering whatever it thinks it should answer right so that instruction following capability is important and that's what we always do.
7:00And then for us, internationalization is important. Google is very international and we want to reach everyone in the world. So that part is important. And I feel like 3.0 Pro, at least, I was talking to Tulsi this morning and she was remarking about how incredible the model is for languages that historically we haven't been really good at, which is awesome to see. So continuously, you have to put the focus on some of these areas. They might not look like, okay, it's the frontier of knowledge, but they're really, really important because you want to be able to interact with the users there. Because as I said, it's all about getting that signal from the users.
7:36And then if you come to actually a little bit more technical domains, function calls, tool calls, agentic actions, and code, these are really important. Function calls and tool calls are important because I think it's a whole different multiplier of intelligence that comes from there. both from the point of view of the models being able to just naturally use all the tools and functions that we have created ourselves and then use it in its own reasoning but also the model writing its own tools right like you can think that like the models are in a way the models are tools in themselves as well so that one is is is a big thing like obviously like code is because not just because we are all software engineers, but also because we know that with that, you can actually build anything that is happening on your laptop.
8:28And on your laptop, it's not just software engineering that happens. Bring anything to life. Bring anything to life, right? So a lot of the things that we do right now happens in digital world, and code is the basis for that, to be able to integrate with anything that happens, pretty much anything that happens in your life. Not everything, but a lot of things. That's why these two things together, I think, like, makes up for a lot of reach, like, for users as well. I give this example of, like, wipe coding, right? I like it. Why? Because, like, a lot of people are creative. They have ideas, and all of a sudden you make them productive, right?
9:05Like, going from creative to productive in a way that, like, you can just write it down. and then like you see the application in front of you and like it is like I mean most of the time it works and when it works it's great right like I mean so cool and that loop I think is great so all of a sudden you enable a lot more people to be builders like building something like I mean like it's great I love it yeah yeah thank you for this is the AI studio pitch I appreciate we'll clip this part out we'll put it out online um one of the interesting threads that you mentioned is like how and actually the as part of this Gemini 3 moment we launched Google anti-gravity a new agent encoding platform how much do you think about like the importance of having this product scaffolding to hill climb on quality from from a model perspective obviously yeah tool calling and coding yeah yeah it's like to me it's very very important and I think like like Like antigravity as a product itself, yes, it's exciting.
10:04But like from a model perspective, if you think about it, so it's like double-sided, right? Like let's talk about first the model perspective. From the model perspective, like being able to have this integration with the end users, in this case software engineers, and learning from them directly to understand where the model needs to improve is really critical for us. I mean, it is important in areas. Like Gemini app is important for the same reason, right? like I mean understanding users directly is very very important. Antigravity is the same way, AI studio is the same way. So having these products that we work really close a bit and then understanding and learning that getting those user signals I think is really massive and Antigravity has been a like a very critical launch partner.
10:49It hasn't been long that they have joined right but in the last two three weeks of our launch process like their feedback has been really really instrumental. The same thing with search AI mode, right? Like, I mean, AI or we use even that we get a lot of feedback from there. So, like, to me, this integration with the products and getting that signal is the main driver that we understand. Like, of course, we have the benchmark. So, like, we know how to push the STEM, the sciences, the math, that kind of intelligence. But it's really important that actually we understand the real-world use cases because, like, I mean, this has to be useful in real world.
11:25Yeah. In your new chief AI architect role, you're now responsible for also making sure that we don't just have good models, but the products actually take the models and implement them and sort of build sort of great product experiences across Google. How much, obviously I think this is the right thing for users, like getting Gemini 3 and all the product services on day one is like an awesome accomplishment for Google. and I think even more so, hopefully more product services in the future. How much additional complexity from the deep mind perspective do you think it adds to try to do it? In some sense, life was simpler a year and a half ago.
12:02But we are building intelligence, right? A lot of people ask me, you have these two roles. I mean, I have these two titles in a way, but they are very much the same thing. If we are going to build intelligence, we have to do it with the products, through the products, connecting with the users. With the Kaya role, what I'm trying to do is make sure that the products in Google have the best technology that is available to them. We are not trying to do the products. We are not product people. We are technology developers. We develop the technology. We do the models. And of course, just like everyone is opinionated on anything, people are opinionated.
12:38But the most important thing for me is making the models, making the technology available in the best way that is possible, and then work with the product teams to enable them to build the best product in this AI world. Because this is a new world, right? There's a new technology that is defining a lot of what users expect and how the products should behave, what information that they should carry over, and all the new things that you can do with this new technology. So to me, it's about enabling that across Google, working with all the products. I think that's exciting, both from the product perspective, from what users getting perspective, but also from the point of view of, like, as I said, that's our main driver.
13:23Like, it's really important for us to be able to feel that user need, to be able to get that user signal. That's critical for us. So that's why I wanted to do it, that, like, this is how we are going to build AGI. This is how we are going to build intelligence, like, with the products. Yeah. That's how I think it's going to happen. This is a great tweet for you to put out at some point, because I do think it's interesting. I share this perspective that we are in some sense co-building AGI with our customers, with the other PAs. It's not like it's this purely research effort that's off in a lab somewhere.
13:56It's a joint effort with us in the world. And I think it is actually a very trusted, tested system as well. It's a very engineering mindset that I think we are adapting more and more. And I think it's important to have an engineering mindset in this one. because when something is nicely engineered, you know that it is robust. It is safe to use. So we are doing something in the real world and we are adapting all the trusted, tested, in a way, ideas of how to build things. And I think that's reflected in how we think about safety, how we think about security. We try to think about it, again, from that engineering mindset of think about it from the ground up, from the beginning.
14:42not something that comes at the end right like we don't like so when we are doing post-training models when we are doing pre-training when we are looking at our data we always have this like everyone needs to think about this like do we have a safety team obviously we have a safety team and they are bringing in all the technology that is related to them we have a security team they're bringing in all the technology but enabling everyone in Gemini to actually also heavily be part of that development process that is that is taking this as a first principle and those teams are themselves part of our post-training teams, right?
15:14So when we are developing these, when we are doing these iterations, release candidates, just like we look at like GPQA, HLE, those kinds of benchmarks. We look at its safety security measures as well. Like that's, I think that is a very, like that engineering mindset is important. Yeah, I completely agree with you. I think it also feels natural to Google, which is also helpful because of how collaborative and like big, how big the effort is now to ship Gemini models out the door. I mean with Gemini 3 I think like we were just reflecting on this like to me one of the important things is like this model has been a very team Google model.
15:49We should look into the data it might be like one of the I mean some of the like maybe the Apollo, NASA programs had a lot of people but like it is I think this massive Google global also global effort across all of our teams to make happen which is crazy. Every Gemini release like takes people from like this continent, Europe, Asia, all around the world. We have teams all around the world and they contribute. Not just GDM teams, all teams across Google. It's a huge collaborative effort. And we sim-shipped with AI mode. We sim-shipped with Gemini app. These are not easy to do because they were together with us during our development.
16:30That's the only way that on day one, we can actually go all together out at the same time the model is ready. and we have been doing that. When we say across Google, it's not just like people actively trying to build the model. All the product teams, they're doing their parts as well. Yeah, I have a, maybe this isn't a controversial question, but, you know, Gemini 3 were sort of soda on many benchmarks, a lot of benchmarks. We're sort of sim shipping across, you know, across the Google product surfaces, our sort of partner ecosystem surfaces. The reception is very positive, sort of the vibes of the model are good.
17:02if you sort of fast forward, knock on wood, if we sort of fast forward to like the next major Google model launch, like are there things that you are like still on your list of you wish we were doing X, Y, and Z? Like how does it get better than the gem or should we just enjoy the moment of Gemini 3? I think we should do more, right? Like we should enjoy the moment because like one day of enjoying the moment is a good thing. This is the launch day and like I think people are appreciating the model so like i'd like the team to enjoy this moment as well right but but at the same time every area we look at we also see gaps right like is it perfect in writing no it's not perfect in writing is it perfect in coding it's not perfect in coding i mean especially i think on the on on the area of like agentic actions and coding i think like um like i think that there's a lot more room there.
17:57That's one of the most exciting growth areas and like we need to identify where we can do more and we'll do more, right? Like I think we have come a long way. The model is like I would say pretty much like maybe 90-95 % of the people who will engage with coding in some ways. Are they software engineers or these are creative people who want to build something? Yeah. I'd like to think that this model is the best thing that they can use, right? But there are some cases probably that is, we still need to do better. Yeah. I have another sort of pointed question for coding and tool use. What do you think has it just been, if you sort of look at the history for Gemini, and obviously we had like a very multimodal focus for 1.0.
18:36And I think for 2.0, we started to make some of the like agentic infrastructure work. Like, do you have a sense of like why we, and I'll make the caveat that like, I think the rate of progress looks really strong, but like, why has it just been like a focus thing? Why we haven't been like state-of-the-art and agentic tool use from the get-go. But for example, multimodal, we have been, literally Gemini 1 was state-of-the-art and multimodal, and we've sort of held that for a long time. I don't think it was a deliberate thing. I think it was like, I mean, honestly, I think, like if anything, when I reflect back, I tie it to using the models, the development environment being closely tied to real world.
19:12The more we are tied, then we are more better understanding these like real requirements that is happening. And I think like in our journey in Gemini, we started from a point where, of course, like, I mean, like the AI research in Google is a huge history, right? Like the amount of amazing researchers that we have and the amazing history of AI research that has been done in Google. I think it's great, but like Gemini is also a journey of moving from that like research environment into this, like as we talk, this engineering mindset and getting into a space where we are really connected with the products, right like when I look at the team like I have to say I feel really proud because like this team is still majority formed by people like including me right like for five years ago we were writing papers like we were researching AI and here we are actually we are at that frontier of that technology and that technology you are developing it via products with the users it's a completely different mindset that we are building models every six months and then we are doing updates every month, month and a half.
20:16It's an amazing shift and like I think like we walked through that that shift. Yeah I love that. Gemini 3 progress has been awesome. Another thread that was top of mind is just generally sort of how we're thinking about like where the GenMedia models which I think historically have like not been a huge, I mean not that they haven't been a focus, they've always been interesting but I feel like we've had with VO3, VO3.1 with the Nano Banana model we've had like so much success from like a product externalization standpoint. And I'm curious how you think about in this like pursuit of we want to build AGI.
20:51Sometimes I think sometimes I can convince myself that like a video model is like not part of that story. I don't think that's true. I think in general you can sort of, you should understand the world and physics and all this other stuff. So I'm curious how you see all these things intertwining together. If you actually go back like 10, 15 years ago, genitive models were mostly on images, right like because like we could we could much better inspect like what is going on in terms of and also this idea of understanding the world understanding the physics as the main driver of doing generative models with images and and and then some like some of the exciting things that we have done with generative models like date back to like 10 years ago like maybe like 10 years ago feels like 20 right 20 years ago we were still doing image models right I mean that's I was hesitating a little bit but during my PhD we were doing like generative image models right like everyone was doing those at that time we walked through that like I mean we had uh we had things called like pixel cnn's right like they were like image generative models in a way what happened was um I think it is also it was also a big realization that text actually was the better domain to have very fast progress but I think it is very natural that the image models are coming back and like at GDM we have had really strong image video audio models for a long time I think that's what I'm trying to explain maybe bringing those together I think is natural so where we are going right now is we have always talked about this multi-modality right and of course naturally like we have always talked about like input output multi-modality and that's where we are going right and when you look at it as the technology progresses the architectures the ideas in between those two different domains have been merging with each other.
22:37It used to be that these architectures were very different, right? But they are getting together quite a lot. So it's not like we are forcing something in. What is happening is naturally the technology is converging. As the technology is converging, it is converging because everyone understands where to get more efficiency from, where the ideas are evolving, and we see a common path, and that common path, I think, is getting together well. So, Nanobanana is one of those first moments, right? Like where you can iterate over images, you can talk to the model. Because what happens is that text models have a lot of world understanding, right?
23:11Like from the text. They have a lot of world understanding. And then the image model has the world understanding from a different perspective. So, like when you merge those two, I think you get exciting things. Because I think people feel that this model understands the neons that they want to get through. I have another question about Nanobanana stuff. Do you think we should just have goofy names for all of our models? Do you think that would help? Not really. Look, I mean, like, I think we didn't do it on purpose. Gemini 3. If we didn't name it Gemini 3, what would you have called it? Something ridiculous.
23:41I don't know. I'm not good at names. I think I like, I mean, it was Rift Runner, right? Like, it was Rift Runner. Like, we actually use Gemini models. Those are code names. We use Gemini models to come up with those code names, too. And Nano Banana was not one of those. Like, we didn't use Gemini, right? There's a story about it. I think like it's published somewhere. Yeah. I mean, as long as these things are natural and like organic, I think I'm happy because I think the teams who are building the models, it's good for them to sort of like have that connection. Yeah. And then when we release them, like I think that just like, I mean, that happened because we were testing the model with the codename, right, on LM Arena.
24:23And people loved it. And I think, I don't know, I'd like to think that it was so organic that sort of it caught on. I'm not sure if you can create a process to generate that. I agree with you. That's my feeling. If you have it, you should use it. If you don't have it, it's good to have standard names. Yeah. We should talk about Nano Banana Pro, which is our new state-of-the-art image generation model built on top of Gemini 3 Pro. And I think the team, I think actually, even as they were sort of finishing, Nano Banana sort of like had early signal that potentially doing this in a pro capacity, like you could sort of get a lot more performance on a bunch of like more nuanced use cases like text rendering and world understanding and stuff like that.
25:09Anything sort of top of mind for, I know we're a lot of stuff going on. I think like this is like probably where we see this like alignment of different technologies is coming into play, right? Like, I mean, because with Gemini models, we have always had like every model version is a family of models. Like we have the Pro, Flash, Flashlight, like this family of models. Because at different sizes you have different compromises in terms of speed, accuracy, cost, those kinds of things. As these things are coming together, of course like we have the same experience on the image side as well. Yeah.
25:42So I think it's natural that the teams like thought about, okay, like there's the 3.0 Pro architecture. can we actually tune this model more to be like generative image using everything that we learned in the first version and increasing the size and i think like where we end up with is something a lot more capable understands really complex like some of the most exciting use cases are you have large set of really complex documents you can feed those in we rely on these models to ask questions you can ask it to generate an infographic about that as well and then it works right? So this is where this natural input modality, input output modality just comes into play and it's great.
26:21Yeah, it feels like magic. I don't know, hopefully folks will have seen the examples by the time this video comes out but I think it's just, it's so cool seeing a bunch of the internal examples being shared around. It's crazy. Yes, I agree. Like it's exciting when you see that all of a sudden, oh my god, yes, like that's sort of huge amount of text and concepts and like complicated things explained in one picture such a nice way yeah like when you see those things like it is it's nice right like you you realize the model is capable and it's yeah and it's it's the there's like so much nuance to it too which is um which is really interesting i i have a parallel question to this which is uh probably december of last year uh december 2024 before, Tulsi was promising how we were going to sort of have these unified Gemini model checkpoints.
27:14And I think what you're describing is like actually that we've gotten really close to that now where like the architecture historically was unified in terms of like image generation and oh I see. Yeah yeah and I'm curious do you think like that I assume that's like a goal is like we want these things actually mainlined into the model and there's like natural things that stop that from happening and I'm curious like if any like context or sort of high level. Yeah look I think as I said the technology the architectures they are aligning. Yeah. Right so we see that happening. At regular intervals people are trying but it's an hypothesis and like you can't be ideology based in this right.
Read the full transcript
27:50The scientific method is the scientific method. Like we try things we have an hypothesis and you see the results sometimes it works sometimes it doesn't but that's the progression that we go through. It's getting closer. I'm pretty sure near future we are going to see something getting together and I think gradually it's going to be more and more like one single model but it will require a lot of innovation right like it is hard like if you think about it the output space is very critical for the models because that's where your learning signal comes from right right now our learning signal comes from code and text that's the most of the driver of like that output space and that's why like you are getting good at there now being able to generate images is like like we are so tuned for the quality in images like it is it is a hard thing to do right like generating really like the quality of the images the pixel perfectness is hard and then images are also conceptually it has to be very coherent like every pixel both the quality matters but also how it fits with the general concept of the picture like it matters right it is harder to train something that does both the way i look at this is to me I think it's definitely possible it will be possible it's just about finding the right innovations in the model to make it happen yeah I love it I'm excited it'll hopefully make our serving situation easier too if we have uh that I don't know a single model chart it's impossible to say it's impossible I agree with you the sort of interesting thread as we sort of sit here and you know DeepMind has a bunch of the world's best AI products hopefully vibe coding and AI studio, Gemini app, anti-gravity, and sort of across Google that's happening now.
29:30We have a great state-of-the-art model with Gemini 3. We have now Banana. We have Vio. We have all these models that are sort of at the frontier. The world looked very different like 10 years ago or even like 15 years ago. And I'm sort of curious like for your personal journey to get to this point. When we were talking yesterday, you had mentioned, which I had no idea, and I mentioned and this is someone else, and they also were like, I had no idea of this. You were the first deep learning researcher at DeepMind. And I think taking that thread to this place that we're at now feels like it's a crazy jump to go from just like the fact that people weren't excited about this technology, I don't know how long ago you started at DeepMind, like 10 years?
30:122012. 13 years? Yeah. That's crazy. 13 years ago, people weren't excited about this technology to the place, or I guess DeepMind was excited about this technology to the place now where like it is literally powering all these products and is like the main thing and I'm curious as you reflect on that um what comes to mind uh well is it surprising or like was it obviously well I mean I think this is the hopeful positive outcome scenario case right like um the way I say it is like like when I was doing my PhD I think it's the same for everyone doing their PhDs you believe that what you do is important or is going to be important right like you're really interested on that topic you think that it's going to make a big impact.
30:50And I think like I was in the same mindset that's why I was really excited about DeepMind when like Demis and Shane reached out and we talked I was really excited to learn that there was a place that was actually that was that was really focused on building intelligence and deep learning was at the at the core of it. And it's actually like like me and like my friend Carl Greger actually like we were both in Jans lab in NYU we joined we joined DeepMind at the same time so just to be very very specific and then at those times it was very unnatural that you would have a deep learning focused and ai focused uh startup even yeah like i think that was very visionary and an amazing place to be like it was really uh it was really exciting and then like like i started the deep learning team it grew i think one of the things that i like i mean my approach to deep learning has always been that like a mentality of how you approach problems and the first principle it's always learning based that's what deep mind was about everything is bet on learning it was an exciting journey to start from where we were at the days and then rl and agents and everything that we have done along the way like you go into these things at least like this is how i think i go into these things hoping that a positive outcome happens but i reflect and I say that like we are lucky right like we are lucky to be living in this in this age because I think a lot of people have worked on AI or the topics that they are really passionate about thinking that this is their age and and and this is when it's going to pan out but it's happening right now and like we have to also realize that AI is happening right now not just because machine learning and deep learning but also because it's like the hardware evolution has come to a certain state, like internet and data has come to a certain state, right?
32:44So there are a lot of things that align together. And I feel lucky to be actually be doing AI and sort of like working up to this moment. I think it is like when I reflect, that's how I feel that like, yes, they were all choices that like we worked on AI and we made and I made like specific choices to work on AI. But also at the same time, I also feel very lucky at this time we are in this position. It's very exciting. Yeah, I agree with you. I love that. I'm curious, like, what are some of the, and I was watching the Thinking Game video and sort of like, see, like learning more about like, I hadn't, and I wasn't, I wasn't around for AlphaFold.
33:23So that's the only context that I have is like reading about it and seeing people talk about it. And I'm curious, like, as you reflect and having lived through a bunch of that, how things are different today versus what they were before. And I'll sort of tee you up with one example, which is what you kind of alluded to off camera right before this, which is, and this is not exactly your words. You were like, we've kind of figured out how to make these models and bring them to the world. It was like sort of an essence of what you're getting at, which I agree with. And I'm curious if that felt like, yeah, how that is similar or not to how things were for some of the previous iterations.
33:53I think how to organize or the cultural traits of what is important to be successful, to turn hard scientific and technical problems into successful outcomes. I think we learned to do that a lot with many of the projects that we have done starting from DQN, AlphaGo, AlphaZero, AlphaFold. All these kinds of things have been quite impactful and in their ways like we learned a lot on how to organize around the particular goal, particular mission, organize as a large team. Like I remember in the early days of DeepMind like we would work on a project with like 25 people and we would write papers with 25 people and then everyone would say to us come on like surely 25 people didn't work on this i would say yes they did they did right like i mean we would organ because in sciences and in research that wasn't common right yeah and i think like that knowledge that mentality i think is is is key we evolved through that i think that is really really important at the same time i think like with the latest like the last two three years as we talked right um what we have been merged like what we have merged this with is like the idea that now this is more like a like an engineering mindset where we have a main line of models that we are developing and we learn how to do exploration on this main line how to do exploration with these models the good example where i see this and i like every time i see this or think about this i feel quite happy is our deep think models those are the models that we go to the imo competition with with to the ispc competition icpc competition with and i think that's a really really cool and good example because like we do the exploration you pick these like big targets like a competition is really important right like it's really hard problems and like like kudos to every student out there who's competing in those competitions amazing stuff really and like being able to put a model there of course like you you have the urge to do something custom for that yeah we sort of what we try to do is use that as an opportunity to evolve what we have or to come up with new ideas that are compatible with the models that we have because we believe in the generality of the technology that we have.
36:05And then that's how things like DeepThink happen and then we come up with something and then we make it available for everyone. So everyone can use a model that is actually the one that is used in the IMO competition. Yeah, just to draw a corollary between what you said, the 25 people in the paper, I think now the today version of that is you look at like, I'm sure there's a Gemini 3 contributors list that will come out or is already out. Probably 2 ,500. And there's like 2 ,500 people. And then I'm sure people are - Conservatively. Yeah, I'm sure people are thinking there's no way that 2 ,500 people contributed to actually - But they did.
36:36But they did, which is crazy. And it is fascinating to see how large scale some of these problems are now. Yeah, they did. And I think it is important for us. And that's one of the great things about Google. There are so many people who are amazing experts in their areas. We benefit from that. Google has this full-stack approach. We benefit from that. So you have experts at every layer, from data centers to chips to networking, to how to run these things at scale. It comes to a state, again, going on this engineering mindset, it comes to a state that these things are not separable. When we design a model, we design it knowing what hardware it's going to run on.
37:17And we design the next hardware knowing where the models will probably go. But this is beautiful. right like i mean but coordinating this yes of course you have thousands of people working together and contributing and i think we need to recognize it and that's a beautiful thing that's great yeah it's not easy to pull off um one of the one of the interesting threads is around back to this sort of like deep mind legacy sort of doing all these different uh scientific approaches and like trying to solve these really interesting problems and today where we actually like sort of know that this technology works in a bunch of capacities and we truly just need to keep scaling it up and like and obviously there's innovation that's required to keep doing that but I'm curious like how you think about deep mind sort of in in today's era balancing you know purely doing scientific exploration versus like we're just trying to scale up Gemini and maybe we can use my favorite example for you which is Gemini diffusion as an as a sort of like an example of like that decision making come to life in some capacity like that is the most critical thing right like finding that balance is really important even now when people ask me like what is the biggest risk for Gemini and of course I think about this a lot the biggest risk for Gemini is running out of innovation because I really don't believe that like we figured out the recipe and like we're just going to execute from here I don't believe in that if our goal is to build intelligence and we're going to do that of course like with the users with the products but the problems out there are very challenging.
38:44Our goal is still very challenging and it's out there. And I don't feel like we have the recipe figured out that it's just scaling up or executing. It is innovation that is going to enable that. And innovation, you can think about it as at different scales or at different tangential directions to what you have right now. Of course, we have Gemini models and inside the Gemini project, we explore a lot. We explore new architectures, we explore new ideas, we explore different ways of doing things. We have to do that. We continue to do that. And that's where all the innovation comes from. But also at the same time, I think DeepMind or Google DeepMind as a whole doing a lot more exploration.
39:28I think it is very critical for us. We have to do those things because, again, there might be some things that the Gemini project itself might be too constraining to explore some things. So I think the best thing that we can do is both in Google DeepMind, also in Google Research. We would explore all sorts of ideas, and we will bring those ideas in. Because at the end of the day, Gemini is not the architecture. Gemini is the goal that you want to achieve. The goal that you want to achieve is the intelligence. And you want to do it with your products enabling goal of Google to really run on this AI engine.
40:06In a way, it doesn't matter what particular architecture it is. We have something currently and we have ways of evolving through that and we will evolve through that. And the engine of that will be innovation. It will always be innovation. So finding that balance or finding opportunities of doing that in different ways, I think is very critical. Yeah, I have a parallel question to that, which is at I.O., I sat down with Sergey and I made the comment to him that sort of when, and I personally felt this at I.O., which is you bring all these people together to launch these models and have this innovation.
40:38You sort of like feel the warmth of humanity as you do that, which is really interesting. And I was referencing this because of, you know, I was sitting next to you also listening to them and I sort of was feeling your warmth. And I mean this very personally because I think this translates into like how DeepMind sort of as a whole operates. I think like Demis has this as well, where it's like this deep scientific roots, but also it's just like people who are like nice and friendly and kind. And there is something interesting where like, I don't know how much people appreciate like how much that culture matters and like how it manifests.
41:15And I'm curious, like, as you think about like, like helping sort of shape and run this, how that, yeah, how that manifests for you. Like, first of all, like, I mean, thank you very much. You're embarrassing. but like i think it is important to be i believe in the team that we have and i believe in giving people like trusting people giving people the opportunity and that team aspect is important and i think like this is something that at least to my part i can say i've learned through working at deep mind as well like because like we were a small team and of course like you like it's Like you build that trust there.
41:53And then like how you maintain that as you grow. I think it is important to have this environment where people feel like, OK, like we really care about solving the challenging technical scientific problem that makes an impact that matters for real world. And I think that is still what we are doing. Right. Like Gemini, as I said, is about that. like building intelligence is a highly technical challenging scientific problem we have to approach it with that way we have to approach it with that humility as well right like we have to always question ourselves like hopefully the team feels like that too and i'm like that's why i always keep saying i'm really proud of the team that they work together amazingly well like we were just talking upstairs at the at the micro kitchen today like i mean i said to them yes it's tiring it's it's hard yes we are all exhausted but this is what it is like we don't have a perfect structure for this everyone is coming together and working together and like supporting each other it is hard but like what makes it fun and enjoyable and also like what makes you tackle really hard problems is i think to a big extent like having the right team together working together the burden is the way i see it is more like be clear about the potential of the technology that we have.
43:10I can't definitely say that 20 years from now it's the exact same LLM architecture. I'm sure it won't be. So I think pushing for new exploration is the right thing to do. As we talked about GDM as a whole together with Google Research we have to be doing with the academic research community. As a whole we have to push many different directions. I think that's perfectly fine. What is right what is wrong is a like I don't think that it's the important conversation. I think like the capabilities and the demonstrations of those capabilities in real world is the real thing that should speak for itself.
43:50Yeah. I have one last question, which is, and I'm curious to have your reflection on this as well. I feel like for me personally, like my first year and a half plus at Google felt like, which I really liked actually, this like uh google underdog story um to a certain extent which you know despite all the infrastructure advantage and all that like for me personally showing like working when when did you join uh april 2024 2024 yeah yeah so like for and also like the ai studio context so like we were building this product and like right oh yes now now i remember we had no users we had or we had 30 000 users we had no revenue we had sort of very early in the the gemini model life cycle and i think fast forward to today and like it's obviously not like I was getting a bunch of pings earlier as sort of the last couple of days as this model has been rolling out and you know from folks across the ecosystem I'm sure you got a bunch of these as well people like very I think they're really finally realizing that like this is happening but I'm curious from your perspective like what did you feel that like again I had belief that's why I joined Google that like we were going to get to this point but like did you did you feel that underdog-ness too and I'm curious like how that how you think the team will that manifest manifest for the team as we turn that corner i definitely did even before that because like um when llms really like became apparent that they're really powerful right like i felt like very very honestly i felt like we were the frontier ai lab right like in deep mind but also at the same time i felt like okay like there's something that we haven't investment invested as much as we should have as researchers and that's a big learning for me as well right like that's why i'm always very careful about like we need to cast a wide net that's really important that exploration is important it's not about this architecture that architecture and i've been very i've been very open with the team that when we started taking llms a lot more seriously and starting with like with the gemini program like two and a half years ago i think we have been always and i've been very honest with the team that like we are nowhere near what is state of the art here like we don't know how to do a lot of things there are a lot of things we know how to do but like we are not at that level yet and it's a catch-up and it has been a catch-up for a long while i feel like nowadays we are at that leadership group i feel really good and positive about the pace that we are operating at we're in a good sort of rhythm we have a good dynamic we have a good rhythm but like um yeah we have been catching up you have to be honest with yourself right like when you are catching up you are catching up you have to you have to see what others are doing and like learn what you can learn but you have to innovate for yourself and that's what we did and that's what i feel like um it's a good underdog story in the in in a sense in that way right like we innovated for ourselves and like we found our own solutions both like technology wise model wise process wise and how we run right and it's unique to us right like we run together with all of google like look at like what we are doing it's a very different scale i never saw these things as like sometimes people also say oh google is big and it is hard i i see that as like we can turn it into our advantage because we have unique things that we can do so like i'm quite i'm quite pleased where we are but we have to learn through and innovate through that that's a good way to achieve uh what we have achieved right now and like there's a lot more to do yeah right like i mean i feel like we are sort of just catching up we are just getting there there's always comparisons but our goal is to build intelligence right like we want to do that we want to do it the right way and that's where we are putting all our minds all our innovation that way yeah i feel like the next uh the next six months are going to be probably just as exciting as the as the last six months and the previous six months before that uh thank you for taking the time to sit down this was a of fun um i hope we get to sit down again before io next year uh which feels like forever but it is going to sneak up and i'm sure there's going to be meetings like next week that are like io 2026 planning to make everything happen so uh thank you for taking the time congrats uh again to to you and the deep mind team and everyone on the model research team for making gemini 3 nano banana pro everything else happen yeah thank you very much it's been amazing having this conversation It's an amazing journey as well.
48:14And glad to have all the team, but also sharing with you as well. It's great. Thank you very much for inviting me. We got a special little gift. Thank you to congratulate you and the team for making this happen. Oh, nice. Thank you very much. Very much on point. 1500 point Elo Club. First model, right? 1501 for... Yeah, first model. Very kind. Thank you very much.
48:42you
From the publisher
Join Logan Kilpatrick and Koray Kavukcuoglu, CTO of Google DeepMind and Chief AI Architect of Google, as they discuss Gemini 3 and the state of AI!
Their conversation includes the reception of Gemini 3, the ongoing advancements in AI research, and the role of benchmarks in pushing new frontiers. They explore critical areas for Gemini's focus, emphasizing instruction following, tool calls, and internationalization, alongside Google's collaborative approach to AI development.
Watch on YouTube: https://www.youtube.com/watch?v=fXtna7UrL44
Chapters:
0:00 - Intro
2:00 - Gemini 3 launch reception
4:16 - Continuous progress and innovation
6:47 - Key areas for Gemini improvement
11:45 - Product scaffolding for model improvement
13:56 - Chief AI architect role
17:04 - Engineering mindset and collaboration
18:37 - Future growth areas for Gemini
20:33 - From research to engineering mindset
23:22 - The rise of generative media
27:22 - Nano Banana Pro capabilities
29:31 - Towards unified model checkpoints
36:26 - Organizing for AI success
38:26 - Balancing exploration and scaling
41:40 - DeepMind's collaborative culture
45:21 - Innovating at Google
48:37 - Closing

