Koray Kavukcuoglu on frontier models, coding agents, and building AGI

3 Sep 2026 · 27 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Cori Kavacholu (DeepMind) discusses frontier AI priorities, Gemini model progress (3.5→3.7), agentic coding/workflows, and why there’s no single “AGI test.” He argues success comes from exploration plus user feedback, not a benchmark threshold.

Guest background

Cori Kavacholu leads Google DeepMind’s frontier AI efforts; previously helped start DeepMind’s deep learning team at the time DeepMind was small, alongside Carl Greffr. He joined DeepMind from Jan’s lab during early research.

Key claims

No definitive AGI test exists; “frontier” is the only meaningful goal. Building AGI requires turning models into agents that can execute software engineering tasks, and improving agentic actions/workflows. Google’s long-term investment and large user-facing product funnel accelerates learning.

Notable examples

Atari as an internal focus project leading to DQN; later agent milestones like Go, Chess, StarCraft, and AlphaFold. Gemini 3.7 launched three weeks after 3.6; Gemini 4 is underway.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Quest for AGI Testing

0:00 to 0:24

Discussion on the absence of a formal test for AGI.

“There's always a lot of discussion about, is there a test for AGI?”

Recent Developments in AI Models

0:43 to 3:03

Cori reflects on recent advancements in AI models and architecture innovations.

“future, Gemini 4 on the horizon, the sort of third iteration of the 3.5, post 3.5 series.”

Importance of Being at the Frontier

3:03 to 6:38

Cori discusses the significance of striving to be at the frontier of AI technology.

“So that's the parallel track effect that you are seeing.”

Co-building AGI with Users

6:38 to 8:51

Exploration of how user interaction guides the development of AGI.

“And I think in this role now that you have leading all of GDM, sort of the frontier AI folks, We actually like have a Frontier AI team and sort of the products and other stuff are now part of this story.”

The Historical Journey of AI

8:51 to 10:48

Cori reflects on the historical milestones and progress in AI development.

“When you look at the history of AI, history of machine learning, it has always been like that.”

DeepMind's Early Days and Breakthroughs

10:48 to 14:00

Cori shares insights from the early days of DeepMind and significant breakthroughs.

“And I don't think that there is a, like a sudden threshold of that.”

The Evolution of Intelligence in AI

14:00 to 15:00

Explore the transition from early AI research to practical applications in agents.

“that you can actually study intelligence and it can scale.”

Challenges and Complexities in Modern AI

15:00 to 17:19

Understand the increasing complexity and challenges faced in AI development today.

“More and more, we started building these agents.”

Surprising Successes in AI Development

17:19 to 18:46

Discover unexpected successes in AI that have emerged over the years.

“It's not like you have hundreds of these kinds of things where different scientific or technical domains are in this exponential takeoff stage.”

The Importance of the Journey in AI Research

18:46 to 20:26

Learn about the journey of AI development and its significance in achieving breakthroughs.

“I guess you have to like the journey to be able to achieve something.”
Show all 13 chapters

Building AGI at Google: A Unique Approach

20:26 to 23:14

Examine the unique context of building AGI within Google's ecosystem.

“what are the elements that will make it successful for us in the next step and the next step after that and the next step after that.”

The Future of AI Models and User Interaction

23:14 to 24:47

Discuss the future capabilities of AI models and their interactions with users.

“And there's a lot of exciting progress happening there too.”

Leadership and Progress in AI Development

24:47 to 26:15

Reflect on the role of leadership in advancing AI and ensuring its ethical use.

“an agent that is partnering with them on any kind of agentic task and workflow.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There's always a lot of discussion about, is there a test for AGI? Like, there's no test for AGI. I don't think anyone has a test that they can say, okay, like, if you get this much on this test, that means, like, you are at AGI. And I don't think that is... If you do have that benchmark, send it to us, please, because we'll use it. We'll run our models on it.

0:24Hey, everyone. Welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today, we're joined by Cori Kavacholu. Cori, I'm super excited for this conversation. We sat down probably like three or four months ago, and so I'm excited to catch up reflecting on a lot of the progress that's been made over the last few months looking into the future, Gemini 4 on the horizon, the sort of third iteration of the 3.5, post 3.5 series. Lots of progress in like literally only three weeks. We launched 3.5 at I.O., 3.6, and then three weeks later landed 3.7. Reception has been super positive.

1:01Congrats to you and the team. Yeah, it's been positive. Do you want to sort of, I don't know if there's anything you want to say about like what made that model. We like teased a little bit of architecture innovation or something like that, which I'm not sure we'll spill the beans on, but. Not really. Yeah. Going from the initial 3.0 launch, my reflection is we learned a lot in terms of understanding what it means to do coding. And not just coding, right? Like what it means to do software engineering, what it means to work with tools or work with the functions that people use every day. Basically turn this whole thing into an agent, from a model to an agent.

1:38And that transition, I think we were sort of walking around it. and that was the thing that we wanted to directly tackle. And of course, software engineering is the most critical domain and environment that you want your systems to be successful in because it's at the root of many, many different things that you can do. During that time, we learned a lot in terms of how to train a model, how to train an agent that can actually code with you. And after that, I think the steps have started becoming faster and like in an energy search project, there are many parallel tracks that is going on at the same time too.

2:15So what we are trying to do right now is combine all these learnings, of course, like add on top of them, understand what it means to do agentic actions and agentic workflows better together with new architectural improvements, together with new ideas that we have been working on for a while. Many of these things that we have done in 3.6, 3.7 go back like a year or more and then those start paying off. and you converge them into the model. And it's great to see the quality impact that we got from that. We were very excited in the run-up to 3.7. When we were doing 3.6, of course, we could see what 3.7 could be or what the next step could be.

2:58And then it came together. And internally, we started enjoying that model a lot. So that's the parallel track effect that you are seeing. There's a lot of parallel tracks. I don't know if we've done this in the past, but a few weeks ago we sort of publicly said that Gemini 4 were working on it. It's sort of the most ambitious pre-training run that we've done so far, which is really exciting. People seem excited by it. I think folks are excited internally. Lots of research going into this. Yeah, I'm curious, sort of, I think this has historically been one of our strengths and the team seems incredible and they've always done a great job.

3:32And so I don't know if there's anything top of mind or anything more you want to say about that? Yes, it's the most ambitious run, correct. And yes, so far, touch wood, it's going well. I am very excited. The team is very excited. I look at these as step by step, right? So far, everything is looking good. We are excited, but it's all about realizing that potential until we get to a moment that we are actually putting the model in front of users, first internally than externally. I'm always in a very cautiously optimistic state. I think maybe to tee you up for a very specific question, there's a sort of meme about whether or not we care about being at the frontier.

4:16I see you in meetings all the time and sort of know where your point of view is, but what's your point of view as far as whether or not, you know, you think it's important for us as GDM and as a team to like trying to be at the frontier? I was surprised. I've seen this comments too, And I can understand, obviously, our quality of current models are a little bit below the frontier. And that's fair enough. To put it very bluntly, there's nothing other than being at the frontier that is important for us. That's it. All our goals, all our focus is always about that. And when we are making decisions, when we are making prioritizations in terms of ideas, in terms of what to do and what goals we should follow, It's all about being at the frontier, and I'm 100 % certain that we will be at the frontier.

5:06We have an amazing team. I have so much trust in the team. There's a great amount of creativity and dedication in the team. Being part of Google, we have amazing resources available to us. We have the whole full stack that we can and we are able to optimize to be able to achieve our goals. If you think about Google as a company, it is a company that has always been at the frontier of investing in the next big technology. It has never shied from investing in a long-term important project, right? AI investment in Google has been not the last five years or ten years. Fifteen years ago when there was no AI, Google was investing in AI chips.

5:59Not chips only, AI chips. We weren't even talking about AI at the time. So that kind of long-term vision and investment has always been part of Google, like the quantum computing way. We talk about these things as if it's normal. but these are very technologically critical and heavily scientific and long-term ambitious endeavors that it's in the DNA of Google. So that's why I feel I have the whole trust for us to be able to achieve that frontier. That's all our goal. That's why we were set up. That's what the whole team feels. Yeah, there's nothing less than that. I love that. And I think in this role now that you have leading all of GDM, sort of the frontier AI folks, We actually like have a Frontier AI team and sort of the products and other stuff are now part of this story.

6:49And I'm curious, like, is there anything from this framing of like the Frontier is sort of what we're focused on that you're sort of like excited now that we're sort of bringing these teams together? I know there'll be a lot of business as usual for a lot of folks, but anything particularly that you're excited about? Gemini, of course, is very much focused on building AGI and delivering the goals. At the end of the day, it relies on two things. One, good execution, right? It's very detailed. It's very complex to build AGI, something that you can put in front of people that they can work with. But also at the same time, it relies on a very large funnel to get ideas from everywhere.

7:28So by bringing together GDM this way, what I'm excited about is, as you said, there's the frontier AI, there's the products, to be able to really orient and steer at least the idea, the knowledge, and the focus towards this. You don't want to continuously streamline everything online because the idea is at the end exploration. Success in building AGI still depends on innovation and exploration and channeling that in the right way through Gemini. I'm excited about being able to do that more. The path to building AGI then really depends on our interaction with the users. and when we say users yes some people are using these models in their daily lives to go through their emails but some scientists are using these models to be better scientists to do actual research to invent things that that has never been invented and getting help from the models so the spectrum of usage of these models basically covers everything that is happening in the world So that usage is our guide, right?

8:38That tells us what are the problems that are worth solving that people want us to solve. And we need to solve those problems. And that leads us towards AGI. That's how we have something general that people like interacting with, that people rely on, that people trust. Yeah. I think last time we were talking, you had mentioned this, like, co-building AGI with our customers framing, which I think has just continued to like it makes so much sense it resonates I think it like makes it clear why you know the deployment matters so much like why we this feedback flywheel matter so much there's always a lot of discussion about is there a test for AGI like there's no test for AGI I don't think anyone has a test that they can say okay like if you get this much on this test that means like you are at AGI and I don't think that is you do have that benchmark send it to us please because we'll use it.

9:27We'll run our models on it. All right. When you look at the history of AI, history of machine learning, it has always been like that. It's the progress. It's the journey that makes the difference, right? Like go back like 15 years ago, like 20 years ago, or even more, there was always some notion, some concept of if we achieve these kinds of tasks, That means that is the true intelligence. And at every time period, each and every one of those were the legitimate targets to follow up on at that point. And that's what leads to this kind of progress that is happening. And that's the progress in itself, right?

10:05Picking the next goals is also a big part of that progress. And that's what we are doing, right? And you can see that the intelligence, like from the very early days of AI being very narrow to now it's becoming very, very general. The goals that we are picking are very general goals. That's why we are building general intelligence. And I think it's going to continue like that. You go back 20 years ago, you look at what we are doing right now. I mean, it feels like, okay, maybe AI, it's already there. What people were thinking 20 years ago in terms of what the goals was, I mean, a lot of things are achieved.

10:39But I think the more important thing is, at the end of the day, building an intelligent entity that you can trust, that you can work with. and if we can get there. And I don't think that there is a, like a sudden threshold of that. It's more of a building up that trust and getting better and better at that. Yeah, I'm with you. You mentioned 20 years ago, so I want to go 20 years back. I don't know, maybe not 20 years back exactly, but you were GDM's first deep learning researcher, which is very interesting. And I want to double click on that. You were at Jan's lab with a bunch of other sort of like folks.

11:14The deep learning community was like quite small at that time. previous iterations, DeepMind CTO leading a bunch of technical efforts. The story that you were telling about how during the sort of GDM due diligence process, Jeff was sort of asking you questions about random places in the GDM code base because you literally reviewed all the code. You had to be the one to sort of explain random things, which seems like a crazy problem now in today's world where AI helps us do that. You've been sort of at the helm and sort of pushing on the frontier now for a long time, specific inside of GDM, before GDM.

11:47And so much of the world has changed. I'm curious, like, if you want to, like, reflect back on any of those, like, moments, any of the papers or technical breakthroughs. Obviously, GDM has done so much that sort of, like, stand out to you as you think back. First of all, thank you very much. Yes, I'm old. I think the journey in GDM has been very exciting. Yeah, there was a time at which me and Carl Greger, we were at Janslav. We joined GDM at the same time. I started the deep learning team and it grew from there. There was no big plan about how all this was going to happen. And just to clarify, other folks were just doing, they weren't doing deep learning.

12:23They were doing like what? Like R, just like RL. I think when you're in a startup, when you are doing research, everyone does everything, right? There were like three, four deep learning labs at the time. Bringing that deep learning background together with Carl into DeepMind. It was an amazing environment. I remember, I think in research, there were like 10, 12 people in research, and everyone is working on some really interesting idea from their own background, like a really, really interesting group of people. At some point, there was a discussion of, okay, let's pick an ambitious task that we can corral the whole research team around and try to show progress on that task.

13:03And that's how Atari started. The Atari games and trying to train an agent that can learn by itself, playing those games and being successful in those games. It was like an internal project. Because Atari at the time, there were RL researchers who were thinking that this can be a good environment to measure success. That was the first example of, you know, we are known for these focus projects in DeepMind and we carry that over to now GDM as well. That we bring a diverse group of researchers and engineers together. to solve a hard problem. That was the first example of where we did that. A large group of researchers that came together every week, we would measure progress, and that's what led to DQN.

13:44That's the first example of how deep learning and RL could come together and create a successful agent that can learn to achieve a task by itself. I think it convinced me and it convinced I think many people this is actually possible. that one, games are the right environment that you can actually study intelligence and it can scale. And two, at the time, deep learning was already very popular. Deep learning and RL together is the basis of the recipe for building intelligent agents. It sort of like opened that door. Was it that obvious? I mean, obviously it's more obvious in Hantai, but like it really felt like the path was lit then, or at least there was like a sliver of light at the end of the tunnel, perhaps?

14:30I'm very cautious about these things. I'm not sure about the path that was lit, but we all felt like it was an important moment because it gave us confidence that we should continue. When you look at the history of DeepMind, history of GDM, bigger and bigger goals have been achieved in this trajectory, right? Like, I mean, DQN was the first step, and then every step after that was bigger and bigger and bigger, right? Like, I mean, Go, Chess, being able to do those from scratch, and then using that knowledge for AlphaFold, like all sorts of things. StarCraft. More and more, we started building these agents.

15:07That was very exciting. Yeah. I want to talk about agents stuff, but I'm actually curious, like, if you contrast sort of maybe, like, some of the stuff that was happening back then to some of the stuff that's happening now. And like one example of this is like, you mentioned this before about, you know, a lot of these ideas were like purely in the research domain. And it was like, can you actually prove that this thing is possible? And it feels like there's been at least some transition to like, we know a lot of it's possible. We sort of have some amount of the recipe. It's like really just a lot of like really difficult research engineering work to like actually bring the thing to life.

15:42And so I'm curious, like, yeah, if you want to talk about the sort of like actually bringing the stuff to life now and like how, like, does it feel? super different it does feel different yeah it's also the same yeah like in a strange way because we are still doing rl we are still doing deep learning we are still pre-training these models with like very similar loss functions to what we used to do before we are still doing rl pretty much using like similar rl algorithms the fundamentals of learning did not change the fundamentals of optimization did not change but the domains kept getting harder right like just from like Atari to like chess to go to Starcraft, the domains kept getting harder and the way they get harder is in real life you have ambiguity, you have depth, you have depth of ambiguity.

16:31That's the biggest change that you need to be able to deal with that. Language in itself is much richer, much complex, much more multi-domain, multi-faceted than any constrained action space that you can work with so like of course like things became a lot more complex a lot more larger scale like the basics of it is still the same yeah but also at the same time there's a lot more room to extract the information and be successful in this domain but everything is a lot more open-ended a lot more open to interpretation and that requires a lot more intelligence to be able to deal with that to to anticipate what someone is saying for the model to be able to partner with that person or to respond to that that that person yeah i want to talk about the the future and stuff but maybe one last question on this anything that's sort of like surprising to you as you sort of like look back and like and i'm sure there's lots of stuff that's surprising new innovations that have come up but like anything that maybe like folks were working on before that like has just turned out to be like extremely successful that was not obvious it was going to be super successful or i don't know what any way you want to take that i take a step back And if you look at it statistically, it is surprising that something has done this exponential takeoff.

17:51It's not like you have hundreds of these kinds of things where different scientific or technical domains are in this exponential takeoff stage. So there's a little bit of an element of I feel lucky that we are living in this age. and all the knowledge buildup has happened up until this point. All the technological progress in terms of like chips and internet and data has happened up until this point. That like we are the ones like living through this. I quite enjoy that. But also at the same time, like, I mean, of course, like it's intense and it's competitive. Yeah. I feel lucky. I feel surprised that a lot of the things that like 20 years ago where we were exploring as students are now actually like some of the biggest practical implications in the world and I'm pretty sure a lot of people who were students at the time are feeling the same way yeah somebody sort of just describes you as somebody who like loves the journey a is that actually true and like be what that what that means in the context of like what we're doing as as GDM because I think it's actually like a important framing from like to like capture your headspace and your belief about what we're doing?

19:06I guess you have to like the journey to be able to achieve something. So yes, I like the journey. As I said, I think progress, especially in technology, like in my experience, sometimes happen through these like stack sigmoids where something happens, something gets unlocked and then you have big progress and then you put a lot of like incremental stages on top, which actually builds a lot of your understanding, which leads to another discovery. and then a little bit more of those and then more understanding and then more discovery. I think about that process a lot, about like what leads to progress, what leads to these jumps and how to achieve that in the context of any of the projects that we do.

19:48Because as much as I always like talking about we are building AGI, like we used to research AGI, we used to write papers about AI, we used to write papers about intelligence and we did that like it was it was pure idea to pure science now we are living in the world that we are actually building AGI because people are using it so you have to you have to build it you have to do it with that responsibility yeah right with that seriousness but at the same time you can't let go of the fact that it is still research and innovation that fuels all this that funnel is the critical thing, is the biggest differentiator.

20:25So I always think a lot about what makes it successful, what are the elements that will make it successful for us in the next step and the next step after that and the next step after that. And that's why I think that journey is very important. But the goal and the mission at Northstar is extremely important, right? I mean, journey only matters. You know that you are going somewhere good and valuable. And I feel very strongly and the whole team here feels very strongly that building AGI and bringing that positive impact to the world, to our users, to the world, I think is the best thing that we can do.

21:04Yeah. And you made this comment. I don't know if this was yesterday, but something about like doing it specifically at Google, I think is like actually like a really interesting way. Like it's like the mission is obviously important, but like in context of Google and all these products we have and the broader work that's happening is like such a unique way of doing this. I think it's the right place. We talked about Google having the DNA of doing these technological investments. But the other part is Google has a very large number of multi-billion people products. This is a company that really identifies itself by providing good services and good value for its users and understands that and users like using our products.

21:51And that is really valuable. And as I said, I see that as a critical element of building AGI. The more we can put Gemini in a place where we understand and learn and get feedback from that interaction, the better we will be. and also being in an organization where bringing value to the user is the utmost important thing. That's why I think it makes it the right place. Yeah. And the context of us being very user-centric, lots of users asking, where is Gemini 3.5 Pro? So I'm curious as you can sort of give us a quick spiel on where the model is as much as we're willing to say. look as i said there are parallel tracks in our model family we always have the pro the flash and the flashlight but at the same time we are seeing really fast progress with our flash models yeah we are working on 3.5 pro at the same time because it's one of those parallel tracks because research happens in multiple scales at multiple times right but also in terms of bringing the best thing forward for the users we are seeing really good progress with flash we can iterate really fast and i can see that like we talked about from 3.5 to 3.6 to 3.7 we have a flash model that is getting better and better and more and more competitive getting closer to that frontier yeah so like that is very exciting for all the researchers internally as well so like we are just using that as our progress path right now and 3.5 pro it's still like uh we are still working on it, but we are working on Gemini 4 as well at the same time.

23:29And there's a lot of exciting progress happening there too. I love it. I'm excited. Another one that was top of mind, Gemini 3 felt like a huge moment for us. We were at the frontier. And then I think, interestingly, it was almost like the frontier shifted and sort of like where there was like a lot of new emergent capabilities and all that stuff happened to be in agentic coding and all these other things. And I think we were just starting to like rev the engine and we were at the frontier in that moment and then sort of like the landscape shifted and how you think about that i like i think there are there are two things that we need to keep in mind one is by definition in a very competitive environment the frontier will always shift there will be ebbs and flows of different things and the cadence and the frequency of like um like which lab is producing their most capable model is gonna change that's number one but number two is is a fair point and we talked about that in the context of 3.5, I think we learned a lot about agentic actions and agentic workflows.

24:32And being able to bring that to life in a model, that's the process that we went through. Where I'm feeling right now, I'm feeling very, very, very comfortable and good right now where we are and our capability of understanding what users need when they are working with an agent that is partnering with them on any kind of agentic task and workflow. but we went through that process of building that i love that um i'll ask you one more which is a little bit uh hypothetical but if you could sort of like wave your magic wand and sort of get the models to do anything and not need to spend a lot of time and energy is there do you have anything on like the top of that list that you would sort of want like a capability sort of like better behavior on something um it can be it can be small or big i don't know i think like like if i had a magic wand like i would i would just make them more intelligent i think like is the models get more intelligent they do everything better yeah everything more intuitively and i think that is uh that would be excellent yeah i love that i sat down with sergey at io i think 2025 or something like that i made this comment to him this was after you you and i were sitting next to each other listening to him and demis do their their talk and i made the comment to sergey that that IO and just like some of these like moments of bringing people together, like it makes me feel the warmth of humanity.

25:50And I was, I was saying that in context of like some of the conversations that we were having. Um, and so I feel like this, it's just exciting to see, or like, I'm excited to see you like take this, this leadership role, because I, I feel that warmth of humanity from you at the same time that I feel the like progress of like your, your focus on the frontier. And so, uh, on a personal level, like I'm, I'm excited and I feel like we we've got the things, uh, going the right direction. I feel like you're, you're well positioned to sort of help us push. So, um, thank you very much. Yeah. I, I no, no, no, no fake, no, no free praise.

26:24I feel like you earned it. It's a lot of fun. Kari, this was an awesome conversation. Um, thank you for sitting down and chatting about all this stuff. I'm excited. Hopefully we'll have more to, to talk about soon. Um, and thank, thank you everyone for tuning in to this episode of release notes. We'll see you in the next one.

26:44you

From the publisher

Google DeepMind SVP and Chief AI Architect Koray Kavukcuoglu joins host Logan Kilpatrick to reflect on the journey from DeepMind's early reinforcement learning milestones to Gemini and what it means to build AGI at Google.

Watch along and learn:
* What made Gemini 3.7 Flash a breakthrough for engineers
* The ambitions of the Gemini 4 pre-training run
* What it takes to stay at the frontier of AI research
* How co-building AGI with users shapes everything the Google DeepMind team does

00:00 Intro
00:24 From models to coding agents
03:07 Gemini 4 pre-training
04:08 Why the frontier is all that matters
06:38 Leading frontier AI at Google
08:53 Why there's no test for AGI
10:58 Early days at DeepMind
15:08 Atari vs. real-world ambiguity
17:18 20 years of deep learning
18:47 The reality of engineering hill climbs
21:05 Why Google is the place to build AGI
22:19 Parallel model development
23:34 The shift to agentic coding
24:57 What makes models more intelligent

Watch on YouTube: https://www.youtube.com/watch?v=Rrr2gdbvNFU

More from Google AI: Release Notes

All 30 episodes
Koray Kavukcuoglu on frontier models, coding agents, and building AGIGoogle AI: Release Notes · 27 min
Listen in VO