Google DeepMind's Logan Kilpatrick: Why the Model Eats the Harness

11 Jun 2026 · 51 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Google’s agentic AI push—Gemini “agent harness” (Antigravity), agentic coding, and Omni world-model video editing; how this changes product strategy, cannibalization fears, and startup opportunity.

Guest backgrounds

Logan Kilpatrick runs Google AI Studio and the Gemini API. He works on next-gen developer tooling and DeepMind/Google agentic systems; he also references internal teams and dogfooding across models.

Key claims

“Agentic Gemini era” is becoming real via an “agent harness” that powers multiple Google products (Search, Gemini app, Cloud, AI Studio). Coding is a general-purpose harness proving ground. Cannibalization is likely positive-sum: AI answers increase search activity. Google aims to maximize customer outcomes, not “eyeball time.” Agent harnesses will be “digested” into models over time (“model eats the harness”).

Notable examples

Omni edited a live stage video in real time, adding a dog that jumped onto Logan’s lap while he kept talking. AI Studio enabled “vibe coding” Android apps; he built a plant-gardening app and cited ~350,000 Android apps created since launch week.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Agentic AI

1:24 to 3:06

Logan explains the concept of agentic AI and its significance for Google.

“Logan runs Google AI Studio and the Gemini API.”

The Role of Antigravity and Product Integration

3:06 to 4:33

Discussion on how the Antigravity framework integrates with Google's ecosystem.

“And sorry, help me with antigravity is the IDE, right?”

Impact of Agentic Properties on User Engagement

4:33 to 7:18

Exploration of how agentic AI may affect user interaction and product engagement.

“Are agent harness and coding harness synonymous or not?”

Future of Product Surfaces and User Experience

7:18 to 11:51

Discussion on the future of AI product surfaces and user experience strategies.

“I have this term stuck in my head, agent-led growth.”

Evolution of Coding Agents and Enterprise Applications

11:51 to 13:49

Logan discusses the evolution of coding agents and their implications for enterprises.

“I think like how Google will end up strategically deciding, like, do our customers want to deal with us having 10 ,000 products or would it be better to only have three?”

Comparing AI Coding Tools

13:49 to 14:00

Logan addresses the competitive landscape of AI coding tools like Codex and Claude.

“From like a, you know, from the DeepMind perspective, do you think Long Horizon agents is like a KPI that matters?”

The Role of Coding Agents in AI

14:00 to 14:22

Explore how coding agents enhance business efficiency and innovation.

“deep mind, like we're doing lots of things, which we can talk more about later.”

Shifts in AI Narratives and Competition

14:22 to 16:34

Discuss the evolving narratives surrounding AI models and competition in coding.

“I'd love to shift gears a bit and talk about coding.”

Challenges in Developing Coding Models

16:34 to 17:52

Understand the complexities of creating effective coding models and their training.

“I think the folks, the group of folks who we have working on code is like, I describe it as like the Avengers of AI internally.”

The Impact of AI on Developer Capabilities

17:52 to 21:59

Learn how AI tools enhance developer productivity and ambition.

“Like 3.5 flash was like all post-training gains, which is really cool.”
Show all 26 chapters

Future of Superintelligence in Coding and Science

21:59 to 24:08

Examine the potential for superintelligence in various domains beyond coding.

“I feel like I can tackle, this is my personal experience, I feel like I can tackle more ambitious problems.”

Building Video Games with AI

24:08 to 25:45

Delve into the intersection of AI and video game development.

“Why did the editors have so many problems?”

The Future of Game Development with AI

25:45 to 28:00

Evaluate how AI could revolutionize the development of video games for everyday users.

“making a lot of video games inside AI Studio and the other developer surfaces that you have?”

Exploring World Models in Gaming

28:00 to 28:40

Learn about the evolving definitions and applications of world models in video games.

“And so this is a good segue into, I want to ask you about world models next, But do you think vibe-coded video games is more likely going to be game engine plus coding agents based?”

The Blurry Line of World Models

28:40 to 29:40

Discuss the changing nature of world models and their scalability.

“And so I think there's, again, there's actually a bunch of interesting startups like doing work, like figuring out what is the scaffolding for world models so that you can take them from these like very open ended.”

The New Era of Model Understanding

29:40 to 30:50

Examine how contemporary models differ from traditional action-conditioned models.

“And I think Demis sort of like framed it to the world, rightfully so, as a world model because of just like the level of understanding that it has of the world.”

Omni Model Capabilities

30:50 to 32:00

Discover the capabilities and iterative improvements of the Omni model.

“but like it can do a lot of those same use cases that you would describe or like visually could create with that same exact world model, which I think is what's most interesting to me.”

Generative Media and Authenticity

32:00 to 33:50

Explore the intersection of generative media and personal authenticity.

“And it's starting with like the use case that works the best right now, which is why it's the one that's available.”

Vibe Coding in AI Studio

33:50 to 35:20

Learn about the introduction of vibe coding for Android apps in AI Studio.

“what it means and i mean one of the things we've thought about for our podcast is the visuals matter as much as the content.”

The Rise of Personal Apps

35:20 to 36:40

Discuss the increase in personal app development through AI tools.

“On the coding side, you launched the ability in AI Studio for people to vibe code Android apps.”

The Changing Landscape of Models

36:40 to 38:40

Analyze how models are evolving beyond traditional structures and capabilities.

“anything vibe coded like really fly in the app store yet?”

Harnessing Models for Flexibility

38:40 to 40:40

Delve into the significance of harnesses and their future in AI models.

“what we have historically thought of as the model is not the model anymore.”

Opportunities in a Model-Dominated World

40:40 to 42:00

Evaluate how startups can thrive amidst advancing AI technologies.

“So a lot of the application companies want flexibility, which is why they're building their own harnesses.”

Exploring Capability Overhang in AI Startups

42:00 to 43:19

Discuss insights on the potential of focused startups in AI over large corporations.

“I think there's like, you know, there's that thread of capability overhang, which I think there's a huge amount of alpha in.”

Inside Google DeepMind's Culture

43:20 to 45:34

Logan shares his firsthand experience and observations of GDM's culture and leadership.

“have like established code bases and all this other stuff, because you can just like run way faster and write software quicker.”

The Role of DeepMind in Google's Ecosystem

45:35 to 49:26

An overview of how DeepMind collaborates within the broader Google ecosystem.

“And you feel that in the DeepMind culture.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So we could edit this set so it looks like we should. Yes. Yeah. I want this where we were talking off camera. Like we should do that for the intro because I think it just like makes all this stuff more capable. I've seen these examples of like such subtle nuance that like make me appreciate that it's like the world understanding playing out. I was I was giving a talk and was on stage with with my friend Tulsi who leads the model team. I had mentioned to someone in the crowd to like edit the video and they literally like took the picture, edit it. with Omni in real time. And this like dog came on the stage in the edited version.

0:34The other guests sort of like look down and see the dog. They like chuckle a little bit. This is while I'm like opining about whatever AI nonsense. Maybe they're just laughing at your jokes. Yeah, it was not my jokes. They laugh at the dog coming up. It jumps onto my lap. I sort of like acknowledge the dog. I keep talking. I'm like petting it or whatever. And just like there's like so much subtle subtlety and getting that right and the model crushed it and it's just it's very interesting and like still trying to like absorb and digest like what that means for you know the way we make content and all these other things that's so interesting

1:23I'm delighted to have Logan on the show. Logan runs Google AI Studio and the Gemini API. You spend a lot of your time thinking and building for the next generation of builders. Yes. So I'm excited to talk to you about everything from agentic AI to AI coding, world models and more today, and right off the heels of Google I.O. So what better timing? Yeah, I'm super excited. Thank you for having me. Wonderful. Let's start with agentic AI. So Sundar opened I.O. by calling this the agentic Gemini era. What does agentic AI mean for Google? Yeah, it's a good question. I think, and we were, we sort of, if you followed closely, we did sort of mention some of these things back with like Gemini 2.0, which I think was like a little bit early.

2:06And so I think this era, this like Gemini 3.5 era feels like it's actually becoming true now. And we're in the era of agentic coding or agentic products and everything agents as far as Gemini goes. I think for us, this agentic layer, and I think we announced this actually at IO, sort of being powered by the anti-gravity agent harness, is this like additional through line for Google that sort of connects all of our products that they're sort of like based on now. And so historically, like prior to Gemini, there actually like wasn't a through line for the, you know, probably sub hundred number of Google products that we have, the 50 Google products we have.

2:43There wasn't a through line. We had Gemini. It became this through line. and everything is now sort of using Gemini in some way, that's now becoming true for antigravity as sort of all of the products rebase to become sort of like agentic native products and like actually taking action on behalf of users and helping them get things done. You see this like new through line emerging, which I think is actually really, really interesting. And sorry, help me with antigravity is the IDE, right? Or the non-IDE. Yeah, antigravity is a lot of things, which I think is sort of, again, is an opportunity for us.

3:17You have sort of a core IDE. You have sort of like the agent first experience if you want it on the web. You have a CLI, you have an SDK. So I actually think, and I don't know how much we've framed it this way, but it really is an ecosystem of stuff that we built and it's designed to sort of like meet developers wherever they are. So you could use it through the Gemini API if you want to and you want a managed agent that you don't have to do any of the sort of infrastructure work for. And then the most interesting bit is like, It's not just the ecosystem of anti-gravity stuff. It's also powering, literally, the same harness is actually powering all the other Google products.

3:49So anti-gravity will be powering a bunch of agent stuff in search, in the Gemini app, across cloud and AI studio, which is really exciting. I see. So it used to be the Gemini API, so the language model was the through line in terms of how AI gets baked into every Google product. Yeah. And now it's not only the API, it's the coding harness. Exactly. that's being used in each of these products. And therefore, it's the coding agent itself that's driving more agentic properties inside the products. Yeah. Fair description? Fair description. I think more generically too, it's just like it is the agent harness.

4:22I think like coding as sort of like a specialized use case of the agent harness, I think is obviously powerful, but it is like coding has proved to be the general purpose agent harness in addition to also working really well for coding. Are agent harness and coding harness synonymous or not? There's definitely nuance. I think there's like optimization that you can squeeze out of like specializing. And actually you see this where like the, you know, technically the agent harness that gets used for the way that AI studio uses it is like a little bit specialized for, you know, the vibe coding use case.

4:51And the way that the Gemini app is using the agent harness is a little bit specialized for the sort of consumer always on 24 by 7 agent. So I think you have that base harness that like probably has like 80 percent of the same stuff. And then you specialize for coding or for whatever the use case is. How do you think about the cannibalization of the existing business, especially now that you are going much more aggressively into agentic properties? Because I could see, for example, if all you're doing is search or summarization, there's not as much of a cannibalization fear. Whereas if you're actually going through my emails, replying to them for me, am I even going through my email anymore?

5:31And so I could imagine that there's actually just fewer human eyeball hours on your products as a result of having more agentic capabilities. Is that fair or how do you think about the cannibalization? Yeah, it's interesting. I think one sort of observation I have is that like at the beginning, and I think Sundar has done a great job of sort of talking through this, is at the beginning of the sort of current AI era, like everyone assumed that AI being able to answer questions for you was going to be like negative sum for search. And actually what's ended up happening is it's been incredibly positive sum for search.

6:08Like people are searching more, people are doing more. And agents are searching too. Yeah, and agent, actually, again, there's like this whole market that spawned at the same time that agents are doing more at the same time that humans are also searching more. And so I think it will be, obviously, there's a finite amount of like human time in the world. But from like my early feelings of how a lot of this is playing out, it does feel like it's very positive. some from like an ecosystem value creation like how the human behavior aspect of it turns out i think is like somewhat clear in the next one to two years much less clear you know three to five years from now when the technology is improved and the products probably look a little bit different than the way that they do but ultimately like that is the success of product i think like we we have a bunch of conversations with demos all the time and it's like the point of building the technology is so that it can go and do stuff for you like that point like success for google like probably doesn't look like, you know, maximizing eyeball time in front of our products.

7:05It's like maximizing outcome for customers to like do the thing that they want to do so that they can go and live their life and do what they want. And so I feel like you'll probably see us go down the route of like maximizing outcomes for customers and like not maximizing eyeballs. Yeah. I have this term stuck in my head, agent-led growth. Like it seems to me, so I'm using coding agents a lot in my personal time. And, you know, I just let the agent make all the infrastructure choices for me. I'm like, I don't care which database you tell me. And the reason I ask is, it's true in coding today.

7:37I would imagine it's maybe going to be generally true for a lot of things, let's say shopping down the line. How do you think that's going to change how advertising works, how value capture works for the aggregators? It feels like it's a very similar trend. This isn't perfectly true, but a lot of these things are just like proxies of each other. Like the way that SEO works, I think like is directly correlated with like the way that like, I forgot what the term now, it's like GEO is like the generative engine optimization or whatever it's called. And so it does feel like there's a lot of correlation between the things.

8:15My guess is it looks like much less of a radical shift than I think maybe what we assume right now, just because these things compound on top of each other. If you were to grade the scale of agenticness in terms of crawl, walk, run, where are we in terms of how agentic the Google suite of products is? Yeah, that's a great question. It's definitely like crawl right now. And I think some of this is like all of the inherent product tension for Google is like you have, what, 13 billion plus user products. And so like, I actually think we have some more like labs like experiences where you're probably closer to running or walking.

8:54But I think like most of the product experience today is definitely closer to crawling. And I think that's just like the stewardship responsibility we have sort of building a product that's being used by lots of people. Like I don't think the long tail of customers are like ready to have AI running and just doing all the things like they probably they want to be in the driver's seat. They're cautiously taking the first step. And I think the Google team and like search is maybe like the most quintessential example of this. Like I think they have a lot of responsibility to actually do that in a way that it brings people along and doesn't just like change everything of how they interact with the Internet and the way they associate with products and stuff like that.

9:28So, yeah. Which products do you think are closest to the walk? That's a good question. I think Gemini app is definitely closest to walk. And so for Spark, I think having a 24-7 always-on agent, like literally going and potentially doing a bunch of actions on your behalf is definitely like one of the frontier use cases. And I think you'll see, I think like anti-gravity is another one where it's like you could have autonomous coding agents, you know, rebuilding operating systems and doing, you know, billions of tokens and spending thousands of dollars on your behalf. And I think those are, again, like more and actually like they're in GDM as well as another angle of this.

10:02So I think like GDM is taking like very much like a frontier look at this where I think like the rest of Google's products, I think, are like more incrementally getting there, which again makes makes reasonable sense to me. Yeah. Do you think that Google ends up with one, two, three product surfaces for using AI or thousands? It's tough. I think a lot of this is actually baked in just like how humans consume products. And my sense is that there's something nice about having this compartmentalization and this specialization of products where it becomes, if you end up with a product that is doing everything for you, inherently there's more work involved in using that version of the product.

10:44I think I think would be like the default state. I think maybe somebody will spin together like the truly magic experience that doesn't make that true. But I think I think the long tail of folks end up having to spend more mental energy and more time to actually like get the general purpose product to do the thing that they actually want to do versus like something nice about I click my calendar app. It just shows me my calendar like I don't need to worry and deal with anything else. This is my hot take for why slide decks have existed for so long of just like, you know, the thing, the piece of information, you want to be exactly in the same place.

11:15And I think we as humans are just actually very used to that as opposed to the idea of a generative interface sounds so cool to me. But it's like, do our brains really, isn't that just more cognitive overhead for us? It definitely is in certain cases. And I think somebody needs to, again, there's a lot of incredibly smart people in the world. And so maybe somebody will find the experience that makes it feel more natural. But to me right now, maybe not 10 ,000 is the extreme version. I'm guessing it looks more like more products going after sort of like different. And maybe the other answer is like, I don't know what it looks like for Google.

11:48For the ecosystem, it looks like a lot more products, I think. And that's really exciting. I think like how Google will end up strategically deciding, like, do our customers want to deal with us having 10 ,000 products or would it be better to only have three? We'll come down to like a strategic decision for us. That totally makes sense. When I talk to companies in the enterprise, they say, you know, everyone's talking about agentic AI, but the only place they've seen agents really working is coding agents. Do you agree or disagree with that take? Yeah, I think it depends what your bar for working is, which I think is a lot of the nuance of this.

12:22Like, I think if you're if you're truly trying to like offload very complicated tasks for for domains in which like it's the models haven't actually crossed the threshold of quality, then like I think that's definitely true. Like the it's not going to solve the problem. But this is something that I want. I wish we could like measure a good example is like open router, for example, is like measuring, you know, the the total token consumption that's happening. And so you can sort of like see these trends play out over time of like how much more intelligence is in the world, you know, now versus a year ago.

12:53In parallel, the thing that I'm actually really interested to measure is like how long is the average like thing, the average like agent run or the average task actually taking place? And I don't think it's something that they publish, but I feel like they probably have interesting data. I'm sure there's others because because I do think you're like seeing these like new model capability lands or new model drop and and it's like spiking up. And maybe the the curve is still like very low right now, but like you're seeing those like early signs of it spiking up or to like long running tasks. And all the model labs are talking about like we released this new model and it did, you know, three days of autonomous work or whatever it is.

13:32that's the extreme. But I think in practice, you're seeing that like trickling up like pretty, pretty quickly, which is really interesting. So even if the enterprises haven't felt it outside of coding, like they are going to like this year as sort of a bunch of those other use cases get much better as well. From like a, you know, from the DeepMind perspective, do you think Long Horizon agents is like a KPI that matters? Is it the KPI that matters? It definitely matters. I think for deep mind, like we're doing lots of things, which we can talk more about later. Like there's, you know, a huge portfolio of different bets that are taking place.

14:08Long-running agents obviously matters a lot. And I think also like specifically coding agents and that matters a lot. Like it clearly is an accelerant of like every other part of your business if you have a great coding model. And so making sure we have that, I think is super top of mind. Got it. I'd love to shift gears a bit and talk about coding. Yeah. Okay, I'm going to ask a hard question. A lot of my developer friends were using Claude for a long time. OpenAI saw that declared code red. Codex is now really good. I would say my friends are maybe split 50-50 now in using Claude and using Codex.

14:43I don't hear a ton of them using Gemini, which has always kind of puzzled me. What's going on with that? Yeah, it's a great question. I think there's one part of the story that I'll add, which is which which makes it even more interesting, which is December, the narrative was that Google had won. And when we landed Gemini 3, I think it was like such a such a profound improvement from a model capability perspective. I think a lot of the narrative was like Google has taken a huge leap forward and made that happen. And I think what was interesting to see sort of as a as an ecosystem participant is like how not how quickly that narrative shifted, but just like the next wind of the narrative obviously was like all the agent encoding stuff that happened over the holidays and then into January and beyond.

15:30And that was not that long ago. And so it is a... I feel like we've been in warp speed ever since. Yeah, for sure. But it's a matter of reminder of just how fast things can change. I think the observation is not unreasonable. I do think what's happening behind the scenes for us is trying to push the frontier as fast as possible on coding. And so I think anti-gravity actually is an important part of that. I think one of the takeaways is that it's actually really hard to make a great coding model for this developer use case of really long running SWE work if you don't actually have a product that does that.

16:08And so I think Google realized that. That's why the sort of like windsurf deal happened. It's why those folks came over and then ultimately built anti-gravity and sort of we've been using internally actually and Sundar showed this at IO, just like the graph of growth of token consumption inside of Google. So you sort of like, you need that engine to spin and sort of the meta comment again is like, the engine is spinning, it takes time in order to like actually make model progress. But I'm super confident. I think the folks, the group of folks who we have working on code is like, I describe it as like the Avengers of AI internally.

16:43And so like, it really is like the, some of the best people inside of google trying to push the rock up the hill on this stuff and taking it super seriously and trying to push and i think three flash um you know notwithstanding like some of the conversation about like the price and stuff like that like is sort of a step towards actually starting to bring a lot of these capabilities um and like the fruits of that labor paying off like it's a flash model that's better than any pro model we've ever released from a coding standpoint and the pro models were really good before. So there's another thread of this also, which is like everyone forgets that there's like pre-training windows.

17:20And I wonder like somebody should like track this online, which would be interesting to see. Meaning like the big run, like what clusters have been available. Exactly. The big runs are like, are an interesting thread of this. And so it might look from an external perspective that like, oh, you're super behind in some way. And And like, actually you, you miss all the context of like where the big runs are and where the large pre-training runs are. So I think that, that also like, obviously there's pre-training has historically been like a massive strength for a deep mind. Like we have some of the best people in the world.

17:52And so excited to see sort of the fruits of that labor and everything else that's happened. Like 3.5 flash was like all post-training gains, which is really cool. So a huge, a huge testament to the team, the work that that team did to actually like make the level of gains and like surpass the previous pro model, um, literally just with post-training, which is awesome. How religious are you all about dogfooding internally? Like our, for example, our deep mind folks still allowed to use other models, or is it like you guys are using the Gemini harness now and we have to make this really, really good.

18:23Yeah. There's, I mean, I think people it's so healthy to be using other models just cause like, it's, it's so sometimes hard to like actually grok what's happening in the ecosystem. If you're not to like, I use all the models, I use all the products. I think like, you know, folks across the rest of DeepMind are doing the same thing. You definitely have to use the Gemini models though. It's just like great from a feedback flywheel perspective. And it's part of how they get better is like DeepMind has and Google more broadly has like 100 ,000 plus incredible engineers who are using the models and giving feedback.

18:54And like, it should be a competitive advantage for Google because we have that scale of sort of engineering resources and like the depth of the talent and can run, you know, A-B tests and live experiments and all that stuff. So I think you have to use all the models, but I think for the majority of folks, it's like Gemini as the daily driver, which is great. Do you believe in this narrative around like a soft takeoff of like once you have a good enough agentic coding model, then it accelerates the pace of research progress and like it's a self-reinforcing cycle? It seems obvious that that's true, but maybe I'm too, I've drank in too much Kool-Aid that that's the case.

19:30Are you seeing the signs of it yet? Yeah, I mean, you definitely see some signs of this. I think the signs that are like still early is doing this from a model perspective. And I think part of the context of that is like the resource allocation for some of these like larger training runs is just like significant. And so like you definitely still have like a human in the driver's seat of making those decisions because like you're not going to accidentally, you know, take 10 ,000 TPUs to go kick off some job that like actually doesn't make that much sense. But from a product perspective, you for sure see it.

20:01I think we're seeing this on our team. We've built mobile apps using anti-gravity, and we'll launch them to the world faster than I think any team at Google has ever built a mobile app. Josh's team did this with the Gemini Mac OS app and sort of end-to-end delivered an app faster than any team had ever delivered a Mac app at Google. And it's because of identity coding. And so it's great from a product perspective. I think you've said in the past that if you could have a system that could build anything with code. Humans can't compete on the same level, and that's narrow superintelligence. Do you think we've reached that point?

20:36It is interesting. I think this narrow superintelligence example is interesting to see. Obviously, it kind of feels that way for coding right now, where coding is just so good that it does kind of feel like narrow superintelligence. I don't know. It depends how you actually end up the details of quantifying this. But I think the important thing is like, to your point earlier, it works incredibly well for code. And so it would be great if it did a bunch of other things, but it's actually just like so impactful that it can be great at code. And so I spent a lot of time just like letting that fact sort of just like wash over me because I think it's like, obviously building AGI is super important and very interesting, but like building AGI, if it sort of like takes away from the story of like the current present capability of the technology, I think is actually like kind of a bad, a bad sort of like trade-off.

21:32And so I'm trying to like always hold these two things in my head equal at the same time, which is we need to build general purpose technology, but obviously it's so impactful to have this thing. And it feels like it hasn't taken away sort of, it's been one of the best positive outcomes is that I feel like it hasn't taken away from like human developers. It really does feel like an accelerant of what human development, like I as a human developer feel like I have more agency in the world. I feel like I can tackle, this is my personal experience, I feel like I can tackle more ambitious problems.

22:03I feel like I used to kick around ideas and they were like slightly out of reach and I would just be like, ah, wouldn't it be nice? And now I have the opposite problem, which is I'm kicking around an idea and I'm like, I could probably make this even more ambitious. And sort of it does it adds a different layer of sort of responsibility or like some different layer of burden, actually, because I'm like, oh, I can't just like do the sort of MVP of this. Like, I actually need to like go 10 steps further because the technology enables me. And like resetting my my level of ambition, I think, is something that I've also spent a bunch of time thinking about.

22:41But I think that will happen in other, these like vertical super intelligence domains, which will be interesting. And it feels like we're going to get a bunch of those before we've like solved, like it's almost like jagged, like jagged super intelligence, I think is what we'll end up with. What verticals do you think we'll get super intelligence at next? That's a great question. I do spend a lot of my time, too much time probably thinking about coding these days. So I'll think for a second of like the other, the other domains. I think part of this is like things that have like better verifiability obviously are like the ones where you'll see the gains happen more quickly.

23:18So like things with like math and finance, actually like science could be a really interesting one. Like it would be fascinating to see like some of these domains where there's some level of verifiability, like actually like really start to take off, which would be cool. And I also think like an important thing in this like broader narrative about just like what a what impact AI is having on the world. Like you almost like want that to be the case in the sequencing of like things that work. You want a lot of these like really, really good, impactful, positive things for the world to happen as early on as humanly possible.

23:54So that like folks understand what the potential positive impact of the technology is. So I think science could be a really interesting one. And yeah, obviously there's all the stuff happening right now with like math proofs and stuff like that, which I'm not a mathematician. So it's somewhat over my head. But I saw a great tweet the other day. Why did the editors have so many problems? Exactly. That's a good one. I like that. That's a good like T-shirt. So funny. Okay. But speaking of Twitter, I went through your Twitter before this. So I'm going to read back another tweet at you. The good thing on Twitter is there's a public record of all your predictions.

Read the full transcript

24:26I need to turn on that auto-tweet deleting feature or whatever it is. um last october you tweeted everyone is going to be able to vibe code video games by the end of 2025 yeah did that end up being true it feels close and i think there's i mean it obviously not triple a games like you're not building a you know the next call of duty or gta yet um but i think it's it feels closer than it's ever been um and i think a lot actually a lot of the interesting bit about video games is you actually need to end up building a lot of this like other stuff like models and we were talking off camera before this like 3js is a great example of this like 3js makes a lot of things possible that weren't before but there's still all these like rough edges that like just a coding agent doesn't solve and so you need like you know sprite generation and like the models aren't very good at doing that natively and so you need like some orchestration layer and tooling in order to make that happen there's a bunch of other things like that that like are core to like the gaming video game experience that need to have a high degree of reliability that I think it feels like it's within reach, but actually like requires a lot of like product scaffolding work in order to create experiences that are like reusable and replayable and sort of like have the level of depth and requires a little bit of taste in there.

25:44Do you see people making a lot of video games inside AI Studio and the other developer surfaces that you have? Yeah. And so this was actually based on like us looking at the early data. And there was something like in AI Studio at the time, it was like 20 % of all apps that folks were making were actually games. Like people were trying to build games. A lot of it. Is that the most popular category? It's not the most popular category anymore. Just because I think like the ecosystem has shifted and like the user base has shifted. But it was a lot of a lot of games. What is the most popular category?

26:13I think it was like it's like 20 % like finance related stuff. Wow. People like counting their money that much. People like I think it's it's something around crypto, actually, I think is what people are doing a lot of stuff with with finance. A lot of like personal productivity things and a lot of gen media stuff, actually, because obviously the Google suite of gen media stuff has done a great job. But I also think GDM has sort of like a obviously Demis cares a ton about games and sort of like started his career and doing AI stuff because of games. and so I think we'll have some interesting swings at this and our team actually in Kaggle which is sort of a bunch of the AI benchmarking stuff we do in GDM sort of works with GDM to build this game arena which is sort of our way of sort of like testing progress towards AGI like using games as a proxy which again is like very deeply rooted in GDM's history so.

27:07How close do you think we are to you know rando off the street with a good idea can vibe code a really fun playable game i want to say this year i actually i think it's i think the model capability makes it possible i think this is where like i've gotten excited on the product side and you know we were again we were also talking off camera about sort of like the startups in this ecosystem because um it feels like it's possible it doesn't feel like there's a gap in model quality it feels like there's a gap in like you someone who knows what it takes to build a great game actually like putting the scaffolding together in the right way to make that possible.

27:40I think there are folks who are doing this right now. And so some of it is like a discoverability and awareness thing that like people just don't even know that they can do that. And some of it is just like maybe certain categories of model capabilities are just like slightly off and we're like, you know, weeks or months away from like that chasm being crossed and then it just like working for most people. And so this is a good segue into, I want to ask you about world models next, But do you think vibe-coded video games is more likely going to be game engine plus coding agents based? Or do you think it's more likely to be world model based?

28:17Yeah, I think what will end up happening is the definition of world models will blur, which we should talk about with Omni. And it will still, I think the like coding agent will look like some sort of world model type system. But you actually do need to make world models useful for like real things. You need like scaffolding. And so I think there's, again, there's actually a bunch of interesting startups like doing work, like figuring out what is the scaffolding for world models so that you can take them from these like very open ended. of the inherent design of world models, very open-ended spaces and like do it in a tangible way so that it's like grounded in a use case that like you could use in a reoccurring way.

29:03That could be somebody maybe will figure out the scaffolding for world models to make games possible. But like the inherent nature of world models right now, I think make it so that it's like actually not well-suited for like games in the current form. But the progress has been crazy. So who knows, maybe in like two years, the versions will be able to, but at least in the short term, It's like coding agent plus some sort of game engine, I think, is like where you'll see way more alpha from a game's perspective. That makes sense. OK, so you said the definitions of world models are blurry. Can we unpack that?

29:33Yeah, I mean, I think like Omni is an example of this. You know, we launched this at I.O. You can sort of take in any input, create any output. And I think Demis sort of like framed it to the world, rightfully so, as a world model because of just like the level of understanding that it has of the world. I think that like technically looks different than, and I'm not an architecture expert on like the way that we've done world models before, but it is different from an architectural standpoint than what's happened in the past, which I think is positive because it's getting closer to like some of the ways in which it might actually be more scalable.

30:09And historically, like it's been like super not scalable. It's like very, very expensive to run traditional like online world models. Yeah, like Genie being like. Yeah, exactly. So if you think of traditional world models as being like an action conditioned video model, then like right now, when we say world model, what we actually mean is a model that has some understanding of the world as opposed to being strictly technically a action conditioned video model. Yeah. And so the interesting thing though, is like it has understanding of the world, but then it also has that like really great. And that's where like the line is blurry to me, where it's like it can do a lot of those same use case.

30:49It's not real time right now, but like it can do a lot of those same use cases that you would describe or like visually could create with that same exact world model, which I think is what's most interesting to me. So I do feel like this like world model, video model thing is gonna change and play out in a different way than was obvious before. And how does it work under the hood? Like whatever you're able to share, like is it Gemini plus video models? Is it something different entirely? It is a single model, which I think is the important part. Like this was actually part of the original desire was like you were training like eight different models to do all of those things historically.

31:30It's like you have a text model with the baseline Gemini model. You have audio. You have music models with Lyria. You have Nano Banana. You have VO video models. We have a whole suite of audio models. And like, it would be great for us, our customers, if you just had a single model to do all those things. So it is like a new setup that sort of makes that possible. It's not like routing to a bunch of different models, which like we, you could have imagined, we could have done something like that actually before and done like a Gemini Omni model. But this is like a true Omni model. And it's starting with like the use case that works the best right now, which is why it's the one that's available.

32:08is this video editing capability. Technically, it's functional with the other things. It's just the quality isn't perfect and is not state-of-the-art, so we haven't rolled that out yet. It's also just the first crank of the model turn on Omni. It's the Omni Flash model, the first iteration. And so we'll have much, much more capable, powerful versions, which will be exciting to see. So we could edit this set so it looks like we're... Good. Yes. Yeah. I want this. Again, we were talking off camera. We should do that for the intro because I think it just makes all this stuff more capable. And I've seen these examples of such subtle nuance that make me appreciate that it's like the world understanding playing out.

32:53I was giving a talk and was on stage with my friend Tulsi, who leads the model team, who I don't know if you've ever had on before, but she's amazing. I love Tulsi. um and in i mentioned to someone in the crowd to like edit the video and they literally like took the picture edited with omni in real time and this like dog came on the stage uh and like the other in the edited version the other guests sort of like look down and see the dog they like chuckle a little bit this is while i'm like opining about whatever ai nonsense yeah that it was not my jokes they laugh at the dog coming up it jumps onto my lap i sort of like acknowledge the dog i keep talking i'm like petting it or whatever and just like there's like so much subtle subtlety in getting that right and the model crushed it and it's just it's very interesting and like still trying to like absorb and digest like what that means for you know the way we make content and all these other things that's so interesting yeah i'm i'm the biggest bull on generative media and what it means and i mean one of the things we've thought about for our podcast is the visuals matter as much as the content.

34:00For sure. That's how you catch people's attention in the first place, right? And so, okay, I'm excited to play with Omni. I'm excited too. And I think you probably feel this way as somebody who makes content, but I've historically been very, for myself personally, I don't use AI to make any content that I produce. It's all my words. It's always my voice. It's always my image and picture showing up. I feel like there's just so much alpha and authenticity. And so like, I would much rather it be me than some AI version of me. What I like so much about Omni is that it's like not changing me. It is like changing a bunch of these other bits, which are not me.

34:38Like I didn't choose any of the like set around us or the coffee table. It's like, so our words can stay the same. And like, you can change these bits that are like not personal and do something more interesting with them, which I think is really, really cool and feels like the version of what I want sort of like Gen Media to be, which is like not a bunch of like AI avatars. No Fruit Island videos? Exactly, truly. Like it really is like it's the original content. It's the person. It's like the personhood is there. It's just different and amplified. Super interesting. Okay, I'm excited to play with it.

35:14Yeah, we should send some prompts right after this and try some things. I don't mind the fruit videos though. I'm happy for a world of both. On the coding side, you launched the ability in AI Studio for people to vibe code Android apps. Yeah, yeah. I'd love to hear how that's going so far and where you plan to take that. Yeah, it's super exciting. I think one of the strategic things for AI Studio, and actually this is based on a lot of the feedback from the ecosystem and actually from developers, from others. It's like so many Google products. There's so many different ways in which you touch Google through all these different journeys of building a startup or bringing an idea to your life.

35:51And so we have this like first class principle of like, how do we bring things into AI Studio that make it so that you are exposed to other parts of the Google ecosystem without having to like go through nine different UIs across Google. And so Androids are like a great example, not only of that, but also of enabling people who wouldn't have otherwise built an Android app. And so I literally built my first Android app in AI Studio. Very cool to see. What is it? Yeah, I just did like a plant, not a crypto app, just a plant one. I was planting trees in my backyard. Yeah. And so it was just like playing around with a gardening app as I as I was kicking the tires.

36:32I haven't had my like breakthrough idea yet of what I want for a mobile app, but I'm going to I'm going to come up with something and see go compete on the app store. Have you seen anything vibe coded like really fly in the app store yet? That's a good question. It should be interesting to like see some analysis. I don't know. I'm sure it's like accelerating a lot of things on the app store, but I don't know how much, like, I don't know anyone like personally who's, who's done that. It is interesting. And I was going to make the observation too, that I think the last time I checked the numbers, we were viewing it this morning, it was like 350 ,000 Android apps built in AI studio since last week, which is crazy.

37:06Um, and like, excitingly, it's like 350 ,000 apps that like, probably no one was going to build before. A lot of these are personal too. And so this is where I think this like, maybe GenUI is like farther out there. But I think like the idea of you building software to solve your personal problem is like very real right now. And like people are doing that. It's like one of the most common use cases of a lot of these products. And being able to like unlock a bunch of the native capabilities of the phone, I think is also really interesting because you just have so much context that's like in different places.

37:39So I'm getting very excited about sort of that opportunity. And Android feels like it's becoming the platform for builders. Does it matter that something is an app versus just like the web is so powerful now? Yeah, it's also very interesting to see that play out. Web is definitely powerful. There are certain things that the operating systems have that like you just can't unlock, like lots of like native richness that actually like make experiences feel so much richer. I think about this for like text messaging, actually, that like the text messaging experience and all of the, and all the main operating systems feel way richer to me than like any AI chat app that I've ever used.

38:20Like if I could just talk to AI and whatever texting app I use, like I would be way happier than having to go to some other app because I think we're also just like conditioned on like the operating systems. So yeah, makes sense. Okay. I want to ask about the model eats the harness or the model eats the scaffolding. What are your thoughts? Yeah, I think it's true. And I think part of this is like what we have historically thought of as the model is not the model anymore. Like when I think like two years ago when LLMs were popular, it was like the model was like actually just a set of weights. It was a set of weights and it was like really like, how can you like as simple as possible send tokens in and get tokens out.

39:01And I think we've just like progressively step by step by step. We still call it the model. We still call it, you know, Gemini 3.5. You still call it GPT, whatever, and Claude, whatever. But like, it's actually not just the weights anymore. It's like an entire expanding sprawling system that's built around the weights that sort of like enable a lot of these like next generation experiences from agentic tool calling to tool, you know, like all these hosted tools, search, code execution, et cetera. You know, the models are now being spun up in containers and sort of have an agent harness and all that stuff.

39:33So the scaffolding is like oftentimes a couple of steps ahead of like where the act, what is like baked directly into the model. And then what ends up happening is like the model eats that scaffolding and it becomes part of like the native model system. And there's still value in having sort of the external scaffolding in certain cases and like search maybe is an example of this like there's lots of folks who use different search providers and there's different like use cases that you want and so like sure maybe the model can natively use search but you also want something else code execution another example of that um but it does feel like like maybe the agent harness is like the quintessential example of this right now where like everyone's like ah we got to go build a harness and like the harness is where the alpha is and like i think that perhaps won't be true at least in the way that we think of the harness today in 12 months.

40:23I think the models will have sort of just like digested a bunch of that. It'll be upstreamed into the model and the alpha will be somewhere else now. It won't be in sort of trying to spin your own harness because the model just like does it natively. But I thought that the part of the reason why people are building their own harnesses is because if you use a harness from any given model provider, you're locked in. Right. So a lot of the application companies want flexibility, which is why they're building their own harnesses. Yeah. And I think that's part of the scaffolding story is like that starts out perhaps true.

40:50but then as the model capability improves like it becomes less true over time actually i think the model that like you you don't have a generalized model if it can't use another harness and so it is and is important and i i mentioned this uh in another conversation with someone a few weeks ago but we need something like harness bench which is like actually measuring like how good are all these different models at adapting to all the different harnesses i feel like that seems like a reasonable thing we should we should measure as an ecosystem um and i'd be curious to see like what models are actually best.

41:20But I think over time, you expect they'd be able to use every harness, unless you're completely out of distribution, which in that case, you're still going to be completely out of distribution, even if you're using your own harness. So not sure it matters much. Fair enough. What about the application layer? How do you think about where independent companies can have a hope of surviving when the model eats the harness and eats the stuff around it. Yeah. It feels like there's, yeah, it's an interesting story that like both of these things feel true. Both on one hand, I, everywhere I look, I'm like, there's never been more opportunity to go and build something.

41:58At the same time, obviously the models are doing more than they've ever done before. I think there's like, you know, there's that thread of capability overhang, which I think there's a huge amount of alpha in. There's the thread of the model companies are like going after these like very general problems. And there's just like so much value in these like verticalized domains if you have expertise in that domain you sort of like know the customers you know the ecosystem like it's just you can really like run laps around even the best model labs because like focus is the like superpower of startups like if you can focus you can do anything and if you look at all of the companies that are big or doing lots of stuff like there's just not a lot of focus and for some for some reasons like rightfully so because you know maybe i'm i'm overly justifying, you know, Google strategy, but like, we just have a lot of products.

42:47We have a lot of users. We have a lot of different things going on. And so like, we actually can't focus in one domain. We have an obligation to do a bunch of things as a big company. I think that's not true for startups. And so I think like 24 months ago, we were all asking ourselves like, oh, wow, it seems like the opportunity space is shifting. And maybe it's possible one of the outcomes is there's less opportunity for startups in the future. That feels like so far in a way, not what has ended up playing out, which is really positive. If anything, it feels like there's just even more opportunity than there was.

43:18Like now coding has helped you like close the gap on like larger companies that have like established code bases and all this other stuff, because you can just like run way faster and write software quicker. The agentic like primitive is like a new category that you can sort of build products around that like actually in a lot of cases to the conversation about like the risks involved with building, like there's risk involved. And so like, what's your like the risk appetite of different companies is different. And so if you're willing to take more risk in some domains, like you can win a user cohort who's like interested in also taking risk.

43:51There's so much opportunity. Awesome. I'd love to talk about Google DeepMind's culture. And I'm curious, what does it feel like to be inside GDM right now? You know, we had Demis at AI Sense. He was so inspiring. I've heard Sergei's back. You guys have Noam Shazir back. Walk me through what it's like to be at GDM right now. It's incredible. I do try to take it all in because it's like a moment. I try to reflect as much as possible in the chaos of all the things that are happening just because there's so much cool stuff going on. GDM's culture is interesting and maybe three observations. One, back to this thread of focus, we're doing a lot of things and so i think you see sort of uh i think about this a lot like from a portfolio perspective i think we have like one of the strongest portfolios which is really exciting but you do see these moments where like another lab or another company whatever it is will like pull ahead in a certain area where like we under invested just like hadn't been focused enough in that domain um and it's cool to see like the the way we go about trying to like close that gap i very much appreciate it.

45:03I think I've watched the Demis Thinking Game documentary a few times. And you see a lot of details of that original culture and just the way that strikes work and all this stuff, which is actually really similar today. You just get a bunch of smart people together and go solve the problem. And I love that. And it's very cool to be a part of. Another one is this, I think you see the culture permeate from who the leaders are. And maybe this isn't a perfect characterization of the ecosystem, but Demis is a Nobel Prize scientist and the sort of OG of a lot of this stuff. And you feel that in the DeepMind culture.

45:48I think like Sam is like the, you know, maybe one of the world's best businessmen ever. And like you sort of see that in the open AI culture and the way that they go about the world. I don't have a strong sense of who Dario is, but like I think Anthropic is a very interesting place. And you sort of, at least as an external observer, like he seems like an interesting guy and so somewhat esoteric. And so it seems like they're sort of like that in the DNA and the culture of the company. you know the other labs are interesting um but i like this like very scientific approach to the world um and the way that like demis looks at like the reason he's doing this and the reason they started this mission was like literally to like solve disease and all these things and it's like so easy to get and again i'm always trying to pull myself out of the moment but like it's so easy to get lost in this like competitive race of who's pushing a number higher on sweet bench or whatever it is, it's very easy to lose sight of like the reason we're doing that is so that like we can solve problems that humans actually have.

46:50And there's a my favorite quote from all of Silicon Valley is something like, you know, we can't let other people make the world a better place more than we can, which is like what this moment feels like. And the Gavin Belson quote, and I think about that all the time. And it's like we're all fighting over who can make the world better, more than the other person, which just like when you frame it like that, it seems really goofy to me. And so it's very much not zero sum. And I think that's like a way of looking at the world. I think the last thing about DeepMind's culture is like we're very, it's sort of the engine room of Google, which I think is like literally the Twitter bio now of the DeepMind Twitter account, which I love.

47:30You man the DeepMind Twitter account? I don't want any responsibility manning other people's accounts online too much responsibility to do that. But it does feel like that too. So it's like on one hand, you have sort of like the deep rooted lab culture. On the other hand, you have sort of like all of these partners across the Google ecosystem that we're collaborating with everybody from Android that we talked about earlier to Google Cloud to Gmail to Workspace, et cetera, et cetera. And so it's an interesting blend of like, I think there's lots of research work happening, but like there's tons of applied work that's happening to like actually like work with some of the like the forefront customers like deploying Gemini to billion user products is a problem that like only two companies in the world have and we have 13 of those products and like we the you know Google goes through this all the time now and it's such an interesting place to like see that happen and see the innovation that takes place in order to make that actually possible and I feel like it's You can only do that inside of Google, which is really cool.

48:30Beautifully said. Did they give them a lot of heartburn when you joined and were tweeting a lot? That's a good question. Did you have to get sign off from comms? I'm very, one of the silver linings to my Google experience has been just like how great that group of like folks across marketing comms are to work with. And I think like, you know, their job is protect Google, make sure we tell the right story and make sure a bunch of bad things don't happen. And so I have a ton of appreciation and partnership with them. But it's been an incredible experience to like be able to go try to tell the story that resonates with developers in a way that feels authentic and not have a huge amount of, you know, I don't have to get my tweets approved all the time and all this stuff.

49:18like it's a very very positive culture and i think hopefully i am always trying to walk the line of uh not not burning uh the the trust and goodwill that that i've accumulated with those folks but it's been super positive because ultimately i think it's like it's really hard for google to tell this like authentic story it's just there's just like it's a big company there's a lot of people there's a lot of opinions and so you take the like magic of google and you water it down through like a lot of people and a lot of process. And you actually, you miss the beautiful story, which is like Google's doing the most interesting technology in the world and like helping our users with some of the hardest problems in the world.

49:54And it feels it's a privilege to like get to help tell that story. So it's a lot of fun. I enjoy it. I love what you're doing. I love what Josh is doing. I think you guys have put a really kind of sincere human touch on, as you put it, the most important problem of our time. So thank you. Well, wonderful, Logan. Thank you so much for joining me today. This is a very far-ranging conversation, everything from agents and coding to world models and harnesses and GDM culture and lots of nuggets here. Thank you for joining me today. This is a ton of fun. Thank you for having me. And I'm excited to see what the folks cook up of where we've been sitting this whole time, maybe in front of us.

50:31Maybe there'll be a dog. A dog something. You can make my dog dreams come true. I love it. Awesome. Thanks, Logan. Of course.

50:45Thank you.

From the publisher

The entire startup ecosystem is racing to build agent harnesses. Logan Kilpatrick, who leads Google AI Studio and the Gemini API, argues that scramble has a roughly 12-month shelf life. Models will absorb the scaffolding and run it natively, so the edge moves elsewhere. Google's own bet runs in parallel: a single agent harness, born from the Windsurf team and now called Antigravity, has become the connective tissue across search, the Gemini app, Cloud, and AI Studio — the role Gemini-the-model used to play. Logan makes the case that coding already feels like narrow superintelligence, and that "jagged" vertical superintelligence (in math, finance, and science) will arrive well before AGI. He argues Google's real goal is maximizing outcomes for users, not eyeball time. He unpacks Omni, the single model built to replace multiple separate systems Google once trained for text, audio, music, image, and video. His throughline: AI is an accelerant for human ambition, not a substitute for it.

Hosted by Sonya Huang, Sequoia Capital

More from Training Data

All 110 episodes
Google DeepMind's Logan Kilpatrick: Why the Model Eats the HarnessTraining Data · 51 min
Listen in VO