Demis Hassabis on shipping momentum, better evals and world models

11 Aug 2025 · 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Google AI: Release Notes - Episode Summary

Episode Title

Demis Hassabis on Shipping Momentum, Better Evals, and World Models

Episode Description

In this episode, host Logan Kilpatrick interviews Demis Hassabis, CEO of Google DeepMind. They discuss the evolution of AI from game-playing systems to today's advanced models, the development of the Genie 3 project, and the need for robust evaluation benchmarks like Kaggle’s Game Arena to assess progress towards Artificial General Intelligence (AGI).

Key Highlights

Introduction

  • Host: Logan Kilpatrick
  • Guest: Demis Hassabis, CEO of Google DeepMind
  • Focus: The rapid advancements and releases within Google DeepMind, including Genie 3 and other projects.

Recent Momentum and Developments

  • Continuous Releases: DeepMind is releasing new advancements at an unprecedented rate, indicating robust progress.
  • DeepThink and Genie 3: These projects are central to DeepMind's current innovations, particularly in building models that understand real-world physics.

Transition from Games to World Models

  • Evolution of AI: From game-playing AI (e.g., AlphaGo) to models that incorporate reasoning and planning capabilities.
  • World Models: Genie 3 aims to create a model that comprehensively understands the physical world, crucial for developing AGI.

Genie 3 and Its Applications

  • Understanding Reality: Genie 3 is designed to generate consistent worlds, indicating a strong underlying model of physical laws.
  • Future Use Cases: Potential applications range from enhancing robotics to creating innovative forms of entertainment.

Evaluation Challenges

  • Kaggle's Game Arena: A new platform to benchmark AI models in gaming contexts, which is expected to foster competition and innovation.
  • Need for Better Benchmarks: Current benchmarks are becoming saturated; new, more complex evaluation methods are needed to test AI systems effectively.

Jagged Intelligence Concept

  • Uneven Performance: AI systems exhibit "jagged intelligence," excelling in certain tasks while struggling in others, highlighting areas that require further development.

Tool Use and System Evolution

  • Importance of Tools: The integration of tool use into AI systems is seen as vital for enhancing their capabilities.
  • From Models to Systems: The transition from traditional models to complex systems that can use tools and reasoning effectively.

Future Directions

  • Omni Models: A vision for future AI systems that can perform multiple tasks with equal proficiency, integrating different aspects of intelligence.

Conclusion

  • Vision for AGI: Continuous work toward developing a comprehensive AI that can understand and interact with the world in a human-like manner.
  • Community Engagement: Plans to increase accessibility to tools like Genie 3 for broader public use.

Key Takeaways

  • The rapid pace of AI development at Google DeepMind signifies a pivotal moment in the industry.
  • The understanding of physics through "world models" is crucial for AGI.
  • Evaluation methods need to evolve to keep pace with advancements in AI capabilities.
  • There is an increasing focus on the use of tools and system integration within AI models.
  • The concept of "jagged intelligence" highlights the current limitations of AI systems, emphasizing the need for further research and development.

---

This episode provides a deep dive into the current landscape of AI as envisioned by one of its leading figures, offering insights into the path towards AGI and the essential elements required for its fruition.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today we're joined by Demis Asabas who's the CEO of Google DeepMind. We're pretty much releasing something every day. It's hard to keep up, even internally. I feel like we're shipping lots of stuff. DeepThink, IMO Gold Medal, Genie 3, which the reception has been super incredible. We want to build what we call a world model, which is a model that actually understands the physics of the world. Building the technology and then also actually putting it into people's hands is like such a beautiful combination. We're starting to see sort of convergence of those models together into kind of what we call an Omni model, which can do everything.

0:34So we also announced our partnership with Kaggle to launch Game Arena. As the name suggests, the best models get pitted against each other. As they get better, the test will get harder automatically. And I think it's just one of probably many new benchmarks that are going to be required as we get closer to AGI. It's an amazing, exciting moment in the industry, I would say.

0:59Hey, everyone. Welcome back to Release Notes. My name is Logan Kilpatrick from the Google DeepMind team. Today we're chatting with Demas Asalbis, who's the CEO of Google DeepMind. Demas, thanks for being here. Excited to chat about all of the releases and progress that we've had in the last few months. Great to be here. Let's start with just the unprecedented amount of momentum. I feel like we're shipping lots of stuff. DeepThink, IMO, Gold Medal, Genie 3, which the reception has been super incredible. Like 50 other things in the last two months that I feel like we've already forgotten about because things are moving so quick.

1:31I'm curious to just like get your high level sense of that that progress and momentum that's happening. Yeah, it's fantastic to see. We've been sort of building up to the speed, I think, of our shipping and progress over the last couple of years. And I think this is that you're seeing the results of that now. But it's an amazing and kind of exciting moment in the industry, I would say. Things to be, you know, things coming out every day. We're pretty much releasing something every day. it's hard to keep up even internally and the field as a whole. So it's exciting to see. And I'm really proud and pleased with some of the latest work we put out there.

2:07Yeah. How are you thinking about DeepThink specifically? Obviously, one of the things I've been most excited about is the model is actually available to a version of the IMO Gold model is available for Gemini app subscribers. People can actually put their hands on the model, which I think has historically, we were pinging about genie stuff. And I think that's building the technology and then also actually putting it into people's hands is such a beautiful combination. So from a DeepThink perspective, how are you thinking about it? Well, look, I think the advent of the thinking models is a little bit a hark back to our original gaming work on things like AlphaGo and AlphaZero.

2:46So we always worked on, from the beginning of DeepMind actually, the history of our work has always been with agent-based systems. So what we mean by that is systems that can complete a whole task, right? And mostly in the early days, playing a game really well. And there's an objective and you basically, you know, you have this model, you know, model of, today we have modeled multimodal models, really powerful ones that model language and everything around us. But in those days, we would have gaming models. And then you need some thinking or planning or reasoning capability on top. And this is obviously the way to get to you know AGI and then of course you could once you have thinking you can do deep thinking or extremely deep thinking and and and then sort of have parallel uh planning in you know you can do sort of planning and thoughts in parallel and then collapse on onto the best one and then move make a decision and then move on to the next one um there's still i think quite a lot of innovation that's required there um but it's exciting to see the the rate of progress even in the thinking part of things.

3:44And obviously for things like maths, for coding, for scientific problems, also for gaming, you're going to need to process and plan and basically do this thinking and not just output the first thing that the model comes up with. That's unlikely to be good enough. So you want to sort of go back and refine your own thought processes, which is in effect what the thinking systems do. Yeah, I had not seen the thinking game and I watched it probably a week and a half ago or something like that. And I was scribbling down notes as I was watching and I was like, wait, Demis and the DeepMind team were ahead of all of this stuff.

4:22I was like, and also there's just so many interesting parallels to like, as you were trying to scale up RL to solve like previous problems, like what that looks like today. And like the like data bottleneck for for um alpha fold is a good example and like similar to like how we have that problem today with like human expert data for some of these domain specific tasks like coding or just like outside of scientific domains very how much of that do you like does it feel like deja vu as we're like now solving these problems with language models and scaling up in other domains yeah i think we were we were always on the uh i i think it's sort of uh become clear we were always on the right track with you know we were the first to sort of use rl seriously basically that was one of the first bets we made back in 2010, along with deep learning.

5:07And of course, our Atari work, which was the first thing, you know, our first sort of landmark result, was the first real deep RL system that could actually do something interesting and useful. In that case, play, you know, Atari games from the 1970s, just from the pixels on the screen and be better than any human can play it. But importantly, play any Atari game, right? So out of the box. So it's really what I think sort of proved to the field that these new techniques were ready for to be scaled up and actually you know be useful. So you know we've always had this kind of thing in mind and thinking is something if you play chess like I do when you're very young you kind of that's all you're thinking about is how to improve your own thought processes.

5:51How does your thought process work and of course that leads at least for me led me to thinking about neuroscience, how the brain works and then AI as this amazing tool and also try to distill intelligence into a digital artifact. And there's still a long way to go. The kinds of systems we have today, they're very good at some things, but they're still pretty flawed at other things and relatively simple things. So they're impressive. They can get gold medals in IMO, which is unbelievable really if you think about it just from the natural language description and these by the way these are just gemini models with deep think and a bit of extra thinking they're not they're not specialized around these these tests and yet they're really really good but on the other hand they can still make simple mistakes in high school maths or simple logic problems or simple games if they're posed in a certain way so that must mean there's still something kind of missing and these are kind of, you know, uneven intelligences or jagged intelligences.

6:54Some dimensions, they're really good. Other dimensions, you know, their weaknesses can be exposed quite easily. Yeah, I want to come back to this. But actually, before we do that, can we double click on Genie 3? And I think there's an interesting segue around models aren't great at playing games. And yet, I saw a bunch of people commenting on Genie 3. And the reaction was like, absolute awe, Like people are like, this is, you know, I saw some very extreme comments being like, we're in a simulation. This is like proof that like anything, anything is possible because the Genie demos are so good.

7:28So how, and it also obviously ties to solving RL with games. Like how much, if, again, if you had to look back and then now reflecting on like the Genie 3 moment, do you think, like, has that turned out how you would have expected? I feel like it's not obvious to me that making models good at playing games results in the world model stuff that we have today. Well, there's several. The extrologe is several branches of kind of research and thinking coming together, ideas coming together. So on the one hand, we've always used board games as a challenging domain to improve our AI algorithm ideas. we used to use computer games a lot for both as challenges but also to create synthetic data so we used to you know we and we still do use a lot of simulated environments very realistic ones traditionally built ones actually like 3d game engines to create more training data for our systems to understand the physical world the reason we're doing that is we want to build what we call a world model which is a model that actually understands the physics of the world, right?

8:36The physical structure, how things work, materials, liquids, and, you know, even behaviors of, you know, living objects, animals, human beings, right? How that's obviously a critical part of our world. We don't just live in language and maths. There's the physical world that we exist in. And so if you want, you know, an AGI clearly needs to understand the physical world, partly also. So it can operate in the physical world, whether that's robotics. I mean, that's what's holding robotics back. It needs a world model or things like Project Astra, you know, our Gemini Live project about having a universal assistant that can assist you in everyday life, maybe exist on your phone or glasses and help you in your everyday life.

9:20Clearly, that also would need to understand the spatial temporal context that you're in. So you need a world model to really understand the world and how it operates. and that's one of the ways to prove that you've got a good world model is to be able to generate the world so there's many ways to test the efficacy and the and the depth of your world model but one great way to just is to just get get it to reverse it and sort of generate something about the world at the world uh like a you know uh you turn on a tap and some liquid comes out of it or there's a mirror and can you see yourself in the mirror all of these things and and and that's what Genie sort of going towards is building that world model and then expressing it and actually be able to generate worlds that are consistent.

10:06And that's the surprising thing about Genie 3 is that, you know, you look away, you come back, and that part of the world is the same as you left it, which is mind-blowing, really. And it shows that it has a very good underlying model of how the world works. How do you think people will use Genie? Is it like, is the intent for us to just like leverage it to help make Gemini and some of our other like robotics initiatives like better and scale that up or like do you think there's actually like uh obviously like maybe the i could see people playing with it in some cases but do you think there's any like yeah it's so it's so exciting in multiple dimensions which is the um uh one of course is we're already using it for our own training so we have a a games playing agent called simmer simulated agent that uh i can out the box sort of take the controls and play an existing computer game right in some cases uh well in some cases not so well uh but actually what's interesting here is that um you can put that simmer agent into genie 3.

11:06so you've got basically one ai playing in the mind of another ai it's it's pretty crazy to think about so you know simmo is deciding what actions to take and what you can give it a goal, like go and find the key in the room. And then it will sort of send out commands as if it's playing a normal computer game. But actually at the other end is Genie 3 generating the world on the fly. So there's one AI generating the world and there's another AI inside of it. So it could be really useful for, of course, for creating sort of unlimited training data so i can imagine that being very useful for things like robotics but just training our general agi systems but also of course it's it's got a lot of potential in an applied sense too uh for the future of interactive entertainment i i've got many ideas that you wouldn't be surprised of what types of next generation incredible games can be made um and then and maybe some new type of entertainment that we haven't really thought of before that somewhere between film and game some new genre of entertainment.

12:13And then finally, maybe the most interesting thing from my point of view as a scientist is, what does this actually tell us about the real world and physics? And maybe things like simulation theory. You know, one has to, when you're working on this like late at night and you're sort of generating these entire worlds and you're sort of thinking through how this technology is working, you have to sort of also consider, and I've always been doing that in my career of like, well, what's going on in the real world? What's the nature of reality? And actually, that's the thing that's driven me through my whole career to build AI as this amazing tool for science.

12:45And I think things like VO3, our video model, and video audio model, and Genie 3 really tell us something about the nature of reality if we look at it with a slightly different lens. Yeah, I love that. I think this is actually a perfect transition to what you were saying before about jagged intelligence and on one hand we have this like mind-blowing system where the worlds are being generated and all this stuff on the other hand you take gemini out of the box and you ask it to play chess and like i am i was telling you off camera like i'm terrible i i know the rules of chess but i'm not actually good at it yeah and i think i could be our models at chess right now so how do you and in some cases like they can't even follow the rules so we also announced gdm's partnership with kaggle to uh launch game arena and have the models have a place to go and play a bunch of different games and test out the capabilities.

13:33I'm curious to get your reaction to that. Well, look, it's really interesting. So this is a wider thing of like, okay, so now all of our systems, ours, Gemini, but also our competitors are getting better and better. Our systems can do some amazing things, generate simulated worlds from text prompts, understand videos, all of these cool things, solve math problems, do things in science. but we all I think intuitively all of us have played around with these chatbots and it's you see the edges of what they can do quite easily right and in my opinion this is one of the things that's missing from these systems being full AGI is the consistency you shouldn't it shouldn't be that easy for the average person to just find a trivial flaw in the in the system you know used to be counting the number of R's in strawberry right I think we managed to solve that now but they're not you know there are some still pretty trivial things that these systems you like a school kid would trivially do that these systems can't.

14:31So why is that is a good question. There's probably some missing capabilities in reasoning and planning and memory that maybe one or two new innovations are still needed in those domains other than just scaling. But it's also partly maybe we need benchmarks, better benchmarks to eke out what these things are good at versus what they're not good at. And they're very general, these systems, but including Gemini, but a lot of the benchmarks we use are starting to get saturated. So you look at some of the standard mass benchmarks like AME, I think the latest result with DeepThink was 99.2%. So you're sort of getting into the region of very diminishing returns, and there may even be an error in the test.

15:16And so they're getting rapidly saturated. And so we're in need of new, harder benchmarks, but also broader, in my opinion. You know, understanding world physics and intuitive physics and other things that we take for granted as humans and we find easy, things like physical intelligence as well, actually. We don't have really good benchmarks for these things and also some safety benchmarks too, you know, testing for traits that you don't want, like, you know, deception, things like this. So I think there's actually really amazing work to be done in creating benchmarks that are really meaningful, that test slightly more complicated or subtle things than the sort of brute force school exam type things that we have today.

16:05And that's why I'm so excited about Game Arena because, and it is going a little bit back to our roots, of course, which is why we came up with it. But a lot of the reasons we started with games still hold today. So firstly, they're very clean testing grounds, right? You get ELOs, you can get scores very easily. It's very objective. They're very objective measures of performance, right? There's no subjectivity, you know, A, B testing with human beings deciding the ratings and so on. You know, I just think it's super scientific in that sense, right? The other thing is they automatically scale with the capabilities of the systems.

16:45So because the systems are playing each other in tournaments, and this is pretty fun to watch, actually, even at this level that they're at. That's the whole point of Game Arena, as the name suggests, is that the best models get pitted against each other. So what we hope is that that will actually drive a lot of progress because the systems are, you know, today, none of the AI systems are very good at gaming and, you know, things like chess, but even simpler games than that. That's an interesting question. Why is that? And I'm sure they'll rapidly progress. Now we have a measure for it with Game Arena.

17:20But as they get better, the test will get harder automatically. So you don't have to, it's not like, you know, Amy or GPQA, where you have to come up with harder science questions, and then who's going to create those questions? Are they leaked on the internet already? You know, each game is unique because it's created by the two players. So there's a kind of uniqueness about that. So that's also nice for testing. And then the final thing is, just like we did with our own early games work, as the systems get better and better, you can introduce more and more complex games into the game arena. So we started with chess for obvious reasons.

17:58It's the classic one we test AI on. It's close to my heart, of course, but the idea is we're going to expand it to potentially thousands of games and then you'll get an overall score. So we're not really looking for systems that just play one game really well. They should be able to play across all games to a good standard. And it could be computer games, as well as board games. And even more interestingly, eventually, maybe the AI system should be inventing their own games and then sort of teaching them to the other AI systems have to learn them. So it's like learning a new game, so it never existed before.

18:33So there's no way you could sort of overfit on training data or anything like that. So I have a lot of ideas of these sort of multi-agent environments that eventually Game Arena could end up supporting. So I think it's going to be a really important benchmark that will kind of be long lasting. And I think it's just one of probably many new benchmarks that are going to be required as we get closer to AGI to make sure that we've actually covered the space of cognitive capabilities. Yeah, one of the big challenges, and this has been my reflection, and I want your reaction to this in the last couple of years, is that as I thought more about evals, it's become very clear that actually most of life's problems are eval problems.

19:14And you don't think about this. Like it's not unless you're doing AI and training models, like I don't, but I'm like, you know, performance in your job is an email problem. Like how you look at all these other things are all email problems. And it is interesting that like we don't, the like human equivalence of solving these problems. Like I, the, the game arena stuff is nice because like there is this empirical, like there's truth and that the system has all these constraints. But as we sort of extrapolate out, you sort of like, I imagine outside of the game domain, like you lose some of this.

19:43Like I was thinking about like, how do we make RL environments for all these other tasks that humans do as an example? And it becomes hard because it's like, what is the source of truth in those cases? And I'm curious, like for non-game environments, like how do we get to start to capture those things? Yeah, I mean, that's always been the hard challenge with reinforcement learning has been, you know, where in domains that are more messy or real world like, how do you specify the reward function or the objective function that you're trying to optimize. And I think in our world and as humans, we don't have single objective functions, right?

20:23It's very messy. In fact, if I was to tell you, what are you optimizing for? On any given day, you might give a different answer, right? And I think we're sort of multi-objective and we're continually weighting those different objectives differently against each other, depending on other states, like your emotional state, your physical environment and, you know, where you are in your career, all of these things. So we but somehow we muddle through, right, with our brains and we sort of figure out roughly what the right north star is. And I think there's our systems, our general systems are going to have to do that, too, where they they sort of learn to interpret may be what the human user's trying to achieve, and then figure out how does that translate into a set of useful reward functions at which to optimize against.

21:15And so there's lots of experiments here about going on about metacognition or meta-RL, where you actually have another system on top that tries to work out what the reward functions are for the secondary system to optimize against. And a lot of those things are still very much research problems. But I think they're again, we used to again research those maybe like, you know, 10 years ago when we were in the middle of our games work with AlphaGo and AlphaZero. But I think a lot of that is going to come back in now. I feel like we should just start doing it now. It feels like all the things that DeepMind was doing 10 years ago are now the like exactly whatever.

21:53Yeah, it's the forefront of what everyone is trying to do. I want to back to this like thinking trend and also related to the game trend. we historically have had all these different like scaling dimensions of models, you have scale pre-training, post-training. The data has been scaling up, the compute was scaling up. Then we had this sort of like reasoning scale up, which is a lot how a lot of the new like DeepThink is enabled by reasoning scaling up. It also feels now like tools are this like new scaling dimension, whereas you give the models tools and more powerful tools and different tools, they're able to do a bunch of stuff.

22:27And I'm curious how that sort of new scaling dimension, like ties to this worldview of what we've been doing with games and some of these simulated RL environments. Like, is there a world where you like give the model a physics simulator and like it. That's a tool that it has access to. Yeah, I think tool use is going to be, you know, and is one of the most important capabilities for these AI systems. A lot of the thinking, the reason the thinking is part of the systems is very important is because you can use tools during the thinking, right? You can call search. You can use some math program.

23:03You can do some coding, come back, and then update your planning on what you're going to do. So I think that's still actually fairly nascent at the moment, but I think that's going to be incredibly powerful once that becomes really reliable and we work out. and the systems become good enough, they can use pretty sophisticated tools very reliably. And then the interesting thing comes is, what do you leave as a tool versus put into the main system, the main brain, so to speak? Now, with humans, it's easy because we're physically, you know, we're physically constrained. So anything that's not in our body is a tool, right?

23:40So there's no question about what's a tool, what's our brain. But with a digital system, you can actually kind of, those things can get blurred. So should it be in the main model, the capability, for example, to play chess or something? Or do you just use Stockfish or AlphaZero as a tool? And that tool could also be an AI system. It doesn't have to be a piece of software. It could actually be something like AlphaFold or whatever. And the question comes, I think, whether that capability helps other capabilities. So, for example, math and coding, we do put in the main model, the main Gemini model, because that seems to lift all boats, right?

24:21If you get good at coding, you're good at math, then your reasoning capabilities are just better across the board. And I suspect that may also happen with things like chess. But on the other hand, you don't want to put too much specialized data in your general model because that might harm other things. So it's very much an empirical question. Does adding that capability in help the other capabilities? If it does, then do it. If it harms the other general capabilities, then maybe consider using it as a tool. Interesting. One of the questions that developers, like through the developer lens, like people who are building with our models are always asking is around sort of like, it's become very clear.

25:01And you actually just said this, like as the model is reasoning, it's actually making use of tools and it's doing all this stuff. The models historically were like weights. You give a token, you get a token out. Now it feels like the models are like becoming these entire systems themselves and how people actually build applications on top of the models is changing because the model is just doing more for you out of the box. I'm curious, like that model transitioning from just being wait somewhere to like actually becoming a system, like resonates with your worldview of like how progress is happening.

25:33and we'll see that continue. And then I don't know if you have suggestions for people who are building stuff as they're thinking about what do I build to this point of what do I do as a tool versus what is the model just going to have empirically as part of it? Yeah. I mean, you're right. The models are improving fast and as they're gaining the tool capability, it's sort of, along with planning and thinking, it's kind of exponentially increasing what the system might be able to do, because obviously it can combine tools in novel ways and combinations. I think one thing you can think about is like, what are the sorts of tools that are going to be incredibly useful for AIs to use and start building those and providing those?

26:16So I think there's a lot of potential there. The agent itself, even with tool use, is not necessarily enough to be a whole product. So there's still a lot of productization, I think that has to be done on top. Now, the hard part, and we've talked about this before, is in this new world is you've got, I think it requires very interesting skills from a product manager or product designer type of role because you've got to sort of design, say your product's coming out in a year, you've got to be really close and understand the technology well to kind of intercept where that technology will be in a year's time and design for that, right?

26:52And I think also whatever product polish you put on top of your product. It has to allow for the engine under the hood to be unplugged and plugged back in with a more advanced system that's coming out every three to six months, basically, maybe even faster than that. It feels like every two weeks these days. So you've got to sort of take that into account. I don't think that's going to change. But I think thinking forward is about the whole web ecosystem and the way that apps work and things like that may change because of agents using these systems and being able to sort of use them as tools effectively.

27:36Yeah, that makes a ton of sense. The progress for Genie 3 has been incredible. And I think people are losing their mind over the system that we currently have. I am hopeful. I'm going to keep pushing you. of like, how do we get the model into more people's hands? So hopefully that's coming. Lots of people are excited. I'm sure you've been getting pings from people being like, how do I use the model? What, like, where do we go from here, from a world model genie perspective? Well, look, of course, we're trying to get it as efficient as possible now so we can give it to many thousands of people.

Read the full transcript

28:06It's in sort of, you know, limited preview currently. But we're also thinking about what's the best way of releasing this as a user experience. you know, can, we would love people to be able to share their creations with each other and allow people to play in the, you know, the creations of other people and things that got voted up, a sort of like a user generator community in a way. And, but the interesting thing is, is to maintain consistency because maybe you capture a lightning in a bottle at one point, you get some great prompt that creates a really compelling world. How do make sure that we can regenerate that world for the next players to come in and experience.

28:45There's a bunch of thinking around there. There'll be a lot more information on that to come soon. I think where it's going all together is if you think about Genie, you think about VO, you think about Gemini, I think a lot of these are separate models currently, but we're starting to see convergence of those models together into what we call an omni model, which can do everything. And we think that that's what an AGI system should be able to do, is really handle all of those different aspects to the same quality level that we see with all of these different specialized models, but perhaps in one big model.

29:25I love that. We were joking off camera about all the chess stuff and it just being a good excuse for us to go to play games. I feel like Genie's a good excuse for us to have a chance to make games and play them and then DeepMind's a video game. Yeah. Well, you know, that's always my secret plan is maybe like post AGI, once that's done safely over the line, you know, go back with these tools and make the greatest game ever. That would be a real dream come true. Is it going to be a roller coaster simulator? Yeah, maybe like the ultimate version of Theme Park, but I have some bigger game ideas in mind.

29:58We're building a bunch of vibe coding stuff in AI Studio. So I feel like the ideal world is pre-AGI, you can actually just start like firing off a bunch of these ideas and we have a whole like Demis game arena that you build yourself. Yeah, I would love to try that out. That's something on my very high up on my to-do list, that's for sure. I wanted to also, and you and I had a couple of tweets going out about this last week, I think, or the week before, we were celebrating like 980 trillion tokens on a monthly basis. And I think we crossed the quadrillion mark. So we had something special made for you.

30:34In your signature blue. Thanks so much. Whoa, fantastic. Awesome. We'll get some derivatives of this as well. Thanks very much. Of course. Thank you. This is a ton of fun, Demis. Thank you for taking the time. I appreciate all the late nights thinking about the future and all the hard work that you and the rest of the DeepMind team are putting in. So it's been a ton of fun. Great. Well, great chatting. Likewise. And thanks for watching Release Notes, everyone. And we'll see you in the next episode.

31:08We really Потом work with a new franchise. That's 붉 currently ongoing. Hey Tyrim?

From the publisher

Demis Hassabis, CEO of Google DeepMind, sits down with host Logan Kilpatrick. In this episode, learn about the evolution from game-playing AI to today's thinking models, how projects like Genie 3 are building world models to help AI understand reality and why new testing grounds like Kaggle’s Game Arena are needed to evaluate progress on the path to AGI.

Watch on YouTube: https://www.youtube.com/watch?v=njDochQ2zHs

Chapters:
00:00 - Intro
01:16 - Recent GDM momentum
02:07 - Deep Think and agent systems
04:11 - Jagged intelligence
07:02 - Genie 3 and world models
10:21 - Future applications of Genie 3
13:01 - The need for better benchmarks and Kaggle Game Arena
19:03 - Evals beyond games
21:47 - Tool use for expanding AI capabilities
24:52 - Shift from models to systems
27:38 - Roadmap for Genie 3 and the omni model
29:25 - The quadrillion token club

 

More from Google AI: Release Notes

All 30 episodes
Demis Hassabis on shipping momentum, better evals and world modelsGoogle AI: Release Notes · 31 min
Listen in VO