DeepMind Is Simulating Entire Worlds - Ready for AI Robots

5 Jun 2026 · 28 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Google DeepMind’s “world models” that simulate how environments evolve given actions, aiming to train and evaluate embodied AI/robots for rare real-world scenarios via simulation rather than only real-world exposure.

Guest backgrounds

Jack Parker Holder, leads the World Models team at Google DeepMind; started the Genie project (2022). DeepMind is known for AlphaGo, AlphaFold, and Gemini.

Key claims

World models predict the next state (e.g., video/audio/scene) conditioned on actions, learning physics and interactions from data. Genie builds a general simulator from broad internet video, not hand-coded games. For robotics, simulation is essential for humanoids and other agents to handle “long-tail” events safely. World models also help evaluate safety and robustness before deployment.

Notable examples

Genie 3 demo with text+image input and real-time interaction; Waymo using Genie 3 plus extra training to simulate rare events like snow on the Golden Gate Bridge, dusty desert roads, and an elephant crossing.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding World Models

0:01 to 0:21

Jack explains the concept of world models in AI and their significance.

“With Red Bull Summer All Day Play, you choose a playlist that fits your summer vibe the best.”

Understanding World Models

1:36 to 2:14

Jack explains the concept of world models in AI and their significance.

“So we're going to talk all about world models.”

The Role of World Models in AI Training

2:15 to 3:20

Discussion on how world models assist AI agents in training and learning.

“And then how the scene evolves as you walk forward, the model is kind of predicting this and producing, is generating this as you go.”

Challenges of Simulating Real-World Complexity

3:21 to 5:48

Exploration of the complexities involved in accurately simulating real-world environments.

“The question is, how are we going to get an environment or a simulator for that task?”

Human-Like Learning in AI

5:49 to 7:21

Jack discusses how world models aim to mimic human cognitive processes.

“will eventually get good enough and close enough that you can interact with something like the real world with this model.”

Genie Project Overview

7:22 to 8:29

Introduction to the Genie project and its goals in world modeling.

Advancements in the Genie Project

8:30 to 10:41

Detailed explanation of the advancements made in the Genie project over time.

“And our team said, well, let's just build a model of environments.”

Interactivity in Genie 3

10:42 to 13:12

Jack explains how users can interact with the Genie 3 model and its capabilities.

“And then once Genie 1 happened, which was published in 2024, we then just had these two jumps where we pushed scale both times.”

The Future of Humanoid Robots

13:13 to 14:00

Discussion on the future implications of humanoid robots and their training needs.

“wonderful that you can do this and impressive i'm wondering what the end game is here like why are you doing this?”

Simulation Scenarios for AI Robots

14:00 to 16:39

Explore the importance of simulating real-world scenarios for AI training.

“Maybe it's sunny in London and they've never seen that before.”
Show all 14 chapters

World Models and General Intelligence

17:19 to 19:51

Discuss how world models contribute to advancements in AI and AGI.

“I suppose one way to sort of sum this up is that, like, at the moment we have these large language models and they're very good at predicting the next word.”

Understanding Language in AI Models

19:51 to 22:36

Delve into the complexities of language understanding in AI models.

“So there's a lot of confusion, isn't there, over to what extent, like, you know, if you're talking to something like Gemini, to what extent does it really understand language?”

Potential Benefits of AI and Simulation

22:36 to 25:07

Examine how AI and simulations can enhance training and learning.

“Exactly, the model has just seen lots of examples, carefully curated examples and some kind of annotations and then from that it's able to extract this kind of intuition of concepts.”

Scientific Applications of AI Simulations

25:07 to 26:30

Discover how AI simulations can be a powerful tool for scientific research.

“I'm wondering if there's a science and research angle for these world models, because presumably you could build incredible...”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Jack Parker Holder:Ready to soundtrack your summer? With Red Bull Summer All Day Play, you choose a playlist that fits your summer vibe the best. Are you a festival fanatic, a deep end DJ, a road dog, or a trail mixer? Just add a song to your chosen playlist and put your summer on track. Red Bull Summer All Day Play. Red Bull gives you wings. Visit redbull.com slash bright summer ahead to learn more. See you this summer.

0:26Joshua Howgego:Think a few years down the line and you actually have humanoid robots in the real world. Think about all of the many different scenarios they may come across. Like maybe a cat jumps in front of them when they're walking down the street and they just don't know how to react, right? There's all these different scenarios that they might come across. Do you want them to see it for the first time in the real world? Or do you want them to have experienced many, many things in simulation before they're being deployed? I personally think the latter is the only way it's going to make sense.

0:50Jack Parker Holder:Welcome to the world, the universe and us from New Scientist. I'm Josh Houdjago and I'm your guest host for this special edition coming to you from South by Southwest London. And I'm joined today by Jack Parker Holder from Google DeepMind. Google DeepMind, of course, one of the world's leading AI companies and it's responsible for AlphaGo, AlphaFold, which people might have heard of. and of course, Gemini, which is Google's kind of flagship large language model. But Jack leads the world models team at Google DeepMind, and he started the Genie Project back in 2022. So Jack, welcome to the podcast.

1:34Jack Parker Holder:Great to have you.

1:35Joshua Howgego:Thank you for having me. Super excited to be here today.

1:37Jack Parker Holder:Yeah, fantastic. So we're going to talk all about world models. So I guess let's just start with the basics. like people might have heard of large language models, right? If they know anything about AI. So this kind of the basis of, you know, like chat GPT, Gemini. What is a world model?

1:56Joshua Howgego:Yeah, I mean, it's a great question that I think at this point we're almost asking ourselves too, because it's been a very hotly debated topic in the last few years. I'll go with my own definition, which is it's a system or a model that can simulate the dynamics of an environment, condition on actions. So imagine that you have some kind of scene and then you say, I don't know, walk forward is an action. And then how the scene evolves as you walk forward, the model is kind of predicting this and producing, is generating this as you go. Now, why would you want to do this? Well, you mentioned AlphaGo, for example.

2:30Joshua Howgego:So AlphaGo was a superhuman level agent that was trained to play Go by playing with itself using a thing method called self-play and reinforcement learning for many many many uh examples now it was able to do this because we have we know the rules of go right so we could give the algorithm the the rules of go and it can simulate go and then it can play it perfectly many many many times and get better and better and learn all these different things um so then after after alpha go one of the next big landmark results from this technology was alpha star which was where we then got agents at deep mind uh learned how to play Starcraft as well.

3:05Jack Parker Holder:And they did this sort of space exploration. Yeah, exactly.

3:08Joshua Howgego:It's a real time gesture game. I'm not probably the main expert to talk about that project, but the idea being it's another extremely complex game that you can learn by playing with it many, many, many times. Right. But as you then start thinking, how are we going to get generalist agents that can really be in the real world? The question is, how are we going to get an environment or a simulator for that task? Because we don't have a simulator of the real world. and so that's where world models come in because essentially if you have a model that can simulate an environment right it can therefore be used by by ai agents to then use it as a simulation to

3:44Jack Parker Holder:then learn new things so help me understand though because if you were being a bit skeptical about this you might say take something like like my son is really into like minecraft or like computer games like we can already like sort of simulate virtual worlds can't we but i assume that a world model is quite a bit beyond just that but help me understand how how exactly yeah so i think the

4:08Joshua Howgego:world world model is a quite an abstract term the way i would use it right so really it means it's a model that simulates an environment so there have been examples where people have actually built world models even from minecraft right in that case it's not doing the thing i'm mostly focus on, which is generating new things. But I think the really exciting example is, well, we have Minecraft, which is a way to build simulation of like more complex worlds, but there's a huge gulf between Minecraft and the real world. Like even subtleties, like, you know, the expressions people have on their faces or the way water, you know, collides with something or like the way trees blow in the wind.

4:46Joshua Howgego:These are extremely complex things. And I think it's quite hard to encode this in some sort of video game, right? Many people have been trying to make video games more realistic, but I think that's a very challenging problem. Instead, we have this amazing progress in generative AI, right? So you mentioned language models, but a few years ago, we started getting really good image generation models. And then shortly after that, we started getting really good video generation models. And these models can actually sort of simulate physics in the real world. Like you drop something heavy and it falls fast, you drop a feather and it kind of floats down.

5:17Joshua Howgego:They had this kind of intuitive understanding of physics. Yes. And so then if you can take one step further and then generate a new world, then they also kind of inherit this capability of being able to actually simulate physics as if it's from the real world, which would be extremely hard to do by building it by hand on top of Minecraft. after I think that's my bet. So it's a bet right and it's still not finished and it may be the case that someone can build with hat by hand with code a much better version of the world than we can but the bet we're making is that learning from data and generating things based on real data will eventually get good enough and close enough that you can interact with something like the real world with this model.

5:58Jack Parker Holder:So is one, tell me if this is a helpful way to think about it, like is there a sense in which we have in our own minds a model of the world yeah and it's based on learning it's based on data and then we have this kind of sense of you know if i throw a ball i can roughly i can throw it so that it will land at a particular spot and that is in a sense what you're trying to build it's not hard coding a video game it's trying to build something which will learn and create a model of a world in that in a similar way to we create one

6:28Joshua Howgego:in our minds is that right yeah so that that's actually a really big question to unpack so so i'll start in a few different ways so firstly the way you can think about this is that language models inherited a lot of the generality through just predicting the next token in a sentence right yeah and so just by doing that right from massive scale training by predicting the next token which roughly is roughly is a word but it's not exactly a word you can learn all of the structure of language without ever telling it about, you know, adjectives or like nouns or whatever, or entities, all these little problems we think about.

7:00Joshua Howgego:Language model can just learn all this structure by predicting the next word. What we do with world models is we predict the next state of the world, which typically is something like maybe video, audio or something like that, but state of the world. And by doing that, the model then learns, you know, how physics works and all these kinds of things, how people interact. So that's the way we're training them. Now to go back to humans, there is a lot of evidence that humans do have world models right uh and they and they are able to like plan and reason in an abstract sense but that brings you to the to the i guess hot debate of what level of abstraction should world models have should they be high fidelity in modeling like right now uh i can in the corner of my eye someone can see over there someone over there is like you know on their phone do i really care about them being on their phone over there when i'm talking to you right now and my task is to deliver an exciting podcast not to worry about what they're doing over there and so maybe humans with their world models actually have a much higher level view right so yeah examples are like if i threw this water on you right now you'd probably be annoyed i don't know exactly what you do i can't model i can't model exactly what you do freak out i can't model exactly what you do but i know you wouldn't like it right so that's what my world needs to know right it doesn't need to know like the fine details you know so what we've definitely done with the genie series has been a bit more in between so we have a model that does actually reproduce like pretty impressive visuals but that doesn't mean that it's like prioritizing that above all else it's actually got like quite a um it's got a reduced representation but under the hood right yeah um whereas other folks are pushing the other way and they're really thinking that world models should be very abstract for for more like more the way humans maybe have it in their

8:39Jack Parker Holder:minds okay yeah okay i'm with you so let's come on to to genie itself yes which um you started in 2022 a few years ago and I think it was there was a version of it released that people can go and

8:51Joshua Howgego:play with online last year I think in January this year yeah so tell me about what is Genie how does

8:57Jack Parker Holder:it work I think it takes in lots and lots of data from all sorts of different sources but you tell

9:03Joshua Howgego:me yeah so Genie was project started in 2022 roughly at this time when we were running out of environments to use reinforcement learning and we were thinking what's the next environment we should use. And our team said, well, let's just build a model of environments. And then once we solve that, we're done. And we can then go back to the reinforcement landing people and they can do amazing stuff with it. So that was the inception of the project. In the first year, we focused on this approach that was sort of saying, what if I have videos, but no action labels? So I'm trying to learn a model that essentially takes in an observation of the world plus an action and then predicts the next observation of the world, right?

9:39Joshua Howgego:That's basically the thing we're trying to model but we don't have actions for internet data because you know you go on the internet you see a video it doesn't have action labels um and so what we did we tried to to find a way where we could like extract them and that was a novel thing that we did and then subsequently trained a model that could accept actions but for arbitrary videos so you could give it a completely new thing it's that's not in its training data and you can still control it with actions because it's learned this like general idea of what actions do so you're talking

10:08Jack Parker Holder:about a video of say a person doing some action and you want to label what that action is exactly

10:13Joshua Howgego:yes but a way of like learning how yeah it's a bit more bit technical but we're learning how to learning how to extract labels for in just a general situation yeah every time someone walks left it should have the same label roughly and if you see enough of videos like of that then you kind of learn this and then the idea with genie was instead of having a world model for just minecraft or just one atari game or just go you have a world model trained from a broad data set That means you can then generate new worlds and new environments. And that was the magic of it. And then once Genie 1 happened, which was published in 2024, we then just had these two jumps where we pushed scale both times.

10:51Joshua Howgego:So Genie 2 was a direct follow-on for Genie 1. And then Genie 3 was a bit different because for Genie 3, we realized that we needed some additional innovations to get to the quality of video. And so we worked very closely with, in GDM, we have a leading video model called VO. So at the time, VO2 was the best video model in the industry. And we worked with some folks from that team and leveraged a lot of their expertise to basically make a big jump in quality as fast as we could.

11:18Jack Parker Holder:So let me make sure I understand this. So basically what's happening is Genie is creating worlds. And you give it some parameters, do you? Or it just creates a world?

11:29Joshua Howgego:Yeah. So the latest version of Genie 3, it has text and an image input. So you could give it an image of this room, and then you could describe in text exactly what's happening in this room, and then the model has both. And then from then on, you just interact with it in real time. That's the difference with Genie 3 is not only is it higher quality visuals, but actually, firstly, it's really consistent. So if you walked around and came back to the same location, it should be pretty much what you saw.

11:55Jack Parker Holder:And who is like, when you say if you walked around, is there like an AI agent that's doing that or how does that work? Is that...

12:04Joshua Howgego:Yeah, that's a great question. So it's a model that just takes input in the form of actions, but that could be you playing with it or it could be an AI agent. So we've actually got both. The original motivation for the project was that once we have the ability to create these worlds, we could then put our AI agents in them and they could learn new things. But by building that and actually by testing it yourself to see if it's working, you realize it's quite fun for humans to play with it and it's kind of a new sort of way of experiencing AI and it's a new way of creating things and so actually we found that humans do enjoy interacting with these models their relationship with them is still yet to be defined right because it's not like it's replacing an existing form of media it's a it's a cliche but it's a new thing um yeah but we have um to jump around again we did release in January this year um a project genie which is basically giving folks the ability to try it out right it's not a fully fledged product but it's like kind of a demo of what these models can do and you can go online yourself with a google ultra account and you can type add an upload an image or you know create an image type in some text and then you just for a minute do whatever you want in this world yeah and i guess it's obviously quite

13:17Jack Parker Holder:wonderful that you can do this and impressive i'm wondering what the end game is here like why are you doing this? Is this about what specific kind of goals do you have in mind for this?

13:28Joshua Howgego:Yeah, that's a great question. So for us, it's always been about this agent's angle. And if you think about it, then in robotics, we've obviously been making really rapid progress in the last few years, right? And so now we've got humanoid robots, right, that can do many physical tasks, and they're starting to be deployed in the real world, right? Now, think a few years down the line, and you actually have humanoid robots in the real world. Think about all of the many different scenarios they may come across. Like maybe a cat jumps in front of them when they're walking down the street, right? Maybe it's sunny in London and they've never seen that before.

14:03Joshua Howgego:Maybe it's like a certain kind of person who's wearing a certain outfit that they've just never seen before and they just don't know how to react, right? There's all these different scenarios that they might come across. Do you want them to see it for the first time in the real world or do you want them to have experienced many, many things in simulation before they're being deployed? I personally think the latter is the only way it's going to make sense and then if you believe that they need to have simulation for these events then how are you going to build the simulator and that goes back to the previous point are you going to show them Minecraft?

14:31Joshua Howgego:I don't think so and that's what you're doing I think you want to build a model that can simulate anything because then you can use it to train and evaluate these kind of humanoid robots that might come into the real world or other robots too but humanoid I think is a very easy to understand modality but there's many other kinds of robots that might, or even virtual agents that might come into the embodied world, like either through as an assistant or something, right? And how are they going to simulate these scenarios? Now, to be a bit more concrete, one of the embodied AI agents that we do have in the real world is Waymo, right?

15:06Joshua Howgego:Self-driving. Exactly. So Waymo is probably the world's most advanced self-driving company. I've used it myself many times in the Bay Area, and it's pretty magical. and now they're testing the cars on the streets in London. So that's really exciting. And they are autonomous cars driving around the road, right? Rain or shine. And there's many things that they have seen, but there's also some things that they can't have possibly seen so far. And so we partnered with Waymo in the second half of last year and they released a blog post in February. They're actually using Genie 3 with some additional training on top to make it know their setting to simulate these long tail rare events.

15:42Joshua Howgego:and we can see in the blog post that they're able to simulate things like the Golden Gate Bridge covered in snow or being on a kind of dusty desert road and an elephant walks in front of the car, right? And these probably aren't in their training data, right? But they could feasibly happen and they want to be able to show the car those situations. And now they can, right? Because they've used this generative world model that can produce these simulations.

16:06Jack Parker Holder:Yeah.

16:07Joshua Howgego:Most of us worry about memory lapses as we age.

16:10Jack Parker Holder:But how do you know when it's something more serious? You can explore this question in our latest New Scientist CoLab feature, where Dr Tim Beanland separates dementia myths from scientific facts. He discusses the breakthrough drugs showing promise in slowing Alzheimer's, the potential for simple, finger-prick diagnostic tests, and the incredible finding that nearly half the dementia cases globally could be prevented through lifestyle changes. It's a roadmap for brain health, and you can't afford to miss it. Search Dementia Everything You Need to Know at newscientist.com or visit alzheimers.org.uk.

16:45Jack Parker Holder:This message was sponsored by the Alzheimer's Society. Study and play. Come together on a Windows 11 PC. And for a limited time, college students get the best of both worlds. Get the Unreal College Deal.

17:01Joshua Howgego:Everything you need to study and play with select Windows 11 PCs. Eligible students get a year of Microsoft 365 Premium and a year of Xbox Game Pass Ultimate with a custom color Xbox wireless controller. Learn more at windows.com slash student offer. While supplies last, ends June 30th, terms at aka.ms slash college PC.

Read the full transcript

17:19Jack Parker Holder:I suppose one way to sort of sum this up is that, like, at the moment we have these large language models and they're very good at predicting the next word. And so they're very good at simulating being able to talk to like. Exactly, yeah. The next, what you're describing here is something which isn't just a model of language, but has a spatial dimension to it. It can model these different worlds. And you've talked about the different applications for that. I'm sort of wondering about the further future, moving towards something that's more like a general intelligence. It sounds like what you're saying is that this is going to be a necessary step, having a world model to get us towards more of an AGI.

17:59Jack Parker Holder:Is that right? And will that be, in itself, will that be enough? Or will we need more on top of that, do you think? I know that's quite a big question to unpack.

18:07Joshua Howgego:Yeah, it's a great question. And I have maybe a slightly hedged answer, but I think a genuinely honest one. I think that you could arguably get to something that many folks would refer to as AGI without this, because if you consider just, for example, some of the things Gemini can do now, you know, really advanced math, coding, have a strong conversation with you in many different languages. It's not my area of expertise, but I think this is inching towards something that I personally would have thought was probably fair to call AGI a few years ago. But G in AGI is doing a lot of work here, right?

18:41Joshua Howgego:Because how general are you going? I personally have been more interested in embodied AGI, which is also kind of loosely defined as a brain that can sort of operate and understand and reason in many sort of situated real environment, right? So it can understand social situations, it can understand physics, it can understand real world situations, it can work well with others in real world, it can understand the many complexities of modern life in the physical sense. And I think you can't solve that without world models. I think you may get something many people refer to as AGI that can't do those things.

19:17Joshua Howgego:But to do those things, I just don't see how it's physically possible that we would be able to achieve an agent without simulation. and I don't see how we're ever going to simulate the complexity of the real world other than learning it from data. So for me, I think it is absolutely the way to get there. And I think that this could have a profound impact on society. But it's obviously a bit longer, it's a bit further out, right? Because most of the progress that's really been deployed in AI has been in this more like, I guess, language-based agents fight. Apart from Waymo, which is actually deployed on the streets.

19:50Jack Parker Holder:Okay, I've got another difficult one for you. Go for it. So there's a lot of confusion, isn't there, over to what extent, like, you know, if you're talking to something like Gemini, to what extent does it really understand language? Or to what extent is it, you know, some people will just say all it's doing is just predicting the next word. It doesn't understand anything. It's just simulating understanding. I'm interested in what you think about that. But I'm also interested in whether or not that debate changes when you have a world model. To what extent can we say that these models have understanding or is it just kind of it appears like it's understanding?

20:27Joshua Howgego:Probably not the one who has the best expertise on the language models. I mean, they do have, you know, you can see thinking traces, for instance, which I personally find super useful. And you can see they have access to tools already.

20:38Jack Parker Holder:And thinking traces means the AI, the chatbot will sort of tell you what its thought process is. Exactly.

20:44Joshua Howgego:And I think sometimes it does feel like what they're doing is intuitive, but I'm not close to the language model. So my view on this is just the same as anybody else's. I have no special expertise there. But for Genie 3, I do think that it has some understanding of concepts. because for instance we have this example where you can prompt I told you you can prompt the model with an image and and text but you can actually also prompt it with a video and text because I mean a video is just the first few frames it's no different you can just place it into the model's context and what happened in the team was someone found by mistake that they could put the wrong prompt in with the wrong video and so they actually prompted the model with a video of like being in some kind of modern city and then text saying you're skiing on like in some in a mountain somewhere and what the model does is it starts off looking at the video but then when you turn to the side you're actually in a cave in a mountain and so it's blended these two prompts together in a very sensible way that I think I think shows a level of generalization and actually understanding of the difference between them and how they could fit together that feels like it feels somewhat intelligent right in terms of the way it's done it and and And then we actually went one step further.

21:56Joshua Howgego:So we have this pretty remarkable example where we have a video of people on the scene playing with Genie on a laptop, right? Because we wanted to show this is a real model, you know, they could have been playing with it real time. And we prompted the model with that video of people playing the model, if you're still following. It's a bit meta. Yes. And then the text was, you know, a jungle with dinosaurs. And then when you start interacting with Genie, what it does is actually it changes the screen of the game that they're or the world that they're playing to a jungle with dinosaurs and then when you turn around you realize they're actually playing it in a little cabin in a jungle and so the model has blended these concepts like and it explodes a little bit yeah i mean you have to i think you have to probably have visual aid for this one but like the idea is that the model even understood that it could simulate things within a simulation essentially yeah i think that

22:46Jack Parker Holder:is that was pretty surprising to me yeah it's sort of an emergent behavior you might say this not been like hard-coded into the model.

22:52Joshua Howgego:Exactly, the model has just seen lots of examples, carefully curated examples and some kind of annotations and then from that it's able to extract this kind of intuition of concepts.

23:05Jack Parker Holder:You've talked about these world models being used to train robots, for example, like humanoid robots and obviously just that could for some people that could be a scary thought, right? There's all sorts of dystopian science fiction about that kind of thing and also just generally AI there's so much concern about it. I mean what can you say to reassure us that all this research is being done for good?

23:29Joshua Howgego:Yeah I mean that's obviously a very poignant question for many of us in the field. Firstly I do believe at GTM in general we follow a very responsible approach to doing things and I'm very proud to work somewhere that has that mindset. I also would go a step further and say I wouldn't feel super comfortable with robots being in the real world if they didn't have this kind of thing because what you can do with Gini is actually evaluate them much more effectively in long tail scenarios that you couldn't possibly expect it to have seen before and I personally really care about this being done so actually this is one of the main motivations for us is if you didn't have the ability to evaluate a huge wide variety of complex but feasible real world scenarios like what happens if it's Halloween and children are wearing costumes and the report's never seen this before.

24:17Joshua Howgego:I'd want to at least simulate that, you know, a thousand times or something, made that number up, and see what it does, right? And make sure it's safe, make sure it works well in that scenario. So I feel like having the ability to evaluate our agents and also train them on the things that they got wrong and, you know, improve them is actually a critical step to make this safe. I think the research is progressing in general pretty fast, but I think this will likely make it so that it can be done more safely. And I'll add as a final point that this not only is useful for training agents, but also humans can actually interact and learn with these worlds too.

24:52Joshua Howgego:And I think in those cases, it's largely beneficial in that it can be used to try new ideas, but also for people to learn new things pretty quickly. So you could experience something that you couldn't possibly experience at your fingertips from a laptop.

25:07Jack Parker Holder:I'm wondering if there's a science and research angle for these world models, because presumably you could build incredible... Like a simulation is a tool that scientists use all the time. We want to simulate weather patterns or the inside of materials at a quantum level. Could we be looking at a future where we use these models to simulate those kind of environments and then set an agent to go and try and understand them? And so could it be a really useful scientific tool? Is that part of what we're...

25:39Joshua Howgego:Yeah, I think so. I think to get, you know, really atomic accuracy, I think it's probably, you know, quite tricky to get to that point. But we always use this example in research a few years ago. We use this example of deploying rovers on Mars, right? So imagine we can't really test things quickly on Mars. Imagine we collected some data and then we could build a world model for that scenario, right? You could then simulate it probably more accurately than you can just guess with the model. And so we probably could start testing things. I mean, in dangerous scenarios too, like in a hurricane or near a volcano erupting, if you wanted to test things or deploy robots there, where I think unequivocally, we don't really want humans to be deployed there.

26:22Joshua Howgego:I think you can test these things in a pretty accurate simulation. That would be quite a beneficial thing to be able to do.

26:30Jack Parker Holder:Thank you, Jack. We're going to leave things there. Thanks so much to our guest, Jack Parker Holder from Google DeepMind. And thanks to you for listening. I'm Josh Howdjago, and do follow us wherever you get your podcasts.

27:02Jack Parker Holder:the sun's harsh rays that can burn and damage your skin. The sun is relentless, but so is our gear. Level up your summer at Columbia.com to spend more time outside and less time slathering on aloe lotion. You're welcome. Columbia. Engineered for whatever.

From the publisher

Episode 374

Google DeepMind is simulating entire worlds using AI - that can be interacted with in real time.

“World models” simulate the environment and physics of the real world. And DeepMind’s Genie 3 model allows people to create these worlds with basic image and text prompts.

The idea is not just to allow people to explore these worlds, but to serve as a testbed for AI agents to learn how to interact with the world before they are deployed in humanoid robotic bodies. 

Could this be the next big step towards artificial general intelligence (AGI)?

Joshua Howgego speaks to Jack Parker Holder, Research Director at Google DeepMind, about the latest developments.

To read more about these stories, visit https://www.newscientist.com/
Learn more about your ad choices. Visit megaphone.fm/adchoices

More from The World, the Universe and Us

All 94 episodes
DeepMind Is Simulating Entire Worlds - Ready for AI RobotsThe World, the Universe and Us · 28 min
Listen in VO