Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building

16 Aug 2025 · 41 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

a16z Podcast Episode Notes: Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building

Episode Overview

  • Title: Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building
  • Description: This episode features discussions with Jack Parker-Holder and Shlomi Fruchter from Google DeepMind about Genie 3, an AI model that generates interactive, persistent worlds in real time from text inputs. The conversation explores its innovative features, applications in various fields, and future developments.

Key Participants

  • Google DeepMind:
  • Jack Parker-Holder (Research Scientist)
  • Shlomi Fruchter (Research Director)
  • a16z Team:
  • Erik Torenberg (Host)
  • Anjney Midha
  • Marco Mascorro
  • Justine Moore

Major Topics Discussed

Introduction to Genie 3

  • Overview of Genie 3:
  • Capable of creating fully interactive environments from minimal text.
  • Significant advancements in real-time generation of worlds.
  • The ability to create environments that appear realistic to users.

Special Features of Genie 3

  • Special Memory Function:
  • A breakthrough enabling consistency across generated environments.
  • Allows for interactions where previously created elements remain intact.
  • Emergent Behaviors:
  • Models show unexpected capabilities as they scale, leading to more complex interactions.

Development Journey

  • Evolution from Genie 1 to Genie 3:
  • Progression involved integrating different projects and insights from previous models.
  • Collaboration between teams at DeepMind was crucial for success.

Applications and Use Cases

  • Potential Applications:
  • Gaming: Interactive and immersive environments that can host user-generated content.
  • Robotics: Enhanced training environments for AI agents to learn from simulated experiences.
  • Education: Creating realistic scenarios for training and learning.
  • User Engagement:
  • The design encourages developers to explore various uses, making it adaptable for future innovations.

Comparing Models

  • Instruction Following and Text Adherence:
  • Genie 3's advancements in understanding text prompts compared to previous models.
  • Comparison with Other Generative Models:
  • Highlighted differences between Genie 3 and image/video generation models like VO2.

Future Directions

  • Next Steps for Genie Models:
  • Discussion of potential future iterations (Genie 4, Genie 5).
  • Focus on expanding capabilities and addressing current limitations.
  • Integration of AI in Robotics:
  • Use of Genie 3 to bridge the sim-to-real gap in robotics, allowing for more realistic training environments.

Key Takeaways

  • Realism in AI: The ability of Genie 3 to create realistic environments marks a significant leap in generative AI technology.
  • Interactive Experiences: The focus on user interaction can lead to novel applications across various industries.
  • Collaborative Innovation: The collaborative nature of DeepMind’s teams contributed greatly to the breakthroughs achieved with Genie 3.
  • Future Potential: There’s a strong emphasis on the unexpected applications that users may discover as they experiment with Genie 3, indicating a shift towards more interactive and user-driven AI development.

Philosophical Reflections

  • The conversation delved into the implications of AI and simulation, pondering whether our reality could be a simulation, showcasing the philosophical depth often intertwined with advancements in technology.

Conclusion The episode provides an in-depth look at Genie 3's capabilities, the research behind it, and its potential impact on various fields, emphasizing the importance of generative AI in shaping the future of interactive experiences.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00All of the applications basically stem from the ability to generate a world of that just from the few words. You look at it and like there's a world that's generated in front of you eyes and it's amazing that it's happening. I was very excited about how far can we push that. And it's at the point where like a human who is not an expert will watch it and think it looks real. Right? And I think that's pretty incredible. GD3 from Google Debtline can create fully interactive persistent worlds in real time from just a few words. Today, we're joined by the team behind it. Shlomi Frickter and Jack Parker Holder from Google DeepMind plus Andean Midha, Marco Mascoro and Justin Moore from A16Z.

0:40We'll talk about how it works, the special memory that keeps worlds consistent, the surprising behaviors of learned, and where world models are headed next. Let's get into it. Jack Schlomi, GD3 is taken over the internet. We're honored to have you on the podcast today. As a response surprised you, reflect a little bit about the reaction. We weren't sure how big it's going to be, but today felt to definitely that way of something that was for a long time coming, basically being able to generate environments in real time. I think a lot of work that was done in Google's DeepMind and outside pointed to that direction, but we really wanted to make it happen and I hope we have.

1:21Team, one we reflect internally a little bit about what we found so game -changing about Gen3 and where we so excited to have this conversation. Yeah, for sure. First of all, it's an amazing model. I think there's a lot of excitement around the special memory, the consistency across all the frames. I think this is the first time I can see like you can have some sort of interactive way of doing this stuff with videos because it used to be like, you would do one problem and you would have 15 seconds of a video. But now you can actually have some sort of interactive kind of element to it, which I think is very exciting.

1:49So can you elaborate a little bit more like you insights on these like how was life, for example? Fear now what data you should collect, how you make it very interactive and keeping the flow of the whole video, which I thought was phenomenal. Sure. Yeah. So I think you kind of highlighted a few capabilities, sort of the length of the generation, the consistency of the world. Maybe diversity as well of the time kind of things you can generate. I think the main thing is that obviously we made progress in quite a few different fronts, right, in separate efforts, right? So we had this Gini2 project that was much more sort of like three environments that it could generate.

2:27And it wasn't super high quality. It felt like it came from Genie 1, but it wasn't the same quality as things like VO2, which, to say to the R video, a model at the time came on December, roughly exactly the same time it came on a week later. And Gini2, and obviously internally, there was a lot of discussion between the two projects about the different directions we're pursuing. And then, Jeremy had also worked on Game Engine, right, which is the Doom paper, as people know it, which I think you guys also wrote a nice piece on straight after that came out. So I think that also attracted a lot of attention.

2:58And so we felt that across these different projects, we had quite a lot of interesting things that would naturally kind of combine. And we could basically take the most ambitious version of the combined project and see if it was possible. And unfortunately it was, and quite, I think the timeline is probably the bit that surprised many of us. Because obviously we sell ourselves these goals and we tried very hard to achieve them, but you're going to be totally sure how it's going to actually feel when you've got that point. I think it ended up being something that resonated with people a lot more than maybe we expected, but we always believe us.

3:34Yeah, I was just said to this that I think there is time, so a component is really important. I'm not many people experience it firsthand, but really try it in their release to at least have a few trusted thirsters Interact with it and also get the feel of it by adding these overlays that show what happens How people can like use the keyboard to control it and I think there is something magical about the real -time aspect I felt it's for the first time when our model like game engine model started working fast enough and we were just like Oh my god, it's actually I can actually walk around and it was a bit of an a while moment And yeah, I think there is something when it responds immediately that is really magical.

4:14I think that's kind of sparked the imagination of many people when the dome simulation came out and here we really wanted to push it somewhere, we weren't sure it's going to work. So it was definitely the edge of what's possible. I think that's how we felt. So we just said, yeah, let's try and see if we can make it happen. I think you guys, I don't know if this was on purpose or not, but you perfectly timed it. when everyone on X and Reddit and everywhere was making those videos of characters walking through games, but they obviously weren't interactive, they weren't real time, and then you guys came out with this release that was like, now this is an actual product, and it blew folks away.

4:50I'm curious, because you can imagine so many different applications for this, right? Like more controllable video generation, or making it much easier to create games, even personal gaming, where someone's just kind of creating their own world, they walk through, like, RL environments for agents, robotics. Are there any particular use cases that you're most excited about? I think all of the applications basically stand from the ability to generate a world just from a few words. And I think, for me, this potential, when I started looking at video models, I think it was pretty early when I think one of the models were like, imagine video, which was modeled by Google research.

5:28But there are a lot of models that they were very basic compared to what you have today. but the ability to simulate something like you look at it and like there's a world that's generated in front of your eyes and it's amazing that it's happening. And I think at this point I was very excited about how far can we push that. So I think there was one way to do it and you need definitely another way to make it a bit more interactive. So I think all of the deep applications basically stem from this core capabilities. So it can be entertainment, of course, as you said, it can be train -ing agents, it can be helping agents to reason about the world.

6:02education, so I don't think any particular application is more important or than others, I think it's really up to how developers in the future will be the long -term stuff. Yeah, I would get basically the same answer in the end with a different journey to get that, right, which is I've been passing myself work to reinforcement learning for a few years before starting the GE project in 2022. The motivation I literally was, like, that in RL at the time, we had this problem where I would say which environment should we try and solve, right? Because once you've already done go, wish you both thought was years or decades away, and then that was solved in 2016, more solved.

6:38But we should be human 11, 2016. And then Starcraft, three years later, which is not particularly long time for something incrementally significant. So it was 2021 time. It was a big question of what should we try and do with RL? We know that the algorithms can learn superhuman capabilities if they have the right environment, but we don't know what they're in. And so we are working on designing our own ones, right, with Colored. But then instead it seemed like the more promising path when you had the first text image models coming, it coming out. Whereas like, what if we just think long term, what's the way to really unlimited environments?

7:11That being said, over the course of the project, and originally we started it, I guess in 2022, it was very focused on that one application, but it seems quite clear now that this could have a big impact on all those other areas you mentioned, right? So I think it's like language models in 2021 maybe you probably wouldn't have guessed like an IMO gold medal a few years later But come that fast as a direct application of that technology, right? It was probably it can help me with my emails or whatever it was and I think it's really cool to build these kind of new Class of foundation models and then see what people can imagine doing with it And that's one of the very exciting things about sharing the research preview right?

7:47You've got this kind of feedback So hoping a lot of these things can happen One of the things in the research preview post, Jack, that blew me away, was this... And it wasn't even your first GIF, I think, in the blog post. It was either second or third. You had this visual of somebody painting the wall with the paintbrush, and then the character moves... Yeah, this push -over. Right, like, out of... ...to a different part of the wall, paints, and then moves back. And the original paint is still there. And I didn't believe it. I was like, there's no way. And then I read, and you read, it was described as a special memory.

8:22So the persistence part for me, I'm not taking away from all the other stuff. The interactivity is amazing, but I think broadly speaking, folks expected that at some point, video generation, for example, would become real time. When I saw the Genie 3 post, I was like, okay, they actually went and did it. But the special memory, the persistence was when I sat up in my chair and I was like, how did that happen? Could you talk a little bit about when did you discover that as an emergent property, or was that a specific design goal? What's the backstory on that? because that feels like a big unlock Jack, why don't we start with you?

8:53Yeah, so that's a great question. I'll say a few things. So the TLDR is it was totally planned for, but still incredibly surprising when it worked that well. So that specific sample, when I saw it, it was hard to believe. I actually wasn't sure that the moral generated for a second. I was like, that took me to watch it a few times and really checked and freeze the frames and look back and check that it was the same. But go back to a few steps. So I'm obviously Genie 2 had some memory, right? So this got kind of lost because I mean Genie 2 came at a time When there were lots of announcements very rising announcements being VO2 only a few days later It was a busy time of the year and the main headline act was that we could generate new worlds at all Right so that was the thing that we wanted to emphasize But it did have a few seconds of memory and we had a couple of examples like they created a robot near a pyramid Looked away looked back in the pyramids there, but it's like kind of blurry It's not perfect, but some other models around the same time or more recently didn't have this feature, right?

9:52So people kind of indexed to that because they didn't notice the early signs of it in the G2 work. And then for Gne3, we basically, when much more ambitious on the same sort of approach, right? And we made it like a headline goal for ourselves is, can we make the memory be what it is? right? We said we want minute plus memory and real time and the higher resolution all in the same model. And those are kind of conflicting objectives, right? So we sell ourselves this kind of technical challenge and we said, if we target this, then it's just about feasible and it'll be pretty incredible. And then you still don't know obviously it's going to pan out.

10:30So then when you get to the end of the research, one seven months later, to see the samples, it still is quite mind to be honest. So yeah, it's kind of planned for, but still pretty cool and exciting when you see it because like, you know, they research projects on like sure things are they? So one thing that we didn't want to do and we didn't want to build an explicit representation, right? So there are definitely methods that are able to achieve consistency and they did that through and explicit some 3D, you know, it's their nerves, I'm sorry, I'm splattering, and other methods that But basically say, okay, if we know how the world looks like, we use this kind of like prior assumptions on how the word remains static pretty much, then we can build representation, then now what we're looking at.

11:17So that's great, I think, for some applications, but we didn't want to go down this path because we felt it's so much limiting. And I think so we can definitely say that the model doesn't do that. And that generates kind of frame by frame. And we think this was, this is really key for the generalization to actually work. every time someone interacts with it for the first time and they like test, they look away and then look back I'm always like holding my breath and then it looks back into the same and like whoa It's still really it's really cool. It's very cool. And how long is this special memory?

11:46I don't know if you can talk about it You mentioned a minute plus, but is there some sort of like Make sure that you have you see like can you keep it for half an hour or what is the limit on that? There is no like fundamentally mutation, but to currently the current design will emit it to one minute of this type of memories. Yeah, it's also a real -time tradeoff for the guests as well. We felt that because of the breadth and the other capabilities that like a minute were sufficient for this version, like it's quite as significant a leap, but obviously eventually you'd want to be serious. One more question related on the between Genie 1, 2, for example, in LEMs like you have Deep Cic are one, they saw on this paper, like the longer they keep it running, they suddenly will see like these interesting behaviors like the models are like reasoning or like would give like a, oh, I'm wrong, it is, I should self correct.

12:37Do you see anything in kind of like this scaling from two to three? Do you see any sort of like interesting behavior that you were not expecting that suddenly just appear by increasing the amount of data and the amount of compute? Yeah, I was just saying, I think there is a bit of Like overall, definitely like many generative models will see that improvements happen with scale. I think that's not secret. And I don't think it's not the same type of intelligence, I would say like an LLMA as I'm not really reasoning is the right term. But we do see that some definitely things like it can infer from view approach like a door.

13:13It's and it makes sense for the agents to maybe open it. So you might see that it's starting to do that, heart sample. Or there's some like a bit toward understanding that happens over time and it just like things look better and more realistic. So I think these are the trends that we've still observed. Yeah and from G2 to 3 it's I think the real world came with this really increased right. So on the physics side some of the water simulations you can see some of the lighting as well like a really breathtaking. I think we have this example of the storm on the blog and that one I I think is super cool.

13:47And it's at the point where a human who is not an expert will watch it and think it looks real, right? And I think that's pretty incredible. Whereas what you need to, it was like, it kind of understands roughly what these things should do. But you know it's not real, right? You can look at it and you can clearly see that it's sort of not very photorealistic. So I think that's quite a big leap on the quality in that side. Yeah, one of the things that was really cool in all the examples was the water is sort of a great way to see, like, does it understand, like, what the world is and how objects interact.

14:19And that example, someone posted the feet going in the puddle was amazing. But then there was also that example of, like, a cartoon character. It was more of an animated style who was, like, running across this kind of green patch of land. And then ran into this blue, kind of, wavy thing that looks like water. And he started swimming, which I thought was really interesting. Like, were there particular things you had to do around that for the model to be able to understand how characters should interact in different environments and different styles. What you're basically describing is the real breadth of different environment terrains and worlds and things like that, water or walking on sand versus going downhill and snow and how the agents' interactions should differ given that they're terrain that they're in.

15:07And I think that that really is a property of scale and breadth of training. This is very much like an emergent thing. I don't think there's anything really specific we do for this, right? You again, like you hope the model hasn't learned this because it should have like a general world knowledge. It doesn't always work perfectly, but in general it's pretty good. Like so for the skiing examples, you do go fast when you go downhill and then when you turn, and try and go back up here, it's very slow, if not at all possible. When you go into water, obviously, you hope, as you said, that the agent will start swimming and slashing, and this does typically happen.

15:45When you look down near a puddle, hopefully you're wearing Wellington boots. Like this kind of stuff does just kind of make sense. And I think it feels pretty magical, because it very much aligns with what you were thinking about the world, and the models just generated it all. So yeah, that's also one of the really exciting things, special. Yeah, and on all that I want one kind of trade of the typically we have is that we want the models to do two things. We want them all to create the world in a way that looks consistent. So, Jack said like if you if you walk in the rain or in piles, then probably wearing boots, but if we provide it with a different description or like the prompt is saying something else, we want it to still follow the prompt.

16:27And there is some tension here because some things are very unlikely, right? You might say I want to wear flip flops and jumping or anything. Then the model still has to try and create something that is very unlikely. And that's where typically video models may be find it more challenging. And that's where our models might find it more challenging, but still successful to a surprising degree to go into this kind of a global mobility area. And I think that's really, in a way, That's what we want. Many people just want to look at the video that looks like they're on this room. But something a bit more exciting.

17:09And that's where I think this is the magic of the models that they can take you to places. It's not so likely to be in reality. The text following is really amazing in this model. That does feel really magical. I think there's something that the VO does really well as well. right? Like pretty much what you asked for. It's really well aligned with text. And so, and we've had that with Genie 3. So you could describe very specific worlds and really kind of like arbitrary silly things. And it pretty much works. Like we actually had this discussion because people were very disappointed to find out that the video I made in my dog actually was not my dog's photograph, I just described her in text.

17:56And yeah, I don't know if that's a big secret, but it looks exactly like her. And the model just kind of knows, right? And I think that's pretty amazing. So I think that that's actually a really important capability that we didn't have with Gini 2 as well, right? Because we relied on image prompting. And so there was some transfer issue like where you rely on imagines of a generate image. And that often does look really good, but it's not necessarily a good image for starting the world. Whereas like going directly from text, you get the controllability from anything you want. Plus it just kind of naturally works because it's in the like correct space for the model to do its thing.

18:35And that's something really powerful. And why is that Jack? What do you think led to such a massive instruction following a Dexter Deurance gain? Because it's a pretty hard thing to do. Well, I mean, our team never really works on this. So Genie 1 and 2 both worked with image prompting. And so obviously like for this next phase, we leveraged a lot of the research done internally on other projects and personal ways. I mean, Xiaomi's obviously worked in co -leading the VO project. And so we were able to kind of build on a lot of other work and ideas internally. And that basically allowed us to kind of like turbo charge progress, right?

19:15So if we've done this sort of by incrementally building like ourselves in in on an island, it would have taken I think a lot longer than being part of Google DeepMind where we have these teams that have a lot of knowledge in different areas and sort of lean and build on which I think is super exciting about our big industry right now is that we have so many experts in different areas that we can like seek out advice and help from. And so I'm here a question for you on that is having led to VO3 work, which is kind of mind blowing. Is there a reason why this is Genie 3 and not like VO3 real time?

19:53So I think it's definitely a bit different, right? Like Genie allows you to navigate the environment and then maybe take actions, right? And that's not something that veil at this point can do. But there are other aspects that are different, that with the journey doesn't have, right? Doesn't journey doesn't have audio, for example. Right. So we just think it's, it's, while definitely there are potential similarities, it's sufficiently different. Also another thing is that this point, journey free is not available, you know, as a product, and we do think about it as like a product that is, kind of makes mainstream and became very, very popular.

20:32And, and, you know, what the future holds, I don't know, but I mean, at this point, we just felt it's sufficiently different in terms of what's capabilities and how we think about this. So, Gene Freys pretty much a research preview, right? It's not something we are really seeing at this point. You know, something we think about a lot is what are the edges of a modality? We're talking about it all the time, which is, you know, the lines start blurring pretty quickly, but we in real -time image and video and then real -time video and interactive, whatever world generation, world model. I don't think we have a good word for what G3 is yet, but you guys called that world model, which is I think a great term.

21:12But in your mind, where does a video generation modalities stop? And real time worlds take a start. And do you think in the future are these converging into basically one modality? Or if you had to predict over the next few years, do you guys think actually, Yeah, these will diverge into completely different disciplines. It seems like they share kind of one parent today, which is in a video generation, but where is the world going, do you think, are these two completely different fields? From my perspective, they're different. So I would say modality is one thing, right? We have text, we have audio, even without within audio, there are different types of submodalities.

21:53Speech is not the same as music. We have different products for music generation. and we have other models for speech generation, speech understanding. So even within one modality, you can have different flavors. And then of course, you have video and other things. So I think basically, I would say the modality is one dimension and another is how fast or how quickly we can create new samples. And completely or so, maybe the direction is or dimension is how much control we have, right? So I think we can take specific direction or a specific vector in the space for Jimmy free. I think different products, different models can try and go in a different different direction.

22:42I think the space is pretty big and there are a lot of trade -offs to be made. So yeah, I don't know. I think it really depends. Some people believe there is one model that's who they're everything, or I think there is still an open end that was the best way. We're in a place where engineering is a big part of our research rights and actually making those, it's not a paper right where we want to build something that people can actually use. So I think this really makes it like an abstract idea, it's go to some to get you to some point, but to actually build things we have to make some concrete decisions.

23:16And I think it forces you to decide what you want to do and what you're doing. Yeah, I think this is a really interesting point, my end. Ultimately, it has to be driven by technical decisions and also the goals. If you look at the models right now, we obviously made a choice that we won VIVVIII and Genie III to be separate projects this year. If you look at them both as they are right now, they have very different capabilities that the other model does not have. And technically to combine all of that already into one model, I think, very challenging to... I mean, the Earthry is totally a higher quality threshold than Gene3, right?

24:00And it has very different priorities, right? So then the natural things, you could say, oh, well, what if we just took these together and combine them? But that may not be the best next step for either of those two models, right? So it may not be the case that the thing that the other one has is actually the most compelling thing for a completely different experience. And I think that given the breadth of interest in both models, right, there's actually quite a small set of people that are like really actively using both. And they tend to be more folks like yourself who are just more broadly interested in AI, right, rather than like really downstream use cases.

24:41So like you mentioned agent training for one which is like a very sort of like high action frequency requires more eco -centric. Because I guess more like worlds where tasks can be achieved. It doesn't require you know that high quality cinema, just our videos you could generate with the B .O. moderates. It's quite different. And then on the filmmaking element, I mean, I'm also sure that Genie 3 is really there at this point. And that would be necessarily the goal. I don't know, I'm filmmaking, Justin Condoves, I'm pretty incredible things with the filmmaking tools today. I would be surprised.

25:18Give me access. I will make amazing films with Jeannie Three. I guess that did kind of get to one of my questions though, which is the work you guys are doing is incredible. And you clearly probably have so much going on in your brains just to coordinate training these models and managing these teams. How much do you also have to think about like what are the downstream use cases of the model when you're training it? Because you could imagine a world in which you're just like, we don't really know where care what people are going to do with it yet. We're just going to go in the research direction.

Read the full transcript

25:48We think we should go and see what happens. But based on how you guys are talking about it, it sounds like you've also been pretty thoughtful around what are the different capabilities or features needed for different potential use cases at least of different models. Yeah, I'll say that basically we have some applications in mind, but that's not what driving the research. It's more about can we how far can we push in this particular direction? Can we make all of that work like really great quality, really fast generation, real time, very controllable. I think that's kind of what drives us. I think they're to have to develop Gini free.

26:29And the applications kind of like follow, and I don't think to be honest, I don't know what would be the applications for, I think we're very surprised. I like to mention like they're free. Well, people find new ways in how it can be useful. And to prompt it, we have like visual stuff, you know, people just discover it, right? We didn't even think about it initially. So I expect kind of the same thing. And I think that's why I am excited for more to be to be able to access it in the future. And in general, our approach is to make sure that that over time, there is more access to them, what does we build.

27:08And I think that's the only way to discover what's the rate potential. I guess one, one, somewhat related to that. How do you think going forward like, Genie 4 or 5 or any other models? Like what is like top of mind right now? Like if you want to, for example, to focus on, I don't know, Like seems like gaming could be one of the applications, having multiplayer type of games where you have two special memories or two different completely views, but at some point they emerge. How are you thinking of like going forward? Like what's next? You sit like scaling these models just on more data more computer, they see creating this sort of like multi -universe type of things where you have multiple players, multiple people looking at the same model, but in different views.

27:48What's the top of mind for you guys? Top of mind, I think for the next few days might be a vacation. After that, maybe walking my dog in the real world. And then I think you mentioned a bunch of re -engineering, thanks, for your list. And I think we are still collecting a lot of feedback on this current model, right? And I think that in general, we are most interested in building just the most capable models, right? And so we would hope to have even broader impact in future and really enable other teams to do cool things with it, right both internally and externally. And for me, it's like, I just started this with like a very, very focused vision about AGI.

28:35And I still think honestly for my what I'm excited about for AGI and which is more embodied agents. I really believe this is the fastest path to getting these agents in the real world. And I think we made a big step towards that. But, and still, like, sometimes even more excited about applications that never thought of that come up from other people seeing the model, right? So, I think it's kind of this trade -off of, obviously you want to focus on some applications, but then you want to be open -minded about others. And I think that's the real joy of building models like this, right, as you get to see all of these people can be way more creative than me with it.

29:12So I think that there's always really cool things that we can do. And honestly, don't really, can't really tell you in one year what the biggest application will be. But we'll definitely be trying to build better models. Yeah, I'm really excited. I think we're only as impressive, you know, maybe in the model is, I think they're very far from actually simulating the world accurately and being able to do, to kind of put a person in there and then do whatever they want. And I mean, when I say far, it doesn't mean it's far in terms of, you know, kind of their time because we are in a really even accelerated timeline, but it feels like there is more work to do to get there.

29:52And I think I just imagine like once we can actually, you know, work whatever the form factor would be, but stepping to this world and just kind of like maybe tell it how we want to what you want to experience. There are so many applications, imagine for example, someone is afraid of talking to people on a stage or in a podcast, right? They can simulate that, right? Or you can have someone who is afraid of spiders, they can maybe actually see themselves getting over that. So that's like, you know, just one example of something that's actually my wife thought about it. It's not my idea. So I think it's really like, there's so many things, So I think this is just, it's all hinges on the ability to simulate the world and put ourselves in it, maybe seeing yourself from the side and potentially having agents interacting with things.

30:46And yeah, the realism and really making it work in the way that it's similar to our world, I think it's really key. I'm actually personally petrified of skiing and the model is already quite good at that. So when things quiet and down, spend some time, because I promised my wife that I would be able to do it. children would grow up knowing how to ski. And we're getting close to the age where I have to live up to my promise and I'm not sure if I want to do it. So we have to improve the model for you, Jack. So you can actually get that in distribution. I hope so. We were just talking about before the we started that we might see applications like in robotics.

31:20I mean, Jack, you were talking about embodied AI and like now like limitation in robotics is the data, right? Like how much data you can collect and now probably you can just generate a little of different scenes that you were not able to do before, purely from just recording, videos or so. So I think that's another thing that is pretty exciting. And I mean, congrats on the model, it's phenomenal. On the robotics application, there was a conversation that I was listening to from Demis yesterday where he was talking about your guys' work on Genie 3. And he mentioned that there's an agent, I think you guys call it SEMA, right?

31:56which can then interact with the Genie agent. And as I was hearing him describe it, which was kind of breaking my mind, which is that you had one simulation agent asking the world, asking the Genie agent to essentially create a real -time environment for it to interact in, right? Which was when I realized, oh, the way you guys have built it, it's composable with other agents. Can you talk a little bit about why that's so important for robotics? Like Marco was saying, and what are the major limited limitations today that you think we'd have to overcome as a space to make the robotics sort of progress the read of progress and robotics much faster than it is now so We designed it to be in our environment rather than an agent right so so gene 3 is very much like an environment model Like we don't see it's like an agent itself that can like think and act in the world It's more just a general purpose sort of simulator in a sense, right?

32:52That can actually simulate experiences for agents. And we know that like learning from experience is a really important paradigm for agents, right? That's how we got AlphaGo because the agent AlphaGo learned by playing Go by itself, trying new things, right? And then learning from feedback, the reinforcement learning, learning to improve itself, and actually discover new things, like it discovered new moves, that move 37 that humans didn't think was a worthwhile move, right? But actually, I'll forget to learn that it was because it could experience and try to write things for itself. And I'm robotics.

33:25We have this paradigm right now, where there's some data driven approaches, right? Where you can collect data in a quite a laborious way. But it looks like the downstream tasks. So it looks real and there's not so much of a mismatch between the two domains. Or you can learn in simulation, right? but the robotics simulations are even the best ones and we have some of the best ones. And I deep by my mouth, I'm a joker, right, which we work with. There's still quite far away from the real world, right? And you have a SIM to real gap. But even the SIM to real gap itself, I think is kind of like poorly named because what people consider to be real in robotics is typically still a lab or some very constrained environment where you've got a bunch of spotlights on a robot and then tons of researchers crowding around watching, you know, whereas really real for me is, it's the ability to walk my dog when I'm too busy.

34:22To hold it, to lead across the street, you know, see someone who's scared of dogs, know to go around them, see someone who'll a ball, change directions, like all these challenging situations in the real world, right? And of course, you still have gripping, you still have these other tasks, But you need to really discover your own behaviors from your own experience, right? And that's that doing that in physical and bodied worlds is super challenging because there's so many reasons why firstly that could be expensive to collect data in those in those settings. You'd have to keep moving the robot back to where it started every time it like doesn't do something right and also it could be unsafe, right?

34:57So there's many reasons why we can't really do learning from experience in the physical world, right? So we do it in simulation. But really what we think with Genie 3 is it's the best of both right because you're taking a real well -data driven approach Right, but then you've got the ability to learn the simulation So it kind of combines the good parts of each of those And so that's why I think it could be super powerful Not just for or about example, but I really love this idea of having when it rains in London a lot I'm not having to take my dog for the second walk. It would be great And as you can see, we build a modern basically for Jack personal vacations.

35:38That's what driving the project is. I just saying clearly, Jack, it's time to move to California. That's the solution.

35:52I mean, I personally love California, but my wife's not. My wife's not convinced. Sorry. We're convinced here. Yeah, just to touch on maybe a final point on the robots, like robotics part, I think it's definitely robotics means it's more than visual. All right, we need to be able to, I think this is an important point. We want, we can drive the decisions of the robot by looking around, but still it has to do, to do extrautions, decide where to move, how to respond to the environment. So I think there are definitely some gaps, but still a decor of the problem being able to reason about the environment, we think this is something that's the, you know, word models, general purpose word models such as Gini free can really help with and maybe with future research, we can actually bridge those gaps of physical, kind of like understanding and actually getting responses, physical responses from the wall, which is a very interesting direction to explain.

36:50One last question from my side, the, another of you can answer this, but like, Is you going to become public, like, a developer's access to that at some point? Or is there, like, some sort of idea on these? So as you can see, we are very excited about having more people accessing it. So we're definitely want to make it happen. There is no kind of a concrete timeline at the moment. But, you know, I'm sure, once we have more to share, we do. Awesome. One of the things I've been thinking about a lot is we see sort of with every, like, modality, like, you know, maybe first LLMs and then image and video and audio.

37:25There was early kind of glimmers of something really exciting in a product or research preview. And then there's a ton of data and compute and researchers kind of poured at the problem. And you hopefully see this sort of exponential progress till you eventually get to the point where you're out of data or the improvements don't come as easily. I'm wondering for your thought, like where we are on sort of that curve for world models. That's a really question. I actually have a super hand wavy, somewhat swerving answer, right? And I think it's actually both. So I think the current capabilities are actually already quite compelling.

38:04And so you could make the case that like if what you wanted was a minute of photorealistic any well generation with memory, that could actually be the end goal, right? And two or three years ago, I probably would have said that was a five year goal. And so at that point, if you just wanted to improve that, I think you probably end up with this maybe like, I think the jump from Genie 2 to Genie 3 was was absolutely massive. And when from being like kind of a cool bit of research that was like showing signs of life, something that could already be very compelling. But I think there's a lot more that you can do with this and and Schlimme kind of references to himself, right?

38:42Like it's not the case that you're dropping yourself in the world, right? And like it's like the real being in the real world, for example. It's actually quite different to that. When you do, you know, take a minute to look away from the screen, it's quite a bit richer out there. And that's just for the real world. We also want this ability to generate completely new things, right? So, I think we've got a huge gap to close, right, with the new capabilities that we want to add. But I think it's maybe a bit different to language models. Well, actually maybe it is similar to language models, But with Lange Tommel's, there's been like lots of new steps that have actually come on top, right?

39:19That maybe we didn't think of a possible, we thought things were plateauing and then a new idea came that made a significant change. And that has happened a couple of times in the past few years. So I think that there's a few more of those left, for sure. My final question for you guys is, are we living in a simulation? Oh, yeah, that's every young, just to add, I think. I am. My thinking about that is actually, yeah, I thought about it a bit. I think that if we live in a simulation, my take is that it doesn't run on our current hardware because it's analog and not like, you know, it's continuous to all of the observations that are continuous and there is nothing like.

40:02But maybe the quantum level is, you know, some limitation of our, you wanted to go philosophical. It's some kind of like a hardware limitation of the simulation we run on. So, yeah, take it or leave it. It's a great answer. Clearly it's all a work for the TPU team to do. Yeah, maybe quantum computing will be actually, we'll be running our actual simulations. So yeah, yeah. That's a great place to wrap. Still on me, Jack. Thank you so much for coming to the podcast. Thank you guys. Thank you guys. Thanks for listening to the A16Z podcast. If you enjoyed the episode, let us know by leaving a review at ratethispodcast .com slash A16Z.

40:45We've got more great conversations coming your way. See you next time. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the company's discussed in this podcast. For more details, including a link to our investments, please see A16Z .com forward slash disclosures.

From the publisher

Genie 3 can generate fully interactive, persistent worlds from just text, in real time.

In this episode, Google DeepMind’s Jack Parker-Holder (Research Scientist) and Shlomi Fruchter (Research Director) join Anjney Midha, Marco Mascorro, and Justine Moore of a16z, with host Erik Torenberg, to discuss how they built it, the breakthrough “special memory” feature, and the future of AI-powered gaming, robotics, and world models.

They share:

  • How Genie 3 generates interactive environments in real time
  • Why its “special memory” feature is such a breakthrough
  • The evolution of generative models and emergent behaviors
  • Instruction following, text adherence, and model comparisons
  • Potential applications in gaming, robotics, simulation, and more
  • What’s next: Genie 4, Genie 5, and the future of world models
     

This conversation offers a first-hand look at one of the most advanced world models ever created.

 

Timecodes: 

0:00 Introduction & The Magic of Genie 3

0:41 Real-Time World Generation Breakthroughs

1:22 The Team’s Journey: From Genie 1 to Genie 3

5:03 Interactive Applications & Use Cases

8:03 Special Memory and World Consistency

12:29 Emergent Behaviors and Model Surprises

18:37 Instruction Following and Text Adherence

19:53 Comparing Genie 3 and Other Models

21:25 The Future of World Models & Modality Convergence

27:35 Downstream Applications and Open Questions

31:42 Robotics, Simulation, and Real-World Impact

39:33 Closing Thoughts & Philosophical Reflections

 

Resources:

Find Shlomi on X: https://x.com/shlomifruchter

Find Jack on X: https://x.com/jparkerholder

Find Anjney on X: https://x.com/anjneymidha

Find Justine on X: https://x.com/venturetwins

Find Marco on X: https://x.com/Mascobot

 

Stay Updated: 

Let us know what you think: https://ratethispodcast.com/a16z

Find a16z on Twitter: https://twitter.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Subscribe on your favorite podcast app: https://a16z.simplecast.com/

Follow our host: https://x.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
Google DeepMind Lead Researchers on Genie 3 & the Future of World-BuildingThe a16z Show · 41 min
Listen in VO