In short
Podcast Summary: Google AI: Release Notes - Episode on Project Genie
Episode Title
Project Genie: Create and Explore Worlds Episode Description This podcast episode dives into Project Genie, a web application developed by Google AI that allows users to create and explore interactive worlds. Host Logan Kilpatrick discusses the technology behind world models and reinforcement learning with guests Diego, Shlomi, and Jack.
---
Key Topics Discussed
- Introduction to World Models
- Definition: A world model generates environments based on textual descriptions, allowing for interactive experiences.
- Origin: Builds upon reinforcement learning where AI agents learn to navigate and interact within these environments.
- Project Genie Overview
- Release: Launched as a web app for Google US ultra subscribers.
- Functionality: Users can create customizable worlds by defining characters and environments. Utilizes the Nano Banana Pro to create canvases before entering the world.
- Demonstrations
- Initial Demo: Showcased an underwater scene with a goldfish and a shark.
- Interactivity: Users can navigate their created worlds, interact with different elements (like bumping into objects), and change environments dynamically.
- User Experience and Feedback
- Trusted Testers: Gathered insights from initial users to improve functionality and explore new features.
- Challenges: Addressing constraints like generation limits and the need for improved interactivity and controls.
- Physics and Realism
- Real-Time Interactivity: The model generates environments dynamically and attempts to simulate realistic physics.
- Challenges in Realism: Maintaining consistency across frames poses a significant challenge as the environment evolves.
- Future Applications and Development
- Potential Use Cases: Education, entertainment, and robotics represent significant areas where world models could be applied.
- Research and Innovation: Continuous improvements and new applications are anticipated as feedback from users is analyzed.
- Collaboration Across Google
- Team Efforts: Multiple teams across Google, including Creative Lab and infrastructure teams, collaborated to bring Project Genie to life.
- Adoption and Market Impact
- Adoption Timeline: Speculations around how quickly world models will become mainstream and their impact on daily life.
- Interactive Media Evolution: Discussion on how Project Genie could redefine media consumption, shifting from passive to interactive experiences.
- Future of World Models
- Progress Trajectory: Imminent advancements in world modeling tied to hardware improvements and more accessible technology.
- Exciting Outlook: The future aims towards higher fidelity and immersive experiences, potentially leading to a full universe simulation.
---
Key Takeaways
- Innovation in AI: Project Genie exemplifies the cutting-edge capabilities within AI technology, showcasing how generative models can create highly interactive environments.
- Importance of User Feedback: Early user experiences and insights are critical in shaping the future developments of Project Genie.
- Cross-disciplinary Collaboration: The successful integration of diverse teams illustrates the collaborative nature of innovation at Google.
- Path Forward: As technology evolves, the potential applications of world models in various sectors, from education to entertainment, are vast and promising.
---
Conclusion The episode provides a comprehensive overview of Project Genie, highlighting its innovative approach to world modeling and interactive environments. As the technology continues to develop, the implications for users and applications are vast, paving the way for a future where immersive experiences could become commonplace.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding World Models
0:45 to 1:57
Exploring what world models are and their interactive capabilities.
“So it's an interactive version of this video generation models.”
Introduction to Project Genie
1:57 to 3:13
Overview of Project Genie, its features, and its release for Google US ultra subscribers.
“I feel like everyone was blown away in that initial, the initial Genie 3 video with Chef's Kiss.”
Creating Your Own World
3:13 to 4:24
A guide on how to create personalized worlds using Project Genie.
“And what happens when you actually like you click the button generate world it's just like sending a request to the...”
World Interaction and Dynamics
4:24 to 6:24
Discussion on how users can interact with worlds and the dynamic elements involved.
“This gives you basically a starting point.”
Physics and Realism in World Models
6:24 to 7:59
Examining the physics aspect of world models and their realistic responses.
“I mean, there's be latency issues, this kind of thing where you're both interacting with the same model.”
Challenges and Limitations
7:59 to 9:46
Addressing the challenges faced by the model and its current limitations.
“and the model is generating a video across time.”
Bringing Characters to Life
9:46 to 11:28
Exploring how characters can be integrated and interacted with in the worlds.
“What other cool examples that sort of showcase?”
Style and Customization Options
11:28 to 13:19
Discussing the customization options available for visual styles in Project Genie.
“And let's just say the room and the toy in the picture.”
Performance and Latency Challenges
13:19 to 14:00
Analyzing performance issues and latency challenges within Project Genie.
Exploring Time Limits in Virtual Worlds
14:00 to 17:00
Learn about the optimal time limits for user engagement in virtual environments.
“I mean, it's maybe not that exciting, but it is.”
Show all 20 chapters
The Evolution of Project Genie
17:00 to 21:00
Discover the development journey of Project Genie and its impact on user interaction.
“I think that's pretty much where the direction we're trying to go.”
Future Applications of Project Genie
21:00 to 24:20
Explore potential applications in education, robotics, and personalized experiences.
“Maybe they're afraid, a child is afraid of doing something, right?”
The Creative Potential of Interactive Worlds
24:20 to 28:00
Understand how interactive environments could reshape media consumption and creativity.
Exploring Creative Potential of Genie Worlds
28:00 to 29:15
Discuss the vision behind generating realistic interactive worlds through AI.
“Like, it's something that we could use for a bunch of things.”
Collaboration Across Google Teams
29:15 to 30:28
Insight into the collaboration among various Google teams to realize Project Genie.
“the little Bob dinosaur example, like, you are sort of recreating your space.”
Timeline for World Model Adoption
30:28 to 32:19
Speculation on when the average person will interact with world models in daily life.
“to now being something that's talked a lot about by different people in media and social media in particular.”
Development of Project Genie and Its Capabilities
32:19 to 35:58
Discussion on the evolution of Project Genie and its interactive capabilities.
“I think that's definitely something that we will, like more and more people will have access to like this year and I think next year, basically.”
Creative Potential of Interactive Worlds
35:58 to 38:29
Exploration of how users can create and modify their own interactive worlds.
“right like i think simr is a project we've partnered with for a lot longer and they were part of genie 2 as well they played with genie 2 so they've already had a release using genie 3 What is Simr for public?”
Hardware Constraints and Future Potential
38:29 to 40:44
Examination of the hardware limitations affecting the development of world models.
“Yeah, this is one of the ideas that we're super excited about.”
Discussing Project Genie Feedback
42:00 to 42:16
Learn about the excitement surrounding future iterations of Project Genie and the importance of feedback.
Transcript
Automatic transcript. May contain errors.0:07Hey everyone, welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today we're talking about Project G &E, G &E3, joined by Diego, Shlomi, Jack. Maybe we can kick off with like, what is a world model for folks who have no idea what we're talking about or what we're going to talk about? So I think one way to think about it is you have video models that basically generates, you know, you put it in text, you get a video of a few seconds, right? So a world model is able to do not just that, but it actually, you give it a description of an environment, and then it starts generating the world one frame at a time.
0:42And at every point, you can decide that you want to go right, you can go left, you want to do something in this world. So it's an interactive version of this video generation models. So that's, I think, like, international allows us to have experiences that we've demoed in the past. that you can just walk around, you can interact with the environment and kind of like be present in this world. That's why we call it the world model because it simulates tools. There's, I guess, some history of this coming from like reinforcement learning. So you would have used world models as a way to have like AI agents learn how to do things in the world.
1:17But what we realized is with this new generative AI technology that the ability to generate new worlds is actually kind of compelling, kind of fun. And so that's, I guess, more where we are today. Yeah. And for Project Genie, what is Project Genie? Yeah. So Project Genie is the web app that we built in partnership with Google Labs and with Creative Lab. We're releasing that today, actually, to all Google US ultra subscribers. And so you can think about it sort of like what Flow is for Vio or is still for Vio. And so it's actually basically just an app where you can build your worlds. And we'll see a demo of that shortly.
1:49But you can actually define your character. You can define your environment. You can use Nano Banana Pro to create a canvas before you step into your world. But we'll walk through that in the demos. I love it. Cool. Well, I'm excited. Let's maybe we look at some demos. Let's do it. And see the model in action. I feel like everyone was blown away in that initial, the initial Genie 3 video with Chef's Kiss. They did a really good job on just the narrator, the voice, all of the stuff. Yeah, that was also work from the Creative Lab team. You can just sue and crush it. Yeah, so the Creative Lab team actually put a ton of effort and like a breadth of worlds that you see in this gallery here, which were super cool.
2:26and we'll walk through one of those in a minute. But let me just walk through how you create your own world. So the cool thing here is with Project Genie and the Power of Genie 3, you can actually just like dive in and create whatever world you want. Let me walk through maybe a bit of a simple one to start off. So let's say a coral reef and then a goldfish. And basically the first step here is you're doing a quality Nano Banana Pro. And what we found with our trusted testers is that actually there's a lot of value in creating this sort of canvas before you dive into the world. so you'll see in a second you'll actually have like an image generation here and then you can also see we have other controls that we worked on in partnership with google labs down here like uploading your photo we'll try that in a minute um and a few others that we'll walk through can we change the make it like a shark instead of a gold fan yeah and i absolutely do that so change this how how complex can we get could we add like uh something else in the background like you know there's a the hull of a ship or something like that i don't know is that possible yeah cool let's try that out intrigued i also feel like this is a cool story of like this energy of uh of the nano banana model like actually impacting this uh this product experience exactly awesome yeah and i think one one of the sort of wow moments that we learned from trusted testers is the second you go from kind of like a 2d asset into this stepping inside that asset is just mind-blowing perfect so this works so the power of nano banana pro yeah now we're going to dive into it.
3:53And what happens when you actually like you click the button generate world it's just like sending a request to the... So there's some behind the scenes basically we're preparing a more elaborate description of this world but you see it doesn't take... oh wow. This is awesome though that is actually super interesting. Can we go to the ship? Let's see if we can get inside this ship. And we didn't get eaten by the shark as well which is nice. Fish and friends not food exactly how much is there like interaction between um like is are the worlds ever like doing something that sort of like proactively against your sort of like controllable character and not really a game i mean it's not really a game right like it's like pretty much um you can navigate and get around the world you can see other characters you can sometimes bump into things and move them you know kind of by navigation but at this point we don't have much happening in terms of like a narrative or anything like that but in principle there is nothing that blocks that or prevents that that's super cool in this example you saw that we were controlling the fish obviously one of the aspects you want to evolve over time is more controls so how you can actually give like new cool things to do in the world as well so let's go back to the gallery here and i'll show you one of my favorite worlds here which our creative lab team put together and i think the cool part about this one is let's say that you jump into Project Genie and you don't know where to start.
5:18This gives you basically a starting point. You can build on top of these worlds. You can use them as inspiration. You can sort of use that as a starting point. This was my, when I was playing around with Genie, this was my biggest question, which is I like, I sort of had the blank prompt enter Vox, and I was like, I don't know what to do. I feel like this is so powerful, but it wasn't clear. So I feel like these gallery examples bring a bunch of unique ideas to life, which is really awesome. It's kind of the thing with like a new kind of model, right? Is you haven't got like your social media feed full of different examples to how to use it.
5:47And I think we saw with VO3 when it came out at IO last year, so many cool things people started doing with it that I personally would never have guessed. And I'm hoping that this next set of users from Project Genie will find cool stuff that we haven't thought of as well. Yeah. Well, we were talking off camera before about just like the, you know, basically for one of these worlds, there's like a dedicated amount of compute to like run this world. Does that mean we're like within reach of, you know, I could join in this world with my paper airplane and like do something else? Is that like, is that possible?
6:19That's definitely something that's like an exciting thing on the roadmap for us. But definitely that's a little bit more complex than a single person. I mean, there's be latency issues, this kind of thing where you're both interacting with the same model. And I think there's many things that we'd have to figure out there, but like definitely that's something that we're excited about. Yeah. The other thing, this is sort of implicitly captured in this world, but I don't know if you saw it, you could see the water actually splashing. So I don't know if you want to talk about the physics here, because I think that's one of the cool aspects of the model as well.
6:46Yeah, basically the environment is responding to the character in this case or the plane. And in general, what the model is trying to do is it tries to create a realistic, as much as possible, depiction of the world that you're creating frame by frame. So if you bump into something, we can show some examples where you bump into a ball or something and just roll. So the model kind of implicitly learns how the world evolves. When we think that's beyond the entertaining aspect here, it's pretty powerful for other applications as well. I'm also curious, like on this, the like physics, realism of physics, is there some interplay between like the research breakthroughs that we've had on VO actually like enabling some of these things?
7:36It feels like somewhat similar-ish maybe? Or no, it's like completely similar. No, it's definitely related. And I think the key point here is that this is, in a way, a harder problem to solve. Because when you create a video and the model can just change, in a way, the frames in the past, in the future, across the entire video, it has more freedom. So it's basically, imagine that you have this canvas and the model is generating a video across time. And it can change a little bit here, a little bit there, and it can make it consistent and looks good. Now, here we have a more difficult problem, by definition, because the world can evolve, right?
8:13And this is like basically from frame to frame and we don't know what the input's gonna be. So it has to be consistent with the past and with the immediate action. So it's kind of like a harder problem. Do we know what the prompt was for like this ball? Yeah, so actually maybe let me walk through that. Yeah, I need this because I'm like my prompts, I've gotten so lazy because the reasoning models just like invent prompts for me behind the scenes. So let's actually use an example. So let's say you wanna go in here and you're like, I really like this world, but I want to change the ball to, let's say, a red color.
8:44So you click on Remix, and then you could just actually add kind of a change. So change the ball to the red one. This is another Nano Banana example. Yeah, exactly. When we started Genie 1, I think we were using Imagine 1. And like, the model obviously wasn't, our model wasn't great, right? But also the images were, it's quite tricky to get one that worked. I think Nano Banana came out after Genie 3. and it's just it's crazy the progress and we have that here maybe one day our worlds will be defined by more than just an image and text right but for now it gives you a lot of flexibility to create the worlds i think we envision even being able to like you have a video or you have some some content you like and you know at some point you can just walk into that scene right and just walk around talk to people it's just immersing yourself in this world yeah i like the red more than the blue.
9:35It's a bit more dramatic. Logan, to your question, you can actually see the prompts here. So after the world, you can click on reuse prompts and then you'll see actually what we used in the background to try out these. Nice. This is awesome. What other cool examples that sort of showcase? And I'm curious like what the, for folks who are like trying this out or just trying to conceptualize, like what's the upper bound of complexity? Like, is it really like, can you add, you know, many, many things happening in parallel or like where does it start to and actually like for nano banana it's a good example like if you want to edit an image like if you have 10 people's faces in an image and you're trying to do it like it's sort of the model starts to trip up in those types of cases i'm curious like what the natural edge is yeah i think the model is definitely not perfect and and it's it's you know it's a review and we want to make like in a way we want people to play with it to see why it's doing well what still doesn't and we can learn from that.
10:31Definitely, it's doing pretty well on more of those creative environment that it can be very elaborate visually. In terms of the dynamics of the world, we sometimes see that it can start a bit more dynamic and then maybe lose some of this dynamism over time, which is something we're also working on improving. But I think it's very surprising. like it's hard to know so so we i do a kind of like um kind of like recommend you to try and play with it and see what works um do you want to maybe try an image yeah so one of the cool examples we saw from trusted testers is this idea of like taking pictures of your world and also of objects and bringing them to life so we have a special guest over here uh bob the nano banana dinosaur um and just now we took a picture so i actually uploaded it uh to the app here and we're going to see a demo of what it's like using this basically as conditioning to the model and then actually generating a world from there.
11:30Nice. I'm excited to see Bob come to life. And let's just say the room and the toy in the picture. Perfect. And now we're going to actually just dive into this world. All right, here we go. So we have Bob here, our nano banana mascot in the library, the picture that we just took. Yeah. So you can see I can actually control the character in this real scene. Get him out of there. Gave him some nostrils. But the cool part is like you can effectively bring any character to life.
12:04I jumped on the chair you can get out of it. This is where we find out Diego's actually bad at video games. If you press K so you're a good wife you can jump. Yeah yeah I'm on the table now. Oh man. This is awesome I like this library more. He wants to escape. Get out the window. It's, I think the mother will let you go through the glass. Could we, like, I'm... Oh, there's no glass.
12:34It's a big highway. It thinks that there's a big highway outside, which thankfully there's not. Oh, nice. Wow. That is honestly crazy. Oh, and it's a double-decker library. That's awesome, actually. This is my favorite example. but the cool part is like you can effectively bring any character to life which is really powerful it's a different way of experiencing a photo for example yeah that's wild and so from I'm curious about like the style of this and I don't know if this is like a the default style for Project Genie is just like trying to do photo realistic-ish but like can you if we wanted to like make this look like a cartoon world instead of this it can do that type of like style transfer between these things yeah we just like because the photo the input photo was photorealistic or the image so so it just goes with this uh style but you can change it on whatever you want to manga or like on the latency topic though there's also there's a bunch of things right i think the constraint for genii 3 was that we wanted to be real-time latency with the certain action frequency which means that the round trip for an action is very very like low latency and then also the memory right and memory is a huge constraint as you also know in other lengths right the more context length you have the generally more expensive and slower the models get right yeah so to do all of these things in one was the challenge of basically competing objectives and on the research side we're kind of making improvements in all of these areas right so we think that the frontier should just keep improving more capability faster cheaper these kind of things it's just trends that will happen how i've heard i have a question how does it um sort of like are we just artificially cleanly cutting it off at 60 seconds or is it like you really could do three to five minutes longer than that even some demos of that but we felt like it's a good time that gives people enough of the environment but um we we can serve it to that like reasonable number of people it's somewhat of trade-off observing costs yeah and check me mentioned before like there's there is this like issue is the further you go into it the world's like the dynamism can mainly gradually reduce so we figured that like the marginal enjoyment of that second minute is maybe more fun to just spend to try two different worlds for a minute each yeah um but if we've got some feedback already people want to spend longer on it then we can of course do that it also depends on the kind of world right if you're like doing downhill skiing maybe two minutes feels great because it's like you're just kind of continuously going whereas if you're just kind of exploring the library.
15:03I mean, it's maybe not that exciting, but it is. I think in the future, we see more like this going beyond into more interesting environments with more kind of like maybe characters you can interact with more, but beyond just navigation. So there are so many directions where we can build on top of that. And I think when this is becoming more engaging, more interesting, then it definitely makes sense to extend this time way longer. Yeah, one of the interesting insights we learned from trusted testers is people want to adapt what they do in the world based on the context of the world so for example if i have a door obviously i want to open the door but maybe if there's an object i want to carry that object and depending on what the object is like it could do different actions um so it's an interesting sort of like non-deterministic way of approaching like interactions in a world it's also i guess worth noticing this is another thing like we've seen other models right um a year ago when we were thinking we'll get a minute of consistency for an autoregressive model like this in real time.
15:59People thought that was like our stretch goal kind of thing. And then once you actually land this, because Genie 2 had come out, you know, just over a year ago, and that was 10 seconds, low resolution, not real time, and only for more like, you know, different kind of vision environments, not really photorealistic. So all these things we added, and then to go for a minute and people say, oh, but a minute actually isn't long enough. I think that's a sign of the progress, right? That you want to see it for more than a minute. yeah people we all got used very quickly to how things work right like and it's just but to me it's still mind-blowing because i'm always working fast enough you asked if we can make it faster so i think there is in a way we're at the limit of it doesn't like because the model is generating the world as we go it doesn't make it doesn't make a difference anymore if it's even faster right because already generates as fast as we can consume it pretty much um so it is a lot around making it like reducing the cost so we can bring it to more people.
16:56I think that's minima, I guess, as we make crowd quality better, I think that's pretty much where the direction we're trying to go. Yeah. Even just like maintaining speed, but adding all these capabilities is going to be a big challenge, right? And making it even better in all aspects. And as we looked at, I'm super curious to talk about like the future iterations of this, but actually just to make sure we hit on the sort of trajectory to get here, we announced the awesome, we had an awesome launch video for Genie 3 in August. And I think we've been doing the trusted tester and sort of iterating and building infrastructure.
17:28We're like, actually, what's happening? Can we talk really quickly about the entirety of what's happened between really cool video early testers to like actually rolling it out as like project Genie for people to experiment with? So I think one aspect to consider is that from the mall, and so we announced them all and we had some demos and we also made it available to a handful of people back then in August because we wanted to get some feedback. It's pretty much a very new application, a very new experience. We want to see how we should responsibly bring it into the world. And since then, a lot of the work has been on the infrastructure and serving infrastructure and on the cost side because we want to make sure that we can bring it to a reasonably large number of people and we think that with the ultra program basically in the us we can basically have enough people that play with it and we can get a first sign of what it's useful for how people interact with it what people find to work well for them and we also during this time invested in and kind of like running a trust in this program yeah i think that's really a core part of this period in this in the model development cycle because it actually enables us to learn sort of breath first from different types of users so everywhere from creatives to folks in education and so on so we got really rich insights on what is the model like actually sort of currently useful for where do they see it going what are those like north star experiences where we can go yeah i guess just to round it off like going back to august um obviously in hindsight it's oh yeah people want to play with this but we were doing it it was it was very much a research project at scale.
19:06And so we have many use cases in mind. And we have the agent ones, we have, you know, more embodied agents, like robotics could be a user of this. And we have the Simma project, which did a launch Simma 2 in the end of last year, and they used Genie 3 for training their gaming agents as well. So these use cases, and then from a more like consumer standpoint, we thought it'll be interesting, and that's why we wanted to get feedback, but we weren't sure that it would be at the point yet where we would want to launch it to more people but we found i mean and diego alludes to it but we found in the cross the tester program that actually there is there is this wow moment when you first play with it uh and we we want to learn more about the different use cases and so this is like a really big step towards learning more right in terms of how people interact with it um maybe maybe a year ago i wouldn't have expected it to be this compelling right but i think it it's already something that's quite fun to play with so we're excited to see what it gets used for i expected and i believe in i believe in you all to to come up with cool sweaters so i got the sweater i was pinging people in august saying we got to get this in the hands of more people i was like we should just set up a physical genie three station somewhere and just people 24 hours a day going there and building worlds i feel like that's uh that's what we need i feel like i feel like it's one of those things that you see online and you see a video and you can sort of be like, oh, that's interesting.
20:26And then you experience it and it's like completely, it's like a very different setup. So actually like, you know, we're in this present moment, Project Genie coming out to more people, Pod by Genie 3. What's the next step? Like, what are we, like, I'm sure there's a million things we can do more on research and product experience wise, but like, I'm curious what the, maybe on like future use cases, but also just like technically from a model perspective, what we're excited about. So on the application side, I think, you know, it's personally I'm very excited about the kind of entertainment or education applications so one thing that we can help to have is that by having people access that kind of like have develop and see what type of applications we can already build today with what we have so we were you know education is one of the things that we are very excited about we can basically have people interacting in some worlds or we can have imagine that you can take like a person that maybe they want to kind of like have a new kind of experience that they couldn't do in another way.
21:30Maybe they're afraid, a child is afraid of doing something, right? And they want to see themselves actually doing that, getting like walking around the room full of spiders, if you want, right? My kids are afraid of spiders. So I think this kind of like new experiences are very personalized and can be very specific. We hope that that's already going to be valuable. So I think that's maybe in the more immediate term. On the other side, I think, as we talked about in the past, robotics and the work models for embodied intelligence is something that I think there is a lot of potential there. There is still more research to it that we have to do in this space.
22:07But personally, I'm very excited about that. In the natural, basically, the idea is that if you have a model that can simulate an environment, maybe we can use that to train robots in this world or, you know, embodied intelligence or actually use it in real time to help the agent decide how it should handle. handle. Project Genie, as incredible as it is, is very much a starting point right now. And so we definitely see this evolving in close partnership with Google Labs in terms of the things that you can do with it, the controls, maybe the scaffolding of how the app is structured and so on. And then also more surfaces, right?
22:44Like eventually we also want to bring this beyond just Project Genie. So we'll be exploring kind of the breadth of this as well. Developer API. Developer API as well. Let's bring it to the world. I think it's also like a, I mean, you know, not to plug developers, but I do think it's like a helpful, you know, developers have lots of like economically impactful use cases that are top of it, which I think is like the thing that's interesting is like what actually is the product market fit other than just, yeah, like beyond entertainment as a use case. Yeah. And I think the other cool thing is like a lot of these features are actually very similar across use cases, right?
23:17There's like broader interactivity, right? I think it's safe to say that like robotics won't be solved just with WASD and arrow keys. I think we might want more controls than this right to to have our like future you know robot assistants or whatever we're going to end up with um but at the same time i think when you walk around there's a lot you people want to do right um and so for now i think the cool thing is we've we've managed to work with or speak to all the different teams by releasing such like being one of the first to have this kind of model in august um and so we're taking all of this feedback on board and basically everything that gets brought up we we have like a list a massive laundry list of things we want to try out in the next iteration models um so i'm excited i'm excited i think i told jack so far like i feel like we have achieved 50 of what we wanted and like because we set we we always set very ambitious goals and i think there is so much more we can do in this space like so many things that the mall is still not doing well enough and that we should um so i think it's it's it's pretty much um like there is such a huge headroom in this space and it's we're only beginning like you said we did achieve you 50 % although we didn't let you 50 % it's a glass half full that's me you know it's quite similar to when you write papers I mean once you finish the project you're immediately thinking like oh but the next one could have all these things and be so much better yeah I think that's exciting right we've we've definitely seen like in the community a lot of interesting world models right and people kind of doing things maybe more similar to Gini 3 but in our mind we're already kind of thinking like I think quite further ahead anything else um top of mind as far as like how we got to this point or like interesting things for folks to to think about as as also as they like experience and and try the model like i'm curious also for you three as you spent way more time than the average person um using the model any sort of like suggestions or uh things directions to push people in i think the personalization or the ability to bring something that just it's yours so you can't do it with any other system so of course like if you um to create a world that looks like um i don't know some gaming environment of course it's interesting it's cool but ultimately um you know there are other systems that do that but i think the ability to bring something that's like um like we have this toy or we have something that's like or you want to see yourself in some specific style walking around some environment that you care about this is very unique and it's kind of reminded me some things that also some people did with like veo in like early days people like we've seen people are really interested in like um even there were some projects to try and recreate memories for like people who suffer from alzheimer and like they actually relieve their um their childhood memories through veo which was a very cool um research project that that um we had so i think this is like one example where bringing your personal things into the or like into life i think that's a very interesting direction so i I would just, one of the things that you could do is that, yeah.
26:12Yeah, I'd say one thing I'd add is like, you'll find for sure that it's not quite as robust on the prompting side, right? But that actually is an opportunity because the people playing with it now, right? You're going to say in a few years when this is very polished. I remember when it was, you know, the prompt mattered so much. Like we do with LLens now, right? But we're used to this. Like with V03, like it's pretty much every video looks good. With Genwai 3, pretty much every prompt kind of works, right? But there was a time when it really mattered. I remember there were people who were experts at like, spending 10, 20 minutes just crafting a prompt, right?
26:43And so, yes, it may not work the first time, but I think it's worth persisting with because it's such a new kind of model, but it might latch on to something in a way you didn't expect and do something that's actually really kind of interesting and fun. And it's kind of like you yourself are experimenting with frontier technology, right? Rather than consuming a product. I love that. Project Genie, prompting is back. exactly yeah i think that one of the interesting product questions for me is what happens when you actually unlock like everyone becomes or passive media consumption becomes interactive that to me is like a super interesting uncharted territory there's been attempts at this in the past but now with this sort of like truly bespoke interactive media narrative it'll be really interesting to see what that does for just the overall kind of meat and entertainment space and how that evolves and we're super excited about seeing that evolve yeah it's gonna be crazy one more thing that i I think you can create challenges in this world.
27:41You can kind of create some world and then send it to someone and tell them, okay, can you do that? Can you get from this point to point to that point? So it's kind of like a basic kind of gaming experience with a goal, but it's already there in a way. Like you mentioned the ball that you kind of draw with it, right? You can say, okay, are you able to write your name? So there's some very primitive things. I think it's nice to Jack's point. like it is very kind of basic in a way but it also brings a lot of potential creativity there's one with the rings too right where you can try to fly there that was actually that was an idea of that yeah another thing i think is people often ask about like how the frontier and like what we're going to do next right and one thing i often do is i look at the genie worlds for a long time like first person real world ones and then just look out the window and be like look at the contrast right and i think ultimately there will be a point when it's like almost indistinguishable for reality right and and we're not going to maybe go into that topic for today but like in terms of model capabilities right which we've been grounded in that right it's clearly a long way off right um but wouldn't it be incredible if you could just generate like really realistic worlds and move around interact you know do lots of other things in there and so that there's clearly that guy yeah that's that's the vision that's ultimately driving i think this line of work is that you know i like to say it's in the way like imagine you had a copy of the universe that you can do whatever you want in it.
28:58Obviously, it's a useful thing, right? Like, it's something that we could use for a bunch of things. So I think that's, to me, like, obviously, it's a very, you know, far-reaching, maybe even impossible goal. But it's like, as a North Star, it keeps driving us, I think. That's why when we were doing, for example, the dinosaur, the little Bob dinosaur example, like, you are sort of recreating your space. And so it's this idea of, like, what happens when you start to bring in your real space and actually do, like, interesting augmentations. I think it'll be really interesting to see. I love it. So by Genie 5, we're going to find out we're really in a simulation.
29:30We're not in Genie 5. Yeah, I mean, I'm not way back. Maybe we can talk for a little bit about this sort of like cross Google collaboration. Obviously, I think it's like just from hearing all three of you talk about this and from seeing some of the behind the scenes, like very difficult to pull all of this off. So like who are the teams across Google who are involved in actually making this happen? I mean, yeah, it's definitely a breadth of folks involved behind the scenes. Google Labs, folks across Creative Lab, which did a lot of the worlds that you'll see in the gallery, folks in the serving teams and then infrastructure teams.
30:04There's basically just a whole team behind the scenes to pull this off. And it's really been a sprint, you know, ever since we announced the model in August, which has been an incredible, like, heroic work from everyone. Yeah, I guess we've also worked a lot with the communications team as well, like, because trying to, like, as we see in this podcast, like trying to explain a completely new class of model team book that maybe they haven't experienced before. I think it's quite a nuanced topic, right? Something that came from maybe a more niche field in reinforcement learning to now being something that's talked a lot about by different people in media and social media in particular.
30:37I think it's been quite important to get this across the right way. Yeah. If you look back to where we started, you know, research on this in this field, right? And so it's been pretty much a very, like an open question, if it's going to work at all, or are we able to make it like real-time interactive and good enough quality and all of that. And I think it's really, for me, it's super exciting that we're able to actually, you know, kind of close this loop and bring it all the way from a research idea to a mall that we announced and now an experience that we can actually bring to people. I think it's, you know, it's not obvious and it really shows kind of like the ability of, you know, groups within Google to work together across the entire stack.
31:19I think that's great, really unique. Yeah. This is my favorite tagline is research to reality. I feel like this is like a perfect example of that, which is like, we had some idea of something we wanted to do truly a research project. Now seeing the light of day as, as project genie. Um, if you had to, I have a, maybe a spicy contentious question, which is if you had to guess, like for, for each of you, when the timeline for like most people to have experienced a world model, do you think it's more so that like world models will impact people, you know, through, you know, better products in their daily lives by like the companies making them, having world models involved in like, you know, improving manufacturing efficiency or some random example like that?
31:59Or like, do you actually think the future is like, you know, the average person is actually interacting in their daily life in some way with a world model? I'm curious if, And if that's the case, like what, what would you see as like a timeline to actually get to that point? It depends on how you define a world model. Basically, like different, I would say if you, we think mostly about interactive kind of like audiovisual experiences, if you want. I think that's definitely something that we will, like more and more people will have access to like this year and I think next year, basically. And then we will see like some applications where it really shines and I think it will become pretty much basic.
32:37like in many applications but as even if you look today right like how many people are actually using vid generation in their day-to-day it's not that common right it's pretty much so i think there is still like there is still a time for for adoption and really finding the right applications for that because video is not the same as images and you know so on that but i i on the on the kind of applications through breakthroughs in in maybe embodied intelligence um it's hard to give a timeline for that but i think it's definitely happening already like we are starting to see some good um good progress in that yeah i imagine it's based on demographic too right so um like i know that some people probably spend a lot of time you know with more interactive media type existing interactive media maybe that they'll jump onto this uh kind of as early adopters and they'll know what to do, right?
33:29Whereas if I were to give this to, you know, maybe a family member who wasn't as in tune with work and generally doesn't ask me too much about work, they probably wouldn't find it as compelling, right? Because they wouldn't know what to do with it. But some of the embodied intelligence things that will probably could even be in the real world, maybe using this kind of technology in the next one to two years, maybe they would interact with that, right? So I think it depends a lot on where you are in the curve. I think the other thing too is, and Project Genie is an example of this, there's also a scenario where you sort of see this as a continuum.
33:59So it's almost like you have Nano Banana Pro, you have VO, and then all of a sudden you have a third filler, which is like interactive media that you can step into and real-time media that you can generate. And so the hope is that over time, more people get to experience like different types of experiences across that continuum. Yeah, that's really interesting. I also think we were talking off camera about this, the model, actually to this point of like the breadth of capabilities across all of our models, but also specifically for Genie, the breadth of capabilities and how that translates to the way that the model was trained.
Read the full transcript
34:33And yeah, Jack, you had some spicy comment about the models not actually being trained for any specific use case and them being able to sort of generalize really, really well across the spectrum. Exactly. I think we didn't really know exactly like one specific use case for this, right? We just, we had the sense towards the end of last year, right? That the Genie team had been working together a very research project as i said i mean jini one was a research paper which is different from um and then what happened in parallel was there was this work on and on on doom the doom um game engine which shlomi was part of right which showed really the power of real time into activity but only for one specific um world or game which is dune and maybe shlomi could talk about that And also in parallel, we also saw VO2 in December 2024, which feels like aeons ago in AI.
35:24But I remember seeing that and thinking like this, okay, this is it. Video just kind of works now, right? It's pretty amazing. And obviously since then there'd been further breakthroughs. But at that moment, it was the visual quality is good enough at that point in the kind of frontier that this is really worth going for it. And so we kind of aligned on how exciting this could be, right? and we we put a team together of people from the different groups that had worked on all these different areas right and held with expertise and we just had this belief that combining it would be pretty pretty incredible but it wasn't for one specific downstream use case it was because of the potential of so many use cases and i think the coolest thing is that like we had some in mind right like i think simr is a project we've partnered with for a lot longer and they were part of genie 2 as well they played with genie 2 so they've already had a release using genie 3 What is Simr for public?
36:13Simr is one of our most capable sort of gaming agents. So it can interact in 3D worlds. It's a Gemini powered agent. You can give it a text instruction in a 3D world and it can achieve these different goals. It's very general and it can also learn by self-improvement. So it's quite an exciting path towards our like, you know, getting towards AGI or embodied intelligence or how you want to phrase it. And we really released on this at the end of last year. and they use Genie 3 Worlds to actually explore the capabilities of their agents. If you can imagine that Simmer was trained in a few games, but now you can just type in text with Genie, create a completely new world, maybe even a photorealistic one, and then you can put the agent in there and see it do things.
36:56It's like a naturally pretty coupled thing between the two projects, but I think there's many others that I didn't envisage being so successful, and hopefully you'll hear more about these in the coming weeks, right where the different news cases in the agent space but also in the interactive space right there's completely new news cases that i just never would have believed were possible right but once someone says it it's just like wow that's amazing and we should we should do that i'm excited to see the trends because i feel like vo and n over n have had this like very starkly where there's like a couple of these like very viral trends that happen and then everyone goes and tries to do those and they're always like things i never would have expected um yeah which just crazy yeah i i think there is something a little different here is because when you use for example veo and you generate a video and then you share it with you can like millions of people millions of people can can watch it right and they they don't they just watch something that you create so there is a bit of this creative experience by a single person and then you kind of broadcast that the same goes for images you can share it maybe your family you marry more people and we are also thinking about things where you can throughout the the exploration you can actually change the world around you.
38:05So it opens up a lot of kind of new avenues, new ways that we haven't really considered. And I think that's really interesting to see where it takes. But it's less obvious, I think, than image generation, video generation, where we pretty much replace something or automate something that was there before. And of course, with a new set of new capabilities and limitations. But here it's something that wasn't even possible before. I think there is some differences, of course, a lot of similarities. Yeah, this is one of the ideas that we're super excited about. And you'll see a hint of this in the demos as well is the idea of people building on top of worlds and then you effectively have all these like interesting branches of worlds that you can attribute as well so there's a lot of interesting potential there to explore.
38:42At launch you'll actually be able to download your video and then over time we're exploring different ways of sharing worlds as well so people can sort of build on top of them in interesting ways. Cool I'm excited I want like a world profile so people can see all the all the good ideas that I that I came up with. Well actually what is the what is the slope of progress from a world model standpoint look like do you think it's like just from like obviously we've seen a huge amount of progress um we've seen a huge amount of progress on image generation we've seen it on vo we've seen it on like core gemini as well is it like world models are like generally following the same trajectory and like there's low-hanging fruit everywhere and like we're benefiting from the scale and reasoning and pretty much yeah i would say like image is clearly a more mature than video right and then i'd say video i mean how far behind i don't know but definitely this kind of model is like a frontier beyond that right and and so you'll see that the latest generations of video models are much higher quality than we will get with g3 so we don't expect this to be producing like really beautiful videos at this point um but the real-time interactivity is the constraint that those models don't have right so i think that this will you know probably lag that but then it will be a new experience um and so i think we're sell on a pretty steep site, to be honest.
39:58I think here we have some situation where hardware is always a big constraint, right? It's like the way, and I think it's true for all malls. The trend that we see over time that we make malls more efficient at the same kind of like, basically the same cost if you want. But ultimately, we also need hardware that is more available. And like one of the things that we really hope to see at some point that people can just use their devices to run these models, right? So there is no latency overhead. You can just have it accessible immediately. Obviously, there is a gap in the accessibility of this type of hardware, right?
40:36TPUs, GPUs are very strong ones. But I think ultimately, you know, hardware is probably somewhat constraints because it kind of progresses more slowly for obvious practical reasons. But I think that's where we're going to see some of it. there you'll have a full universe simulation on your phone yeah in genie 5 so that's powerful but this is an interesting point we've discussed this as well and like google's vertical stack ownership advantage there yeah i think this is something that we're also uh you know part of the magic of being in google and a deep mind it's like we can play to that strength as well um and you know really be at the frontier of models but also leveraging the best hardware to support that yeah i feel like a simulation hardware seems like world sim hardware seems very interesting i feel like i'd just be kind of magical to you know it's like it's like almost the portal into another dimension you sort of click on it and you sort of go you know i feel like a lot of novelty um yeah this was this was awesome i'm i'm super excited for users to get their hands i feel like i've seen at least a sliver of the behind the scenes the work that you all have been doing to like actually and and the broader team um to bring this model as project genie now to the world so uh thank you for the hardware guy hopefully people's minds are going to be blown you all are also on the internet so if there's uh complaints do not send them to me you three are accountable to to fix all the problems do not tag me you've got the sweater so i think you're you're wearing this responsible yeah i do have this sweater um yeah but i'm super excited i feel like uh hopefully we'll get to sit down and talk about uh project genie 2 or genie 4 whatever yeah who knows what we'll call it but yeah we definitely want to get that feedback i think and that's something that's a great activity for Air Studio.
42:18Yeah, I'm excited. Thank you all. Thanks for tuning in to this episode of Release Notes, and we'll see you in the next one.
From the publisher
Chapters:
00:00 - Intro and defining world models and RL roots
01:51 - Demo: Goldfish and shark in underwater world
04:59 - Project Genie gallery
06:31 - Physics, remixing, and UI prompts
11:00 - Demo: Nano Banana mascot “Bob”
13:20 - Constraints, generation limits, and infrastructure
17:04 - Trusted testers and robotics future
28:34 - Frontier prompting and universal simulation
29:27 - Cross-Google collaboration
31:16 - Adoption timelines and impact
34:16 - Model generalization and historical context
38:52 - Hardware limits and the slope of progress
