Genie: Generative Interactive Environments with Ashley Edwards - #696

5 Aug 2024 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast Episode #696: Genie: Generative Interactive Environments with Ashley Edwards

Podcast Overview The TWIML AI Podcast explores the transformative impact of machine learning (ML) and artificial intelligence (AI) on businesses and daily life. Hosted by Sam Charrington, this episode features Ashley Edwards from Runway, discussing Genie—an innovative system for creating generative interactive environments for training deep reinforcement learning (RL) agents.

Episode Summary In this episode, Ashley Edwards describes Genie, which enables the creation of "playable" video environments for RL agents in an unsupervised manner. The discussion covers:

  • Motivation and Background: The traditional challenges faced by RL researchers in acquiring diverse training environments for agents.
  • Core Components of Genie:
  • Latent Action Model: Learns actions from videos without needing explicit action data.
  • Video Tokenizer: Encodes video representations to facilitate frame prediction.
  • Dynamics Model: Predicts future frames based on learned representations.
  • Model Architecture and Techniques:
  • The use of spatiotemporal transformers and MaskGIT techniques for efficient prediction and representation.
  • Training Strategies: Including how the components are trained together and the benchmarks utilized for performance evaluation.
  • Practical Implications: How Genie compares to other video generation models and its potential applications in various fields, including education and creative media.

Key Concepts and Discussions

Motivation for Genie

  • The need for unlimited training environments to develop generalist RL agents.
  • The limitations of current methods requiring extensive manual creation of environments.

Core Components

  1. Latent Action Model:
  2. Learns actions in an unsupervised way from video sequences.
  3. Represents actions using a discrete codebook, simplifying interactions.
  1. Video Tokenizer:
  2. Tokenizes video into patches for efficient processing.
  3. Enables better predictions by breaking down the video into manageable parts.
  1. Dynamics Model:
  2. Predicts future frames based on the context of previous frames and actions.
  3. Integrates temporal and spatial information to maintain consistency in predictions.

Technical Insights

  • Training the Models:
  • The tokenizer is trained first, followed by the simultaneous training of the latent action and dynamics models.
  • Employs masked token prediction to enhance the model's ability to infer missing information.
  • Performance Evaluation:
  • Uses metrics like FVD (Frechet Video Distance) to assess video quality.
  • Compared against existing architectures, revealing strengths and weaknesses in computational efficiency.

Broader Implications and Future Directions

  • Applications in Education:
  • Interest from educators to use Genie as an interactive teaching tool.
  • Creative Media:
  • The potential for artists and creators to leverage Genie for generating interactive content.
  • Future Improvements:
  • Enhancing inference speed to enable real-time interactions.
  • Exploring diffusion models to enhance generation quality.

Conclusion Ashley Edwards provides a compelling look into Genie and its capabilities for creating interactive environments for RL agents. The episode highlights both the technical components and the broader implications of such technology in various fields.

Additional Resources

  • Complete show notes can be found at [TWIML AI Podcast Episode 696](https://twimlai.com/go/696).
  • For those interested, further exploration into Genie and related topics can lead to valuable insights into the future of interactive media and reinforcement learning technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Typically, as reinforcement learning researchers, we often get a little bit annoyed with how we have to bring in more and more environments. Like typically, if you want to train an agent, you're like, OK, well, maybe I can bring this in or bring this in. But this doesn't really scale. And I would say nowadays, you know, we're interested in generalist agents. And if we want generalist reinforcement learning agents, then like where do the environments come from? So one of the motivations behind that work was how can we basically learn an unlimited source of environments for training agents? It gives you the sort of interactability that you would have in an RL environment without actually having to place the agent within the environment itself.

0:51All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Ashley Edwards. Ashley is currently a member of technical staff at Runway ML, which she joined just a few weeks ago after several years at Google DeepMind. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Ashley, welcome to the podcast. Hi, thank you. It's nice to be here. I'm excited for our conversation. We are going to dig into your work in video generation and the genie environment in particular. But generally, I'm really excited to have you on the show.

1:34We've been following one another and interacting on Twitter for quite a while. And I'm kind of amazed that we've not spoken yet on the podcast. So definitely welcome. Yes, thanks for having me. Absolutely, absolutely. Let's get started by having you share a little bit about your background. Yeah. So as you mentioned, I'm currently at Runway. I was at DeepMind before that. Before that, I was at Uber. But in general, my research has been in doing reinforcement learning from videos. And more recently, I've sort of shifted into how we can actually do video generation itself. Give us an overview of the types of projects that you've been focused on.

2:13I mean, I would say recently, probably Jeannie. Yeah. But I think in general, I guess what people sort of know me for is doing imitation from videos. So often when you're doing imitation learning, you typically need a combination of states and actions that tell you what to do. But often it's kind of difficult to come up with those demonstrations. It requires a lot of effort on the demonstrator. And so my research has been about how we can leverage videos to train agents. So how we can learn actions, policies, rewards for training reinforcement learning agents. And that is one of the areas where Genie is unique because you're learning all of these actions in a totally unsupervised way.

3:03Yeah, yeah, exactly. Exactly. So why don't you share, you know, broadly, you know, Jeannie and kind of the motivation for that work? So as I mentioned, my background is in reinforcement learning. I've been working in that area for quite some time. I think typically as like reinforcement learning researchers, we often get a little bit annoyed with how we have to bring in more and more environments. Like typically, if you want to train an agent, you're like, OK, well, maybe I can bring this in or bring this in. but this doesn't really scale. And I would say nowadays, you know, we're interested in generalist agents.

3:37And if we want generalist reinforcement learning agents, then like, where do the environments come from? So one of the motivations behind that work was how can we basically learn an unlimited source of environments or training agents so we can at least remove that sort of aspect of reinforcement learning where we're having to come up with the environments, unless you're doing robotics, I guess, where you have the environment itself. But yeah, I think in general, it gives you the sort of interactability that you would have in an RL environment without actually having to place the agent within the environment itself.

4:09And so what's the data source that you're learning these environments from? In the paper, we basically use two different kinds of sources of data. The main one would be videos of 2D platformer games. In reinforcement learning, we love training over games. So you kind of look into that sort of environment. so a very diverse data set with lots of different actions and different agents moving around and that sort of thing and then the other one we have is like a robotics data set coming from i think it was like rt1 qt off that we we extracted the observations without actions from that environment and trained as well just to show that we can also use this approach on real world environments as well and talk a little bit about the end result that you've created with Genie?

4:58What does it allow you to do? I mean, I think the thing that really excites me the most is that you can learn this sort of world model without using any actions. You can learn it completely unsupervised from videos, which is quite general. You don't need many inductive biases here. You just give it a bunch of videos and it learns a world model that you can interact with. And so what we found was, in particular, when we were training over the 2D platformer games, that we can take text generated images or sketches even or real world photos, and we can sort of just step into those environments as if they were like a real game, like we could just basically play it like it was like a real world model.

5:39Again, even though we didn't use actions at all here, or learn, or I guess ground truth actions. Yeah. What makes it a world model? Is it, you know, what is that really speaking to? Is it a degree of consistency between frames or how do you kind of capture the essence of that? Yeah, I would say, I mean, when I think of the world model, I really think of something that you can step into and interact with. So if you I mean, if you remember, I guess, like Schmidt Hoover and David Haw's paper from a while ago on, I think they did Doom. They had this world model paper and then you have like you've learned a simulator of Doom.

6:16Of course, this is just a single environment. In our case, we're looking for learning a more general world model, like basically a world model of 2D platformer games that, again, enables you to just like interact and sort of through your actions, you can change the environment essentially. And so what are the major components of Genie? So we have, I guess, three big components. One is the latent action model where we're learning the actions from videos. Another is the dynamics model, which we're actually using to generate the future. And then the other one is our video tokenizer. So this is how we're actually learning the representations of our videos, which we can then use for predicting with our dynamics model.

7:04Okay, let's start with that latent action model. it's super interesting that you're able to ingest all these videos the videos have I guess generally I guess sometimes you see on YouTube videos like the you know Twitch streamers and and the like will have like these labels that pop up that say like either what direction they've pressed on their joystick or you know keypad or whatever our keyboard for that matter but But I didn't get the impression that most of your videos had that and you were somehow leveraging that kind of information. It was more just playthrough videos. Is that right? Yeah, yeah, exactly.

7:43I mean, I'm sure, I mean, that sort of information could have been in there. But I would say it wasn't something that we were explicitly trying to extract. Mainly what we were trying to figure out is, can you learn a representation of what's going to change between the scenes such that you can predict the next frame? And this is basically what our latent actions are. Are there a set of assumptions that you've made about actions, meaning like, you know, that there's going to be, you know, an up, down, left, right kind of thing, and that's your action vocabulary? Or is it, you know, more open-ended than that and kind of fully learned end to end?

8:22I think one thing I would say is we try to restrict the action space so that we can actually interact with it. So we're basically representing our actions using a discrete codebook, using vector quantization. So one, we assumed that we had a small set of actions, which is definitely not the case, but it just made it easier for us to interact with the models. The other thing that we sort of assumed was that there was a common action space across the different environments that we were training on. And this was going to allow us to learn this sort of general representation. because we're learning over all of these environments.

8:58We wanted to learn what was common. So there must be different agents that are going to move to the left or move to the right or jump. And we wanted to learn the representation that captured that. Are these three components, the latent action model, the video tokenizer, the dynamics model, are they each trained independently and then come together in the final system or are they trained in concert with one another? Yeah, yeah, good question. So we train the tokenizer first And then we train the latent action model and the dynamics model at the same time, although we're not using that dynamics model for training the latent action.

9:33So we could actually also train the latent action model separately and then use it for the dynamics model, but we just train those two at the same time. I saw a reference in the paper to spatiotemporal transformers. Tell us what that is and how it plays into Jeannie. Essentially, what we need to do is we need a representation of our images. So what we do is we represent them through patches, again, using vector quantization. And so each image is represented by a patch, essentially. So you have a sequence of images, which is your video. If we were going to try to plug in each of those patches into our transformer and basically have attention over all of those different patches, it would be very computationally intractable.

10:20And so what we end up doing is instead also we incorporate time, that's a temporal part. So the spatial part is the image patches and the temporal part is time. But instead of having each patch attend to each patch in the video, what we can do is instead take a single patch and attend to it across time. So instead of having this like quadratic blow up and resources, what we instead have is like this linear scaling across time. That makes me think of a couple of conversations that I've had recently. One, very recently, I think we just published this interview with Albert Gu talking about state space models.

10:57And one of the main things that we talked about was how tokenizing for video and other kind of more native modalities, non-language modalities, like introduces artifacts that somehow make training models more complex and make the use of attention more inefficient. and i'm just curious you when you think about models that you could have used or might use in a future like how do you think about the different roles of uh alternatives to the spatiotemporal transformer yeah yeah that's a great question i mean i think so the context window of genie i think was maybe like 16 frames so if i think if we wanted to like roll out or once we start looking into more longer generations, I think it might make sense to actually look into these more efficient approaches like we do in these state space models.

11:56One thing that I find like, it's interesting that we sort of have to put our images into these patches just so that we can have tokens for the transformer to predict. It feels very, I mean, it works. I think it works well. But I do think that there's probably other ways that we can represent images. I mean, also you have diffusion, which you can do over latent space rather than having to use to predict over patches necessarily. So I think in general, when we when we built this model, we were looking at, OK, this is a thing that's probably going to work. This is a thing that's probably going to work.

12:27This is, you know, and we put those things together. But now I think there's probably different ways that we can improve upon each of the different components that we had in the model in terms of making it more efficient. And so how did you structure the training of the latent action model so that those actions were kind of pulled out of the training data? Let's start with just imagine we have two frames. So say you have two frames. What you want to do is learn a representation, but you're going to compress such that if you take that representation and your previous frame, you can predict the next frame.

13:05So the latent action model gets to see both the past and the future. But we compress it in such a way that you can't actually extract the information about the state itself. You really just have to try to extract the change that occurs between states so that you can use it for predicting the future. So the way we compress it is we essentially, we place it through a bottleneck and then we put it through this BQ module, which says we have like 10 codes. And we're trying to basically make our representations be close to those codes that we're also learning at the same time. I know those 10 codes, the size of your action space, you can represent 10 possible actions?

13:45Yeah, exactly. Okay. One other thing that I forgot to mention is that, well, I talked about it in terms of two frames, but you can actually use all of your previous frames up to that point and then use that as your context for your latent action model. But you only typically, we only see a single next frame for it. And what are some examples of actions beyond up, down, left, right, jump that are common enough across games that, you know, that will fill out the rest of this 10? Yeah, I mean, I think one thing that I that we that we saw was that it would learn different, I guess, velocities for the character.

14:26So it wouldn't just be like move left. It'll be like move left quickly and move left, you know, with this amount of speed. also we there was like explosions and score change uh lighting you know yeah some of them actually learned to like change the score which was quite cool when it's changing the score is a model essentially just changing pixels or is that something else yeah okay yeah i think well when i say score probably more like health meter like if if something gets hit then you see like it's health go down essentially um which was yeah so that that's it's all in pixel space so yeah yeah Yeah.

15:02Yeah. So you train this model, you you've got this latent action model and then you are generating a video. Is there anything that, or you're generating an environment? Is there anything that maps the action model to that particular environment? Or is it, you know, always, you know, code four is up code three is, you know, fast left, you know, whatever. Yeah, great question. Surprisingly, it does learn the sort of consistent mapping across different inputs. So yeah, latent action zero might always be moved to the left, right might always be latent action four. We didn't do anything to actually ground the action space in this, but I think it's just through the fact that the model needs to learn this representation in order to predict the future.

15:53And it's easier for it to learn the same thing across different environments. Did you have any benchmarks that you're able to compare this latent action model to? Or how do you think about and measure its performance? The thing with interactive models is it's really hard to measure things like controllability, but you can measure things like video quality. So we're using metrics like FVD, for example, to measure how good the videos are. We looked at a couple of different representations compared to using a spatial temporal transformer. So one of them was like the C-Vivit architecture, which was used in Finaki.

16:39And in general, we found that because they don't do this like spatial temporal trick where you only like attend to one token over time. They actually have to attend over multiple frames. So it was quite compute intensive. but that was like i think the main method we like architecture i think we compared to in general we were just like trying out different things in terms of how we scale um our back size and model size and how that had an impact on things the latent action model as well i think we had some experiments i'm probably gonna have to look back at the paper to remember exactly what they were In terms of the video tokenizer and the approach to patching, how custom does that need to be for, you know, a particular problem?

17:31Like is video, like I guess maybe another way to ask the question is like, why is that a major component? Okay, yeah. So I would say one thing is that it enables us to predict tokens with the Dynamics model. and because we were doing video generation if we are predicting pixels themselves then we would end up getting blurry generations so we probably wouldn't want to use a convolutional neural network in this case but because we put it into patches we can predict like a discrete patch so we use a vector quantization of that so we can predict a discrete patch which we can then plug in the in the decoder to reconstruct the image rather than having to predict the image itself so that let us get sharp predictions from the model.

18:19The other thing is that the spatial temporal component of it was important because what we found was that if we were tokenizing only in space, then it might not represent the sorts of changes that you need to represent when you're predicting the next frame. Like if we're only using this for image generation, then we might not need, I mean, obviously we wouldn't need the temporal part because you don't have it. But because we're generating videos, it's useful to actually have an architecture that can capture changes across time. So I think those were the main reasons we looked into using this sort of tokenization scheme.

18:58I would say, so the spatial temporal model is what we ended up using in every component. So the video tokenizer actually wasn't, we didn't use that in the latent action model, We just use it in the decoder. But the spatial temporal architecture we used in the video tokenizer and the latent action model and the dynamics model. With that in mind, like, do you think that it's possible or it would be possible to have a single end to end model since you've used the same underlying model for all three of these components? Do you see it as being possible to kind of collapse this down into a single end to end train model?

19:39at some point in the future? And, you know, what do you think the biggest challenges in doing that? Or is it like architecturally unlikely because of, you know, whatever reasons? It's a good question because it's something that we wanted to get working from throughout the project. Particularly, I think the thing that was most like frustrating, or I mean, it was annoying, I guess, was that our latent action model has its own decoder. So it takes state, next state, goes into action model, predicts the next state. That decoder is different from our dynamics model. I was going to say, that sounds a lot like what the dynamics model needs to do.

20:17The dynamics model, yes, it's taking the state's latent action, predicting the next state. We can never get that model to train our latent actions. It was very difficult. I think, I mean, one thing that we're doing is we're training it in a different way. So that model, first of all, takes in tokens. we found that training directly on pixels for the latent action model performed better than training on tokens for it that model outputs tokens and we found that predicting pixels perform like helped the latent action model perform better so it might just be a difference in the kinds of losses that we had like the dynamics model uses a cross entropy loss whereas as the latent action model uses mean squared error laws.

21:03But I think it could maybe be possible to train like a tokenizer at the same time as the latent action model. But I mean, I'm sure there's a way to make it end to end. We just, we found the way that that worked. I think the next step maybe would be to make it more elegant. So, you know, I wish that it was end to end, but it's not. Got it, got it. And can you dig into a little bit more of the dynamics model? And you said the dynamics model takes as input patches and outputs patches or pixels? It outputs patches, but it's predicting the tokens. So basically, the way our video tokenizer works is you get patches.

21:50We place that through a VQ module, which gives it a discrete code for each of those kinds of patches. And then what we can do for our dynamics model is now we're going to predict patches, but they're represented through tokens. And so you can use a cross entropy loss over those different patches. I guess one thing is that we're doing masked token prediction rather than predicting all the tokens at the same time, at least during training. During inference time, of course, you have to predict all of the tokens. Yeah. uh elaborate on that what do you mean by uh doing mass token prediction yeah so essentially what we do is we can uh mask some of our uh previous token so basically give it a value of zero so that you don't have to predict it um so it gets essentially gets to see is this in in time for a single kind of coordinate or is this like an infilling kind of thing in a where you're you're masking in 2d space right yeah so essentially it gets a sequence of tokens which represents its patches um does that part make sense so these are all of your frames So a sequence of tokens representing patches at multiple time steps or help me tune up my vocabulary here.

23:23So or the way I'm articulating this. So you've got, you know, you've got multiple patches in a two dimensional representation of a frame. Right. Yes. Yes. And then you've got multiple frames in time. And so when you get a sequence of patches, is it like the same patch, meaning representing the same part of an image across time? Or is it? I see. Right. So yeah, okay. So yeah, so a single image, let's say you have a patch. How would you say all that? Like, what's the right way to talk about, like, where the patch is situated in the image versus? Yeah, that's a good question. Let's think about this.

24:11Okay. So let's say you have an image. This is my image, by the way. I keep doing this. This has always been an image in my head. Let's say you have, like, a 16 by 16 patch. So what you're going to try to do is you're going to say, okay, I'm going to go through this image. And each of these parts are going to correspond to a patch. you can then represent that as a sequence. So that particular image is represented by a sequence of patches. So I mentioned we have the spatial part and the temporal part. So the spatial part gets to see the spatial transformer is just attending to across patches in a single image.

24:50And the temporal part attends to a particular patch across time, but it's only attending to that single patch, that same patch over time. Yeah. That's what I was asking. Yeah. Yeah. Yeah, it definitely gets confusing. And I might've even explained some parts of it incorrectly prior to this part. So the main idea is that you can mask out some of those tokens. Okay. So you don't get to see all of the context. You maybe like zero out some of those tokens. Okay. And I think this is where my question came in. are you masking out like is it a random masking and it could be um you know different patches within an image you know from one wait okay so let's take a step back yeah it's random masking yeah it is random masking and it is um in the spatial dimension or temporal or both that we're talking about i would say yeah it's just well it's random masking in a particular uh in a vector which is spatial and then you have multiple of these uh over time is that the right yeah yeah exactly yeah okay um so you get mass tokens and then you're going to predict the next tokens um and we're using a technique known as mask it during inference time inference time um So rather than predicting a single token one by one, you predict them in parallel.

26:25So you have my image again. You unmask some tokens. You keep the ones that you're most confident about. And then you condition your model again on those. And then you predict the next tokens, essentially. So this is mask it as compared to, I guess, what most people are using these days is the fusion. And so you're predicting... tokens based on the masked sequence of patches and that is then used to generate the kind of the t plus one frame of the video right yeah exactly so you can take those tokens plug them into your decoder from the tokenizer that we were training before um that will generate your image so we're again we're predicting the tokens not the not the image itself and then we can reconstruct it i'm curious about the like the computational complexity of all this and i guess the thought that preceded that in my head was like did you publish code for this like can someone pull up a notebook and just do it and then the question was like what does that even mean like how much compute was required to do this yeah uh so no code however there is a there is a um a reproducibility uh section in the paper uh about how people someone could reproduce this on like um like a mid-range tpu over coin run data um so coin run is like a reinforcement learning environment uh so we give instructions for how to yeah reproduce it in that setting we haven't um open source anything, but there have been people who have released code.

28:11So yeah, I encourage folks to check that out. How playable are these interactive environments? Like, do you, you mentioned something about a 16 frame context that's relative to training. Once you've got the environment, you could play that, you know, as long as you would like, because you're just generating a next frame correct yeah yeah that's that's true and then i'm wondering how um you just kind of the noisiness of like transformers and prediction and all this like how that manifests in in videos and playability like is it you could you can play it but you really wouldn't want to like is you know how how interesting is it i think it's interesting in the fact that once you see it you can see that you can change the environment for like a couple of steps then i mean you yeah i i would say that you would never at least with this this model you would never see like another character coming into the environment or um yeah interesting things don't really happen but i think it's still cool like we took my and And that was not to take it all away from what you've done.

29:35It was really trying to get at like, you know, how well a job, the model or the set of models do kind of maintaining consistency in the generation from, you know, one frame to the next over time. And if that results in something that kind of feels playable or, you know, do random things happen that, you know, make it, you know, less playable and reflect a degree of uncontrollability in the generations. Yeah, I would say like, I mean, the model, the actions are able to remain consistent. It does tend to, I think, sort of maybe overfit as you press the same action. So for example, if you've been taking action right, it kind of over time, I would say the actions might end up all going right.

30:31But we found that there's a way there's like an action that would say, okay, stop doing that and like reset, which is kind of weird. But there's like an action you can take that would let you start moving left again. I guess it's like maybe it's like a stay still and then change direction style thing. um but the the thing is so like if you're moving right for a long time you basically generally just see a pattern of the background moving over and over and over again like you you keep seeing the same things you it's like i don't think i've ever seen anything um in like our outer distribution environments like i don't think i've ever seen like anything interesting pop up out of nowhere essentially.

31:11I don't think I understand that. Why does the, it sounded like you were saying the, the semantics of an action changes over time or one of these action codes. Is that what you're saying? And why would that happen? I thought we were saying earlier that they're, they're essentially, you know, static post-training. Yeah. So the actions like given the same or sorry given different prompt frames um taking latent action zero typically would move you to the left or right or whatever across different initial frames over time if you if you've played the environment and you've been taking like the same action like certain things that might just have to do with like velocity like it kind of is just like it's learned the model of how the world's going to work so if you're going up and you're continuing to go up It's very difficult, for example, to start going down or like take the action that's going to go down because the model has predicted.

Read the full transcript

32:06And it's, I would say, model of maybe the physics of the world, but you can't suddenly start going down. So it might be the same thing. Like if you've been pressing the right action, like it doesn't matter. Like the latent action predictions get completely ignored because it knows that it's very rare that you would suddenly maybe start moving left. So it might just be, it's two things. One, it could be just ignoring because of the physics of the world. One, it could be ignoring because of the policies that it's trained with. Maybe people typically don't suddenly go left when they've been going right for a while.

32:41That still doesn't make sense to me. And so there's something that is in there that I don't understand. So my kind of mental framework of this is that you've trained these components. you've got this latent action model that you've trained that kind of pulls out the actions then you've got this dynamics model that um predicts frames and you know once you're at playing you know it is you know you're giving it an action code to you're giving a previous frame or some set of frames and action code to oh i guess it's the previous frames if you and maybe another way to think about this is that if you the dynamics model is maybe balancing the the information it gets from the latent action and from the previous frames and if you have too many previous frames that are saying one thing it kind of ignores the latent action model yeah yeah yeah exactly it's It's the spatial temporal part of it again.

33:51It's paying attention to the previous time steps, actions and frames that it's seen. Got it. So then in theory, one could tweak the dynamics model so that it pays more attention to the latent actions relative to the prior frames or give some additional weight to them, either in some reward function or something else. And that might fix that problem. Yeah. Yeah, that's a good point. Or even just shortening your context window and seeing what happens. Oh, yeah, yeah. But it probably wouldn't be the most interesting generations, but then it wouldn't have the past to pay attention to. Mm-hmm. Yeah.

34:33And so did you need to do anything different or special to accommodate the variety of sources that you're able to accommodate? So you mentioned you can do, you can start with a image from like a text to image generation. You can start with a photo. You can start with a sketch. Did you have to do anything from data collection perspective or from a training perspective? Or did it all just kind of work and it was cool to see? Yeah, it was, it just kind of worked. It was, it was actually really crazy. So we used to have this thing called the genie bottle challenge, essentially, where we would all get together and play our environments and see if we can reach the genie lamp that was generated from a text generated image um so basically we're like trying to play the model in the beginning was really bad um eventually we were able to reach the reach the genie lamp uh and so once we sort of solved that challenge we started thinking about what other sources of data would just work um and And so at some point we asked Jeff Klune's children, Caspian and Seneca, to actually make these sketches that we have in the paper.

35:45Somebody played in those environments. And I think Richie also drew a really cool sketch. And you saw that you can climb up the ladder. And so, yeah, I think we did that. And then at some point someone was like, well, what if we plugged in real photos? I was like, why would that work? We trained over platformer games. um but then we were able to control jack's dog doris so for whatever reason yeah i mean it's very you know she looks like a character i guess in the middle of a screen so we were able to like move her around and stuff but we didn't do anything different we just plugged them in um although i would say actually we did find that like adding um uh sidebars to some of the images like you know what you would see in like a video like adding these black sidebars to some of them kind of helped it figure out what to control um so yeah that was interesting yeah oh one other thing i forgot to mention is um it can sometimes be difficult to figure out what you want to control in the environment um like maybe you want to control the clouds in the sky rather than the dog but if you take an action that controls the right thing then all of the actions start controlling the right thing um which is interesting because they immediately or sorry initially they might be controlling different parts of the scene but as soon as you take an action that's controlling the thing you want to control all of a sudden they start controlling it too um if that makes sense uh it does i think going back to our previous conversation about the relationship between the actions and the prior frames to the dynamics model so if you somehow choose an action that happens to result in controlling a cloud, then the dynamics model will see that in the kind of the frame history and kind of key in on that.

37:31Is that the idea? Yeah, that's a good point. Yeah, yeah, exactly. Yeah. Actually, if you even, if you plug in a video, it's much easier to control given part of that. Like if you've given it some video, you can figure out what to control if you do it. So yes, that's basically it. What do you mean plug in a video? Is that an input mode? In In addition to static images, you can also give it some frames of, I guess, I don't see why not, if you're just kind of feeding into the dynamics model. Yeah, exactly. I mean, essentially what you have to do is infer the actions still. So you take a frame, you've inferred all the actions, you roll that out into the model.

38:06And then you can predict the next frame given that context. Yeah, I don't think we actually put any results in the paper about this, but it was something that we saw. Yeah. Mm-hmm. And I think this came out around roughly or shortly after Sora. I forget the specific timing relationship, but there were kind of some obvious comparisons between the two just from a video generation perspective. but the last comment you made about kind of providing an input a video as input as opposed to you know static image I think maybe accentuates the comparison and suggests that you know this could be like a playable you know so I don't know why that's any different than what and what we've been talking about thus far, but I think it kind of maybe suggests that, you know, maybe the distinction is like, you know, with Sora and, you know, Gen 3 for Runway and these other video generation models, like, you know, they're beyond platformer games and like the idea that like you can do those things and also like have control over, you know, actors in these videos is kind of really interesting.

39:30Yeah, no, I agree. I think it would be cool. I mean, yeah, you can take any of these video generation techniques and then plug them into the model, I would say. I mean, if you train on that sort of data, it wouldn't work on the 2D platformer one. I mean, we controlled the dog, but I doubt it would work on 3D. But yeah, I think it's a good point. We can sort of build upon these things and I guess leverage them to get the kind of controllability that you wouldn't necessarily get initially from the model. But now you have like a nice generated scenes that you can then take over. Interesting. So beyond the idea that now you've got these interactive environments for reinforcement learning researchers, like what are the broader implications of Genie, do you think?

40:18When we started, we were definitely interested in using it for reinforcement learning. I think over time, we realized it was fun to actually interact with it. I mean, this is why we started looking into sketches and people's drawings. I think using it as like a tool in classrooms. Like when we put out the paper, we got people reached out to us from like places I wouldn't have expected. So like teachers interested in using this sort of model for their students, people who add data like video only data and they were trying to learn simulations of that data um that they can use maybe also as teaching tools um for people in the industry um just as something fun for creatives to work with i think there there's many different kinds of um yeah areas that it could be used for and what do you think are the biggest gaps in terms of you know if you wanted to use this to make actually a playable games like do you see that as how far do you see a is like this the right thing to do for that or is that you know or not really or um you know and be like how far do you think that is and see like what are the biggest you know holes or gaps or how would you approach that i mean i think one of them is imprint speed like you can't really interact with these in real time.

41:40I don't, I, I mean, we, the videos that we generated were, you know, not in real time. They were just, you know, we showed a video of us interacting with them, but it was, I think probably one frame per second. So that's one, one thing trying, trying to speed up inference. But also, yeah, I think the question of whether, you know, is this something you would want to do is, is an important one. And I think making sure that you're having conversations with people who would also be impacted by this technology early on would also be important um but yeah it's a good question i mean i don't think i mean genie was i wouldn't play it as a game at this point in time um but it's fun i would say it's a very different kind of game though so it's not something that's going to take over like i don't know whatever your favorite video game is um but i think it's like a new sort of it could be a new form of of i guess interactive um media that people can interact or interactive media you can interact with yeah yeah but yeah it's just it's just like a new form of i would say it could be art it could be i don't know uh some technology but it's different i think it's not necessarily or doesn't have to be the same as a game would be yeah i think referring to it as interactive kind of glossed over the inference speed issue for me and i was thinking someone was like pressing up down left right and controlling this scene in real time and i mean yeah it depends on how much time you've got yeah it takes some time for sure yeah and so were you doing it uh were you doing it in batch where meaning you just kind of picked a sequence of actions and fed that to the model and let it generate a bunch of frames or you know did you run experiments where you were sitting there and pressing up and pressing were up and waiting for it to roll up That was how we generated most of them.

43:31Unfortunately, yeah, like we have this video in the paper, or sorry, on the website that shows like a zooming out on all the environments you can generate with Sheena. And that was all like real gameplay of people like stepping into the environment and interacting. Because random behavior doesn't look interesting. Like if you were just to plug in like a random sequence of actions, you would just see like things jittering around, like it doesn't look that interesting. um so we wanted to have things that looked like real gameplay um the alternative would be to maybe generate like a game tree over time but that would take a really long time um to do uh so yeah no we did it ourselves how does this tie into kind of your other areas of interest around video generation and kind of where do you see this going what are you excited to poke around with next Yeah.

44:21Yeah. Great question. I think it's been interesting to see, I guess, my sort of shift from doing reinforcement learning to working on video generation. Although I would say I've always done reinforcement learning from videos, actually. So it's not that crazy of a shift. But I think it's really cool to start looking into the different kinds of methods that people have been using. Like I mentioned, this approach was mask it. uh now I'm getting more interested in diffusion which you know it's like it's a completely different area you know from what I'm used to but it's it's been fun to learn um about that space um I think also just a creative aspect of it and like I mean I think it's nice uh like I you mentioned I just joined Runaway I think it's quite cool to actually have like people who can interact with the models that you're you're working on and have that sort of immediate impact I think it's quite exciting to be in that space now.

45:18So yeah, those are the things. Awesome. Should we expect Runway to kind of continue working in this direction? I can't say what we can expect. I only joined a couple of weeks ago.

45:34You know, we did just, well, you know, they released it. Like the week that I joined, they released this like Gen 3 model, which was, yeah, it was really, really, really cool to see. Very fun time to join. So I encourage people to try that one out. Awesome. Awesome. Well, Ashley, thanks so much for jumping on and sharing a bit about Genie and what you've been working on. Thank you as well. It's been really fun to chat. Absolutely. Thank you.

From the publisher

Today, we're joined by Ashley Edwards, a member of technical staff at Runway, to discuss Genie: Generative Interactive Environments, a system for creating ‘playable’ video environments for training deep reinforcement learning (RL) agents at scale in a completely unsupervised manner. We explore the motivations behind Genie, the challenges of data acquisition for RL, and Genie’s capability to learn world models from videos without explicit action data, enabling seamless interaction and frame prediction. Ashley walks us through Genie’s core components—the latent action model, video tokenizer, and dynamics model—and explains how these elements collaborate to predict future frames in video sequences. We discuss the model architecture, training strategies, benchmarks used, as well as the application of spatiotemporal transformers and the MaskGIT techniques used for efficient token prediction and representation. Finally, we touched on Genie’s practical implications, its comparison to other video generation models like “Sora,” and potential future directions in video generation and diffusion models.

The complete show notes for this episode can be found at https://twimlai.com/go/696.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Genie: Generative Interactive Environments with Ashley Edwards - #696The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 47 min
Listen in VO