In short
The TWIML AI Podcast Episode #676: Video as a Universal Interface for AI Reasoning with Sherry Yang
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington engages with Sherry Yang, a senior research scientist at Google DeepMind and a Ph.D. candidate at UC Berkeley. They discuss Sherry's paper titled "Video as the New Language for Real-World Decision Making," focusing on the implications of generative video models as a new medium for AI reasoning and decision-making.
Key Topics Discussed
The Role of Video in AI
- Video as a Unified Representation: Sherry draws an analogy between video and natural language, suggesting that video can serve as a unified format for conveying information, similar to how text operates for language models.
- Generative Video Models: These models can perform various real-world tasks, similar to how language models can address a range of language-based tasks.
Historical Context
- Reinforcement Learning Background: Sherry's initial focus was on reinforcement learning and its limitations when applied to real-world scenarios lacking effective simulators.
- The Shift to Video: Recognizing that video can mimic the physical world more accurately than previous simulation methods led to this exploration of video models.
Challenges in Video Modeling
- Data Coverage: The vastness of potential scenarios that can be represented in video makes it challenging to obtain comprehensive training data.
- Lack of Labels: Unlike text, where later words can serve as labels for earlier ones, videos require precise labels to guide generative processes.
- Architecture Diversity: Various video generation models (e.g., diffusion, autoregressive) lead to a fragmented understanding of the best approach.
Advantages of Video
- Richness of Information: Video contains detailed physical dynamics, which makes it more effective for teaching tasks (e.g., cooking or repairing) compared to text.
- Visual Reasoning Capability: Video can be used for complex reasoning tasks, such as understanding spatial relationships and physical interactions.
Video vs. Text
- Direct Interaction: Video allows for a direct and richer interaction model, bypassing the limitations of translating tasks from text to video.
- Unified Task Interface: Sherry emphasizes that video could serve as a single task interface similar to language generation, encompassing diverse applications from robotics to education.
Future Applications
- Interactive Simulators: Sherry demonstrates an interactive demo (UniSim) that showcases how generative video models can simulate real-world environments, allowing users to perform various tasks.
- Potential in Diverse Fields: This technology may find applications in areas such as education, robotics, and scientific simulations.
Open Questions and Future Directions
- Improving Generalization: Sherry notes that improving the fidelity and realism of generated videos is essential for practical applications.
- Combining Knowledge Across Fields: She discusses the potential for interdisciplinary collaboration to enhance video generation models by integrating knowledge from graphics, machine learning, and specific domain expertise.
Key Takeaways
- Video has the potential to serve as a universal medium for AI reasoning, similar to text in language.
- Generative video models can facilitate complex decision-making and real-world task execution.
- There are significant technical challenges that need to be addressed for practical implementations of video-based AI.
- Interdisciplinary collaboration may enhance the performance and application of video generation technologies.
Demo Overview The episode concludes with a demonstration of UniSim, an interactive video generation model where users can simulate real-world tasks by choosing specific actions in a pre-defined scene. This showcases the capabilities of generative video models in creating dynamic, interactive environments.
---
For more information, visit the complete show notes at [twimlai.com/go/676](https://twimlai.com/go/676).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:07Hey, what's up, everyone? Thanks for joining us for today's episode of the Twinwall AI Podcast. You know me. I am, of course, your host, Sam Charrington. And today I'm joined by Sherry Yang. Sherry is a senior research scientist at Google DeepMind, as well as a PhD candidate at UC Berkeley, where she is advised by longtime friend of the show, Peter Reveal, who I interviewed back in 2021 and 2017 before that. If you're watching us on YouTube, be sure to take a moment to hit that subscribe button. And if you're listening in on Apple Podcasts or Spotify, It is the follow button that you're looking for to make sure that you don't miss any of the great episodes that we've got in store for you.
0:44Sherry, thanks so much for joining me and welcome to the podcast. Thanks, Sam. Very happy to be here. I'm looking forward to digging into our conversation. We'll be talking about your recent paper titled Video as the New Language for Real-World Decision Making. Sometimes I ask guests for their controversial opinions. And I think in this case, it's pretty clear that yours would be that video is underappreciated and that its chat GPT moments should arrive very quickly. Is that fair? Yes, that's correct. I think gradually more and more so people started to realize maybe after Sora was released, but this realization could have really came much earlier.
1:24How did you get started in video? So I guess a little bit of background. When I started in research, I was actually doing a lot of reinforcement learning. So this is about, you know, the technology that empowered AlphaGo, that actually this is what enabled AlphaGo to search beyond human capabilities and eventually outperform humans. But then for AlphaGo to work, there is actually something very special about it, which is the simulator of the game that AlphaGo uses to imagine into the future is the ground truth chess or goal simulator. Back then in 2016, people were like, well, AGI moment has arrived, but nothing happened afterwards, right?
1:59Because oftentimes in real world, there is no simulators. And the simulators that people use for robotics or dynamical simulation is actually quite different from how real robots look like in practice. So as a result, you know, any algorithms that people develop in simulation wouldn't work in the real world. So what's so special about video generation models and large language models is that now with these internet scale data sets, we see that these large foundation models, video models and language models being turned into agents and environments and world models and simulators so that we can actually run these algorithms for searching for better decisions in simulation.
2:37But then because how simulators look so realistic to the real world, then any of the decisions we can find can also apply or transfer to the real world. And we can achieve superhuman performance potentially in a lot of real world scenarios. This is why I'm very excited about this direction. And maybe the question is why specifically for videos? So I guess that goes into really the motivation of this paper, which is kind of the recognition that video is, in fact, a unified data format for a lot of different information. So, you know, for language models, one of the reasons that they work so well is because text is a unified format, right?
3:14Because Wikipedia articles are in text, code is in text, math formulas are in text. So when people have a unified representation of information, they can train a single model using a unified objective, which is just language generation. Translation tasks are language generation. Summarization tasks are language generation. And solving math problems are also language generation. So with a unified representation of information and a unified task interface, this is why people can use language models to solve any tasks using kind of Internet scale knowledge. So I think this realization wasn't so clear for people working on video generation.
3:50Most of the time, you know, there was Pika Labs, there was stable diffusion, there was runway. A lot of people were looking into videos, but it was mostly for entertainment purposes to generate video according to user instruction, but usually some cute panda running on the street or something like that. So back in 2022, in September, a few of us from DeepMind and also from Berkeley started to look into, can we really use video generation models to solve real world tasks? Because video is also a unified representation of information, right? So I don't know how long I should go down along this path before you ask another question.
4:25But if we think about it, because I'm just like so excited about all of this. No, it's clear. I think what I heard there and got from the paper was really your excitement about video and almost like a sense of awe. Like everybody, you're all getting so excited about what we got out of text. Like we have these quote unquote emerging behaviors like reasoning and all these other things. Well, video is so much more rich than text. Imagine if, you know, what we can get if we can fully take advantage of video. I think the big question is then how do we do that and how do we overcome some of the challenges associated with taking advantage of video?
5:10And maybe we can start with some of those challenges and why it's been so difficult. For example, with text, I think a big part of the breakthrough is figuring out that we can come up with self-supervised training approaches. And some of the same things are available in video, but also it's different and harder in a lot of ways. Can you talk a little bit about some of the challenges with the way we currently treat video in terms of achieving some of the goals that you're looking to achieve? Yeah, there are many challenges. I guess one of the challenges is data coverage. So any of the things that we're saying now is unlikely to be completely out of the distribution of internet text.
5:51It's a little sad to say, but it's probably true. You know, a lot of the words I'm saying now are probably somewhere hidden in the corner of internet text. But for videos, I'm pretty sure like the bedroom I'm in right now is not on the internet. So in some sense, like the image of my bedroom is indeed out of distribution. So this is kind of like the data coverage issue. If we want to generate videos starting from arbitrary image frames, that is actually much harder than trying to generate text, which is just to complete a sentence. And data coverage is one such issue, but we also run into this kind of lack of labels.
6:25So for language, there is something actually quite nice about it, which is the words that come later in a sentence is somewhat of a label or supervision for words that are earlier in a sentence. If we just cast all of the tasks into a language generation task. But for videos, oftentimes they are really only useful when we do conditional generation, meaning that condition on the image frame would give it some input, like move forward, and then we generate this motion, which means that we will need to collect labels for specifically controlling the type of generation that we want. We probably have a lot of internet scale videos from YouTube, for example, but oftentimes there aren't really good labels for these videos.
7:04Maybe there are captions, like some transcriptions of what we're saying right now, but it's not really characterizing what is happening in the scene, right? It's just a transcription. Why doesn't something like infilling or frame interpolation as a task or objective solve the problem of self-supervision for video? Yeah, so it's not very clear. Just predicting the next frame without any external input is actually doing anything useful. Whereas for predicting next words, because a lot of different tasks can be casted into predicting next word. So predicting next word already solves a lot of the tasks.
7:42But for videos, it probably learns good representations, but how do we use those representations for downstream tasks is not very clear. I guess there's another challenge, which is for language models, people were working on different architectures, LSTMs and transformers, but pretty quickly people have converged on autoregressive transformers being the architecture and people start to work together, iterate on the same architecture. But as of now, video generation models are quite heterogeneous. There are diffusion models, there are mass models, and there are autoregressive models. And it's not super clear whether any one of these models will win or a combination of these models would be the best.
8:21So each person or each organization is working on their own kind of best model architecture. So we haven't really come together collectively as a whole to discover the best architecture, Even though perhaps more recently, there might be a little bit more convergence happening where these higher level semantics are sort of tokenized and people use autoregressive models to generate these high level tokens and use diffusion models to really turn these tokens into high resolution pixels. It just took us a while to come here. When you look across the space and the research, how close are we getting to achieving some of what we're seeing on the tech side with video?
9:00Yeah, so if we were to ask what videos can empower us to do that text cannot, and this oftentimes is the type of information in videos is oftentimes much richer. It contains very detailed physical laws of the world, like friction and torque, implicitly all captured in the pixel space. Language is okay when we use it to describe a car driving down the street, but it's fairly high level. It's not really telling any details of how we can do it. And oftentimes people might prefer watching a YouTube video to learn how to peel a pineapple or how to replace the tire on their car as opposed to reading like a step by step manual.
9:38It's just that because the video has a lot of implicit information that helps people understand. So I would say, you know, one such idea would be can we have chat GPT, but generate videos to answer people's questions in specific scenarios. Right. I don't care about a panda running on the street. I want to know, you know, I'm halfway through cooking this meal, but I'm kind of stuck. So what do I need to do next? Or I'm halfway through fixing my laptop, but I'm not sure what to do next. So now I want to customize answers in the form of videos that I can watch and understand how to perform these tasks.
10:10So that is something that text could only help us to a certain degree, but watching a customized generated video as answers is much more helpful. And there are also these visual reasoning tasks. For example, solving these puzzles, you know, you have a circle, circle, square, and what comes next? Visual reasoning is not very easy with language models as well. My background is in reinforcement learning decision making. So there's been a lot of excitement in robotics about leveraging these large models. The issue is that these language outputs from these models are generally pretty high level. If the goal is to put the fruits in the drawer, then language model will output something like open the drawer.
10:49This is still quite some distance away from the low-level controls the robot needs to know exactly how to open the drawer. If we can generate a video of a robot hand opening a drawer, that actually already has all the information pretty much we need for the robot to know this is how much I need to move in the pixel space to reach the handle of the drawer, for example. So video bridges this gap between these very high-level information from language and these very low-level controls and these continuous values that people have to derive in order to control these systems. Continuing your argument from earlier, one might say that if we do a good job of turning the picture into the thousand words and we can do a good job creating pictures or videos from thousands of words, then why do we need to work natively in video at all?
11:38Are there very concrete, succinctly stated examples of what you could never achieve if you have to use text as an intermediary? I mean, it's actually really hard to say because it's hard to define what is a text because, you know, byte streams are text, but people are not doing language models using byte streams. Text has some structure that language model leverage to solve tasks as opposed to just using raw data in its own format. And the same with videos, there is prior knowledge about videos and also shape and rotation invariances in images that this prior knowledge kind of defines this set of data being a special format that is potentially useful for tasks.
12:21That's why, you know, as humans, we have visual system and also language systems to solve a lot of different tasks. Yeah. I think what I hear you saying is that in the end, it's going from some pixels to some representation of those pixels that's expressed as vectors. And whether those vectors are numbers or letters, you still need to get back to some pixels or actions and so you know whether it's text per se or not it's all about the transformations that are getting you from the pixels to the to the actions or the outputs that you're looking for yeah exactly it's really kind of how we represent these data formats right that comes back to the first part of the point that i was making is well as you said if we just convert every pixel into some discrete token describing that pixel, then there is no difference between videos and language.
13:14But it is indeed kind of like the level of abstraction in the sense that language is not raw bytes. It is some structured information. And same with videos. It is some structured information. Yeah. And so the paper, is it a survey of the challenges and kind of a projection of the road ahead? Or are you proposing specific methods that you think will help get us to richer representations for video and being able to the specific task performance that we're targeting here? I would say it's both. It's really to scoping out this problem space to first really recognize what is the thing that made language model work so well.
14:00And the point that the paper makes is this notion of unified representation of information using text and a unified task interface using text generation. And then it's really like the recognition that we can do the same with videos because as we said, videos has all of these information, detailed spatial and visual information about colors and shapes of objects, distances and orientations, for example, they're all captured in video and also dynamics, interactions, behaviors, actions from humans, from robots. And all of this information is in the video as a unified data format. I wanted to drill in on just that point.
14:38When you say unified task interface, what exactly does that mean in the context of text? And then extend that to video. Yeah, I'm glad you asked. So for language, if we look at how natural language processing community has evolved over the years, it's very interesting. So initially, people had, you know, people focusing on specific topics within natural language processing, there are people working on semantic parsing, there are people working on translation, there are people working on summarization, sentiment analysis. And oftentimes, you know, the focus groups probably don't even talk to each other all that much.
15:09And it turns out that all of these different tasks at the end of the day is a language generation task. Because we have a single model and a single objective, then that's why we can just pour internet data into the single model, single objective, and as a result, get a very general model that can perform different kinds of tasks. So that, in my opinion, is really kind of like the takeaway from the success of language models. And now coming back to videos, why is this true for videos? Well, let's think about what kind of things are in videos, right? Maybe there is how to cook or how to fix something.
15:43Those kind of tasks could be characterized as a video generation task. And there are also tasks like, how can I teach my robot how to open the drawer? And that is a video generation task. And also simulation, physical simulation, dynamic simulation, fluid dynamics, and cloud movement. All of those tasks, knowing what happens next if I give the model some control input, these are all video generation tasks. So again, we see the common theme. Once we have a unified task interface, we can pour all of the videos into this video generation task. Yeah. And do you see the interesting problems in this space as inherently multimodal, meaning you've got some video input and then some conditioning that's outlined or provided as text and then a video generation?
16:36Or do you think a single-moded video in, video out has equally useful use cases? I see it as a multimodal problem because, as I said, there are information at different abstraction levels. There are these pixel information that are very detailed and there are these high level abstractions using text. And oftentimes, if we were to look at it from a decision making perspective, which is what a lot of my research focuses on, we have an agent or actor in this setup that's deciding on actions with a policy, which is kind of like a strategy for choosing actions. And we have some environment, which is taking these actions and tells the agent what happens next if I were to take this action and whether this action was good or not.
17:25So clearly we see kind of two different components in the system. One is an agent deciding what to do. Another is an environment telling the agent, what happens next. So people have gradually realized, oh, we can actually make language models into agents. So there's been a lot of work on using language models as agents to interact with external tools like websites or with humans, for example. So I think there is not too much of a recognition that, well, these large models are actually also good world models in a sense that they can actually provide information on what happens next if I were to take an action and whether this action is good or not.
18:04So there has been some work perhaps of using language model to provide feedback for itself so that it can self-improve from this feedback. But that really boils down to this basic setting of decision making, which is when we have some agents interacting with environments and learn on its own, which is how we get superhuman performance. Because if we just do supervised learning or unsupervised learning using the data set, maybe the best we can get is the behavior in the data set. But once we're in the setting of having an agent, interacting with the environment, and learning to self-improve in the same time, this is how we can get superhuman performance.
18:38And this is really kind of becoming more and more possible because we have language model as agents and perhaps video as world models. And this is really kind of like enabling new applications that in the past wouldn't have been possible. Yeah. Do you see specifically an analogy in the video domain to tools within an LLM? Yeah, there's definitely the analogy in the sense that, for example, video models also need to know how to retrieve information. Like, you know, if I'm generating a video, I put a fruit in a drawer and I closed the drawer like five hours ago. Then when we generate the video of opening the drawer, the food should be in the drawer.
19:24So does this mean I have to store five hours worth of video in the context of the video generation model? Or do I need an external database and have a video model to retrieve information from the database? If that's the case, then it's quite similar to this kind of retrieval augmented language models. Reg style. Exactly. Yeah. So there are more challenges associated with it because if we look at the interface is probably some kind of programming language. And as a result, language models are probably more well-equipped for directly outputting those programming language code to kind of use external tools.
20:02But for videos, there is additional information and there has been work on converting images or videos into code. It's a different form of input with different form of information. And we will still have to require a language model to generate some calls to the external tools, but at the end of the day, we still have more information in the video generation model. One example that just came to mind for me was, say you show a video that is part of tossing a ball and you stop after some period of time, the ball is still going up. The generator might want to call out to some tool to run the physics model on what that parabola looks like can provide some parameters and use that to condition the completed generation.
20:49Is that something that you think is, is that the direction that we're headed? Or do you think that that's a viable idea? Yeah, totally. So I think video generation actually has a lot of implications for domains, like you said, science and engineering, right? Because a lot of the scientists I've talked to were like, well, what do we need to do now? Do I need to convert everything I have into text and try to use a language model, right? It's really difficult for them to believe that. But then oftentimes, you know, a lot of the tools that people have built for simulating fluid dynamics or taking, you know, electromicroscopy scans of samples from biology and chemistry and materials, a lot of these data actually come in the form of images and videos.
21:32And a lot of the end result is being able to simulate the dynamics accurately and as a result, being able to optimize some sort of control input to the dynamics well. So in these domains, perhaps video is a very much applicable form of modality than text, because oftentimes a lot of things are just not in text, right? The first observation is in the form of images and videos. So it's a lot easier in these environments to learn a visual dynamics. So condition on some control input, some turbulence that is asserted into the system. And then to try to simulate the visual effects of these external inputs.
22:11So this is really kind of coming back to the question of like, why isn't text alone enough? It's because, well, a lot of these scientific domains are just interested in the visual information. And I think you had a really good point about these physics-based simulators. So as a person who worked in some robotics questions, I realized the sim-to-real gap is big because there's a certain extent to which people can simulate physics using human knowledge. So then the good question is, well, instead of the graphics people and the machine learning people working in separate areas, while using data-driven learning approach, while using physics prior knowledge to build simulators, maybe it makes more sense to kind of bring together data set that has different information or different pieces of the puzzles to solve the problem of simulating the world.
23:00Because for the graphics engine people, maybe they have very precise control of what's being simulated, right? Because they build things from bottom up so they know exactly this is a table and this is a chair. Whereas for more end-to-end learning people, we have a lot of data from the internet. A lot of them are not labeled. A lot of them don't have precise motions or controllable motions in the video. So can we kind of, you know, somehow combine knowledge from both research communities and think about jointly together, how could the graphics community and the video generation community kind of come together to simulate these highly realistic interactions with the real world, but are also very controllable.
23:44So right now, like people use text to control simulation, right? But if I were to specify, okay, make this person run, you know, 1.36 faster, like it's not clear if we can achieve that level of specificity in controlling the video generation. I think that calls to mind Sora and what we have all seen and got excited about with regards to that model. Talk about, in light of this paper and research direction, how did you parse Sora and the degree to which it is a competing idea, a fulfilling idea, a proof point? Talk a little bit about Sora in this context. Yeah, so it's quite interesting because this paper was written before Sora became available.
24:33There were some delays in terms of putting it on archive. So that's why you actually cannot find any information about Sora from the paper. I think the research community, which is what this paper represents, and Sora, which is probably what the industry community represents, are actually heading towards the same direction, but from different angles. From the research perspective is how can we learn good simulators for training agents for improving algorithms? And from OpenAI perhaps is how can we have hyper-realistic generative models for videos? But then I think this realization that generative models can indeed serve as a real world simulator has kind of been unanimously agreed on across the industry and academia.
25:15Because once it can simulate the world very faithfully, there are just very diverse set of applications, implications that has been opened up, like I said in science and engineering, robotics, filmmaking that has just not been possible before. So I guess this is also somewhat why Sora's framing is a little different from previous work in video generation, which is more targeted towards entertainment media, whereas Sora is more targeting towards treating video generation as a real world simulator so that we can do various things with it. so I had another paper called learning interactive real world simulators it's called Unisim it is really kind of conveying the same idea as Sora last October on how video generation is the real world simulator and the paper went on to like using this simulator to train robot policies and of course maybe that's something that Sora hasn't been doing or maybe Sora will be doing as OpenAI is investing in maybe robotics companies.
26:16But I would say it is kind of using similar lines of ideas to kind of tackle the same set of problems, which is really kind of using video generation as a simulator. I wanted to go back to one of the other parallels you mentioned in the paper between language models and video models, and that is the idea of reasoning and chain of thought in particular. Can you talk about how you see that playing out from a video perspective? Yeah, so definitely. So I give some examples of like maybe solving visual puzzles using video generation models. For example, geometry problems, right? Oftentimes, drawing the right auxiliary line is pretty much equivalent to solving the geometric problem.
27:03So in that case, we could ask a language model to describe, you know, these are A, B, C, D, E points, between which two points should I draw the auxiliary line. But oftentimes, you know, when people solve these geometric problems is to reason, well, once I have a line here, I have a 90 degree angle here and so on and so forth. And then being able to kind of generate that as an answer to this geometric problem is one way of video models can reason and solve these geometric problems. There are also kind of other approaches, which I have done some of the work. So some of the algorithms, when we actually execute them, for example, breadth-first search, right?
Read the full transcript
27:41So we actually have this kind of search trajectories rolling out in space. And that itself is actually a video also, where each of the locations are storing the information about whether this location has been visited and what action has visited it. So if we were to roll out these search space using videos, that is equivalent to a video model has learned how to search. So in some sense, this is actually quite analogous to using language models as kind of a computer algorithm to solve problems. Another thing about this in-context learning or reasoning is the ability to kind of self-improve. So we've seen that in language models, right?
28:21So when we interact with ChatGPT and maybe it didn't get it right the first time, we have some external feedback, it can self-improve. And right now, I think for video models, it's not quite there yet in terms of self-improvement. In a sense, well, if I generate a video and a user tells the model, this is why it's unrealistic because, you know, the chairs don't jump by themselves. People have to move them around. and incorporating this feedback to generate a better video through self-improvement is some kind of reasoning capabilities in terms of self-improvement that we haven't seen in video generation models.
28:59When you think about the fact that we haven't seen things like that kind of reasoning, do you think that we're just not on the right place on the scale curve with regard to video? You know, compute has always been a big issue with regards to video. Or do you think there are other more fundamental factors? I think there are definitely these technical challenges like data set coverage and also the amount of labels we have for videos, like I said earlier. But also in some sense, a lot of the typical reasoning problems that people consider are more or less in a form of a textual or language-based reasoning.
29:39While there is an argument that people actually imagine into the future for example like athletes probably does this visualization of how they imagine the curve of the ball would go if they were to like kick the ball in certain directions but it's not super clear if you know imagining things in the pixel space is the right way to reason about these kind of activities whereas for text very clearly reasoning is a big problem but for videos it's like we know videos can also reason especially for these more visual domain tasks like geometric problems and so on. But as of now, maybe the problem space is at least a little bit limited compared to the amount of problems we have in using text to reason.
30:23While reasoning is a goal that we want to apply these video models to, we don't really have a complete concept of what that means. We don't have data sets and benchmarks and all the things that we have on the tech side because it's much more clear. Exactly. And sometimes maybe people use different terminologies. In a sense, a simulator, a world model, is a reasoning model. Why? Because it's really reasoning about real-world physics. It's just not directly, explicitly reasoning in language. But being able to simulate the next frame accurately according to physics is a form of reasoning. So it's just the definition of reasoning is different and the problem space is different.
31:09The formulation is different. So it's just so new that people haven't even gotten a chance to think about it yet. And so when you talk about the types of models that we have for video diffusion, autoregressive, mass models, better future models, and you review these in the paper, Are you talking about the degrees to which they are doing some of these things? Are you talking about where they're lacking?
31:40How do you think about where we are today relative to what you want to see happening in the field? I think the heterogeneity around video models is still going to persist in the sense that research in better diffusion models or research in better mass models for videos will still probably continue to be published because people haven't extremely been convinced that this kind of transformer diffusion is the solution. although I'm sure there will be more researchers working in this direction. And of course, each video model has its advantages and disadvantages, such as efficiency during sampling time or whether it's easy to have long context videos.
32:23One thing I would point out in this direction is when we think about better models, it's always with respect to the data. So when we only try to develop better video models, is not so clear about what we mean by that. Because only simulating extremely long videos is not that useful because we want to simulate videos that can be controlled, that can take actions and can serve as world models. So in that case, we should really also be thinking about the data side more. For example, we have hours of YouTube videos, for example, of uneventful scenes. And how should we deal with these videos? Should we only train these models if there is something happening in the video?
33:07Or what if something happening is just kind of waves going through water surfaces? Those videos are also useful for simulating fluid dynamics, for example. So it's just not clear that because video is such like a dense format of information and how should we be dealing with it? So when people are developing better model architectures, somewhat it's kind of not really taking into account of the data problem, which is, you know, video probably has a lot of rich information, a lot of data, but is not extremely useful for the modeling purposes and how should we, like one example would be to learn more compact representations across time, right?
33:49We've always been modeling kind of like a fixed number of frames as outputs, but then it's kind of like, well, if there is some interesting thing is happening in a particular frame and nothing's happening in the other frames, should the encoding of the video be a little different to address the events in the video, as opposed to a long sequence of frames of uneventful frames? Is a kind of pixel first approach that, you know, isn't trying to identify lower dimensionality or compressed representations or semantic representations or something like that? Is that, Do you see that as being the feature or are the models going to need to adapt to specific use cases and featureize the video to pull out what's most important for a given use case?
34:40So I think this will kind of go back and forth a little bit, right? So just similar to how language models, you know, people were working on different natural language processing tasks and come together to develop language models, but then trying to fine tune or efficient fine tune or prompt the language model to do specific things. I think the same will happen for videos. Now we're only at the very first stage of recognizing that video is a unified format to represent a lot of different information. And let's try to train a large video just using all the video data we have. But eventually, right, when these videos are being deployed, I'm not sure if we should use the exact same pre-trained model for modeling the cloud versus for modeling, you know, a person trying to fix their car because the dynamics are just completely different.
35:25So in some scenarios, maybe there would be some task-specific adaptation and fine-tuning in particular downstream domains. But, well, ultimately, I think the idea of having a unified data format to consolidate information using a large video generation model, and then which information we elicit downstream for solving tasks is kind of like a second stage kind of process. but it seems like the same trend has happened for language generation where we consolidated information first and then elicited relevant information for solving tasks. Because if we only use the relevant information at the beginning, then we never get the knowledge from data sets that are vaguely related to this task.
36:10So this is really like how can I use data or knowledge or information that is in tasks that are vaguely related but still be able to help a particular downstream task. This is how kind of we get into this space of foundation models being kind of fine-tuned or adapted to task-specific settings. And I will also want to make another point, like, even though this paper was talking about videos, I've also been thinking about kind of this unified representation of information, for example, structured data, right? Because there are things that video cannot capture, which is this micro-level information.
36:47Like we just we don't have that many kind of high, you know, like microscopes to look at these information. So then maybe at that level, a different representation of information will be required, which would be something like atoms in 3D space. Right. And then again, the same kind of philosophy applies, which is, well, we have people working on different types of materials from different research groups. Maybe someone at a university is working on batteries and someone at an industry is working on semiconductors. And if we don't have a unified representation of information or unified data format, these two groups might never talk to each other because they care about different kind of materials.
37:24But fundamentally, the knowledge about quantum mechanics, quantum physics, how atoms interact with each other is actually shared across different data sets. So this is really kind of like why having a foundation model built upon these shared representations or shared data formats. In this case, it could just be atoms in a 3D space, organized in 3D space. And once we can learn a generative model on this, we can fine tune this to discover new materials that has never been seen before. So this is kind of like a trend that we've seen across language, across video and across science, biology or materials.
37:58But then the idea of gathering together all of the data that are, you know, even though they're vaguely related, there is still some fundamental knowledge about these data that we can use to downstream to elicit relevant information. How do you expect to build on this paper or how do you hope that others will build on it? There is definitely a lot of room for solving video generation on its own. Like, for example, these models hallucinate a lot. You know, both my videos hallucinate and, you know, Sora also has these hallucination examples. How can we have these videos to more faithfully output real world physics?
38:33So there are techniques that we can think about. I've been working on that to have these videos hallucinate less. And I think these videos are also not, doesn't generalize super well. That's partially why maybe companies are a little bit more careful about releasing kind of like the APIs to the video generation model because, well, of course there is safety and ethics concerns, but also if I just come up with a prompt and at least for the video models I've worked with, you know, for some of the samples, if I sample a large amount, a lot of times something could go wrong in these generated videos that just doesn't look as good.
39:10So there's still a lot of room to improve in terms of generalizing to better user input, right? Because ultimately a real world model should allow me to upload a picture of my bedroom and start to interact with the object in my bedroom. But, you know, that's not quite the case yet. So generalization is another challenge. Then there are also challenges around efficient fine tuning because these models are also going to get quite big. And what are the fine tuning strategies we have for these models? So these are all kind of follow up work. And I think the more exciting work is really in the applications of video generation models to solve, say, robotics tasks or science tasks.
39:44Yeah. Up until now, we've talked somewhat broadly and abstractly about video and world models and some of the concepts that come up in your paper. But we're going to switch now to a demo that you've produced with the models that you've created. If you are listening to this on Apple Podcasts or Spotify, hit the show notes page and we'll link you over to the YouTube video where you can catch and follow along with this demo. We'll also continue to kind of roll with this so you can continue listening as well. So what are you going to show us, Sherry? Yeah, so this is one of the work of learning an interactive real world simulator or learning a world model, the work that I've done with some collaborators.
40:30And in this demo, to really demonstrate to the listeners what a world model should look like, we first have a few different kind of scenes that a user can choose from, right? So I have a set of pre-stored scenes, but in an ideal world, you should just be able to upload any images you would like and start to interact with it. So Sam, could you choose from one of these scenes? So right now we're looking at a cutting board on a table with a carrot. What is the third scene? They're actually very small on my screen. The third one is actually like a switch on the wall. Oh, switches. So, you know, we just clicked on the switches.
41:10Exactly. So, well, since you chose this one, you are responsible for choosing an action to interact with the switches. So which one would you choose? Let's push the middle switch. press the middle switch. Okay, let's see what happens if we were to simulate this action. And you see like a person's hand actually shows up and went on to press the switch, right? Do you want to choose another one just to, you know, really make sure it's working to the audience? Let's plug in a cable. And let's see what happens in this case, right? We see like a person trying to plug in the cable. So you see the scene is still the same scene and but we can interact with it using different actions like pressing different buttons.
41:48You see like this time someone uses a pen to press it and you know it's really to demonstrate that starting from the same initial frame we can use different actions to interact with a world model do you want to choose another i'm judging by the fact that the the switches are moving around in the frame that it's uh i i guess not particularly strongly conditioned on that first frame it's not like you're superimposing. It's all generated on the fly. Is that the idea? Yeah. So all of these videos are generated. And one of the reasons that the video kind of moves around a little bit is because in the pre-training data, a lot of the pre-training data was collected by humans wearing a GoPro on their head while recording videos and doing different activities.
42:38That's why there's like a lot of movement, as you can see in these videos. And so let's talk about the relationship between the training data and what we're seeing here. So So is there an imitation component of this wherein at some point someone was pressing these switches? And the answer is no. So I didn't really dig up to see what the training image looked like. And very likely it could just be an image of the switch because the model is also trained on internet scale image data. So the point is that, you know, none of these videos actually exist in practice and all of them are simulated. Right.
43:16Right. And is the input or control, is it just the phrases that we see on the buttons here? Press left, press middle, press right? So you can actually really input any phrases you would like, but I don't have the server running. That's why I'm showing you like a fixed set of actions I've generated in the past. But, you know, it's really diverse. So if I were to come back next time, you can tell me exactly what you want to simulate in text and I can try to simulate it. There are also low-level controls because this model is also trained to control robots. So there are robot actions which are delta X, delta Y values or torque forces.
44:03So those ones just doesn't have too much of semantic meaning. So it's not as exciting in the demo. We can really have these kind of egocentric motions, have a person maybe wash hands and you see a sink shows up and the person goes on to wash hands. But we can also have another step followed by washing hands, asking the person to shut off water, for example. And oftentimes manipulation and navigation are considered separate tasks in robotics or in these kind of simulators. But in this case, it's really we've just seen some manipulation tasks. We can also use this same model for navigation, which is to ask the person to turn right.
44:42To turn left, you see a table shows up and the person's moving away from the kitchen and so on and so forth. So I guess if we have the time, I really want to show this one last demo, which really answers your question about whether this video exists on the Internet or was pre-recorded. So in this example, we have the same piece of napkin sitting on the table, but we can uncover different kinds of objects. and could you pick one of the objects from this list? Maybe not spider because it's a little disturbing. You know that was the one that I was going to pick. I should just make this one. Okay, let's see what happens when we click on the bottle.
45:22We see a hand actually shows up and uncovers the napkin and a bottle shows up. And again, we can simulate interactions with a bottle of moving the bottle to different directions or maybe put another cup next to the scene. So the point I'm trying to make here is really because we have internet data, like pen, bottle, toothpaste, plates. So this is why we can specify using text what objects we would like to uncover when we remove the napkin. And it's just not possible to be pre-recorded because no one is sitting next to the napkin and doing all these tasks. So this is really showing a fixed set of objects.
46:07It's really almost infinite because we have infinite internet images with infinite number of objects almost. So we can really create an environment in which humans or robots can interact with different objects to complete different tasks. And this is maybe the ultimate augmented reality of simulating real world. So put another way, in the demo, there are some fixed choices that the user makes, but that could be a text box given the server's up and running, sufficient compute, all that kind of thing. Exactly. And you can enter an arbitrary item to show up behind the napkin and even arbitrary actions to interact with the item as well.
46:56Exactly. Yeah, because, you know, there are many images or many scenes I have, and it just takes me a while to just generate all of these videos. That's why it seems like it's limited. It's really could be a text box where you can just tell me what you want to see on the fly. Of course, it has to be something reasonable given the observation. So if you have a plate here, but you say something completely out of the ordinary, then maybe the generation wouldn't be so realistic. But this has a lot of implications even in remote tourism, for example. So this is like an image of the Sistine Chapel. And if a tourist wants to really navigate in a space by taking different actions, it could do so and looking at different directions and zooming and so on and so forth.
47:41And because we have internet scale data, so we can really do this for kind of any scenes once we scale up the model enough and have enough of these actions. We're kind of flying around the Golden Gate Bridge right now. Exactly. And you can choose to which direction you would like to fly around or if you would like to see more or see less of the Golden Gate Bridge. So all of this is generated by a single model. And unlike what we've seen in terms of generated images to date, for the most part, or videos rather, where you start with a static prompt and you generate a video, here it's interactive.
48:17so that, like you mentioned, that could be a text box. You could say, put the apple on the plate, cut the apple, take away two of the pieces, and continuing to build on through prompting this video that you're creating. Exactly, exactly. So this is what a real kind of a world model should look like. It should support diverse actions starting from the same observation and also repeated interactions. So after the first interaction, the model should remember the previous state of the interaction and allow us to interact with it again. Well, thank you so much, Sherry. Yeah, thank you so much, Sam.
From the publisher
Today we’re joined by Sherry Yang, senior research scientist at Google DeepMind and a PhD student at UC Berkeley. In this interview, we discuss her new paper, "Video as the New Language for Real-World Decision Making,” which explores how generative video models can play a role similar to language models as a way to solve tasks in the real world. Sherry draws the analogy between natural language as a unified representation of information and text prediction as a common task interface and demonstrates how video as a medium and generative video as a task exhibit similar properties. This formulation enables video generation models to play a variety of real-world roles as planners, agents, compute engines, and environment simulators. Finally, we explore UniSim, an interactive demo of Sherry's work and a preview of her vision for interacting with AI-generated environments.
The complete show notes for this episode can be found at twimlai.com/go/676.




