World Models, Explained

17 Jul 2026 · 1 h 14 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains “world models” as a promising solution to AI’s sample-efficiency problem—how to learn new tasks from far fewer training examples. It frames the goal as learning a transition function (predicting next state given current state and action) so policies can be improved via test-time planning, especially in robotics and self-driving where model-free approaches struggle.

Guest backgrounds

The transcript is a two-person discussion; no guest names or bios are provided. One speaker references researchers such as Francis Chalets/Francis (for “intelligence as skill acquisition rate”), Richard Sutton (cited as “Richardson” for a 1967 basketball imagery study), Shaw Druckmann (Stanford neuroscientist), and Danijar Hafner (Dreamer). No personal guest credentials are stated.

Key claims

Perfect world models could require zero environment samples (example: Newtonian physics + model predictive control). World models enable cheaper, faster planning than separate policy/value models. Robotics/self-driving break down “implicit” understanding because dynamics are non-differentiable and other agents adapt.

Notable examples

Newton’s second law for convex optimal control; adversarial drone leading to non-differentiable RL; chess/Go with MCTS and small action spaces; self-driving with enormous (steering/brake/gas) action cardinality and infinite-like pixel state; Dreamer/DreamerV4 using synthetic rollouts to mine diamonds in Minecraft.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Defining Sample Efficiency and Its Implications

0:39 to 4:27

Explore the definition of sample efficiency and how it compares between humans and AI.

“You and I have talked a lot about the various ways people are training models and the sample efficiency of them.”

Extreme Cases of Sample Efficiency

4:27 to 6:10

Discuss scenarios where perfect sample efficiency could exist and real-world applications.

“Um, that maybe Bill Gates, uh, Steve jobs and Jensen have 50 years of, of, you know, world modeling experience to know what people want.”

Model Predictive Control Explained

6:10 to 8:52

Learn about model predictive control and its significance in achieving optimal actions.

“But it seems like in certain domains, especially in robotics and self-driving, as we'll talk about, that sort of breaks down.”

Challenges in Non-Differentiable Scenarios

8:52 to 13:22

Examine challenges in reinforcement learning when faced with non-differentiable processes.

“And so we know that the position PT plus one is going to equal PT plus delta T VT plus one half delta T squared.”

Building Better World Models

13:22 to 14:00

Discuss methods to improve world models and the role of policies in reinforcement learning.

“And you'll hear things like when you study initial reinforcement learning called value iteration or policy iteration.”

Understanding World Models and Policy Training

14:00 to 18:00

Explore how world models are trained and their role in reinforcement learning.

“And that we're going to train it over many, many instantiations of this.”

Applications of Reinforcement Learning in Games

18:00 to 21:00

Discuss the use of reinforcement learning in chess and Go, comparing their complexities.

“And that's far fewer samples to do this post action conditioning if I already have a really good ST to ST plus one world model.”

Monte Carlo Tree Search in Game Strategy

21:00 to 27:20

Learn about Monte Carlo Tree Search and its effectiveness in planning strategies for games.

“Why don't we talk about that for just a second?”

Estimating Values and Actions in RL

27:20 to 28:00

Understand how to estimate values and select actions using reinforcement learning.

“And so this trend in RL is just called test time planning.”

Understanding Monte Carlo Tree Search (MCTS)

28:00 to 31:52

Learn about the MCTS process and how it applies to action selection in games.

“This will give me 361 numbers that sum to one.”
Show all 24 chapters

Challenges of Scaling MCTS in Complex Games

31:52 to 34:36

Explore the limitations of MCTS when applied to larger action spaces, such as in Go.

“And then for all 800, I have to go through this whole process and I have to invoke the model at least 30 times to get through all here.”

Self-Driving Cars vs. Game AI

34:36 to 39:52

Discuss the differences between self-driving car challenges and game AI like AlphaGo.

“The important thing to pick up is that we did 800 MCTS simulations and to cover 361 possible actions on average.”

Action Spaces in Self-Driving Cars

40:04 to 42:06

Delve into the complexities of action spaces and data collection for self-driving cars.

“One way to look at the action space is that it seems relatively small.”

Understanding Model-Free vs. Model-Based Reinforcement Learning

42:06 to 48:50

Learn about the distinctions between model-free and model-based reinforcement learning and their implications for AI development.

“I mean, the amount of work that they're doing at FSD is like incredible.”

World Models and Their Role in Robotics

48:50 to 55:08

Discover how world models can enhance robotic learning and planning through synthetic data and action conditioning.

“is so big, why don't we talk a little bit about how world models actually fit into this?”

Insights from Neuroscience on World Modeling

55:08 to 56:00

Explore the parallels between human cognition and robotic world modeling, highlighting recent research that supports this connection.

“What's also cool is there's a bunch of applications of this, the things outside of robotics too.”

Understanding World Modeling and its Implications

56:00 to 57:00

Learn about how the human brain develops world modeling and its relevance to robotics.

“And then basically predict the next state of the world based on those things with this lingive and diffusion rollouts.”

Introducing JEPA and Its Role in World Models

57:00 to 58:20

Discover the JEPA concept in world modeling and its implications for reinforcement learning.

“Because I think there's been a number of papers that use JEPA as an element of their, I guess, architecture.”

Techniques for Enhancing World Modeling

58:20 to 1:01:30

Explore advanced techniques and challenges in world modeling including self-supervised learning.

“I think my first paper was basically doing something like this, basically turning like a grid into like a bunch of like pyramids.”

Challenges and Open Problems in World Modeling

1:01:30 to 1:04:00

Identify key challenges in world modeling and issues with current techniques, such as the limitations of physics-informed neural networks.

“And I tokenized this into a bunch of different tokens here.”

The Squint Test and Brain Function

1:04:00 to 1:10:00

Discuss the concept of the squint test and how it relates to understanding the human brain and world models.

“problems maybe we can emphasize here that the community can go emphasize working yeah so um I think the first one is that pins doesn't really work.”

The Role of Sleep in Learning and Memory

1:10:00 to 1:11:40

Explore how sleep contributes to memory formation and learning processes in the brain.

“feels intuitively like what our brain is doing.”

Advancements in World Models and Robotics

1:11:40 to 1:13:05

Discuss the future of world models in robotics and their potential applications.

“We have all this work happening with world models.”

Challenges in Robotic Sensing and Control

1:13:05 to 1:14:14

Examine the limitations of robotic sensory systems and the need for more natural feedback.

“And then on the robotic side, there's real issues.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Francois Chaubard:One of the biggest open problems in AI right now is how to solve sample efficiency. That is, how do you get models to quickly learn new tasks or skills from relatively small amounts of training data?

0:08Ankit Gupta:Humans do this incredibly well. We can learn new games, concepts, and skills, often after just a handful of tries. Our best models, on the other hand, often need tens of thousands of data points just to learn.

0:18Francois Chaubard:So today we're going to discuss what many top researchers believe is the most promising path to closing that gap, world models.

0:24Ankit Gupta:We're going to discuss the motivation and math behind world models, current applications, and why this approach might be the key to unlocking AGI.

0:39You and I have talked a lot about the various ways people are training models and the sample

0:44Francois Chaubard:efficiency of them. Why don't we start by just defining sample efficiency and how we intuitively think about it as humans?

0:49Ankit Gupta:Yeah. Yeah. So I think from my perspective, the two major problems that we have left to solve is intelligence per watt and intelligence per sample. Intelligence per watt is like how many valve perplexity points we get per watt of spend. And then intelligence per sample is basically if I have one additional sample in my data set, how much more intelligent am I getting? And so if I imagine I have a new tasks like RKGI, for example, I think like really Francis Chalet has been on the forefront of this thinking. and talking about intelligence as a rate of skill acquisition versus skill acquisition.

1:22Ankit Gupta:And that's very different. And so how fast do we get smarter with more and more samples? And these things are incredibly poor at getting smarter with fewer and fewer samples.

1:33Francois Chaubard:And for context, the RKGI test sets are a really good example of cases where humans are intuitively very good at them. Most humans can intuitively solve those puzzles with some amount of thinking and effort. but our current state-of-the-art AI systems, what people consider frontier intelligence, basically can't do them. Right.

1:52Ankit Gupta:I mean, we come into new problems with such inductive bias from K through 12, like all these math and school that we've had that, you know, these models are kind of getting from the entire, compressing the entire internet. And so when we come in, we're not coming in tabula rasa just like bare bones, but even so that they have, you know, I don't know what percent of the internet you've read. I've read very little percent of the internet. But despite that, and having read the entire internet, it still can't really do well in generalizing to these new tasks.

2:22Francois Chaubard:So now let's think about this in the extreme cases. In the extreme case where, let's say, we were perfectly sample efficient, we were as sample efficient as possible. What would that mean in terms of a model that is taking a set of actions in the world?

2:37Ankit Gupta:Well, I guess the perfect sample efficiency would be zero samples. And like, there are examples of this. And that sounds absurd to say, but the example, the hypothetical I'll give on this is imagine I had a perfect world model. Then I should never go to the environment to go and collect samples to train on. And well, that can't possibly happen. No, it actually can happen. We do it all the time. It's called Newton's second law of motion. It's like Newton mechanics. Like we basically know how to like get an object from point A to point B with a rocket quite easily just by following like Newton's laws of motion.

3:14Francois Chaubard:Yeah, like when NASA plans to intercept an asteroid and is planning it, you know, years in advance and can set it off on a trajectory where it just glides to the right thing and intersects to the right point, that is an example of a perfect world model we've built where we're then just letting that world model act. And that system does not need to intelligently collect new samples from the environment to decide which direction to go next. It's already been pre-programmed and it can perfectly do it.

3:39Ankit Gupta:Yeah, can you imagine if like we needed to collect to 1 million training examples of like us shooting spaceships to the moon to like know how to do it it's like this complete it would be we definitely wouldn't have the Apollo missions right um but we do have that that ability because the the real world is differentiable and we can do something called model predictive control that we're going to talk about in a little bit um but even in our own brain I was just uh uh you know thinking about this on the drive up but like there's so many ways that like I can basically think about the things that you are going to say or what a VC is going to say when I'm, when I was pitching them or what a customer might say, uh, and even product being, having taste, what is taste is like predicting that other people are going to like this thing.

4:20Ankit Gupta:And so we've built this world model over years of entrepreneurship, 10 years of like getting it wrong. Right. Um, that maybe Bill Gates, uh, Steve jobs and Jensen have 50 years of, of, you know, world modeling experience to know what people want. And, uh, and, and basically this is actually proven in the 1967 uh cog sci study uh by richardson that basically showed that if you take a cohort of three different people three groups of people and you uh have one go practice layups in basketball and they go and they shoot they they improve it for one hour they improve by like i think it was like 24 or something like that and then if you take the other one and they just blindfold them and they imagine laying up basketball they improve it 23 percent interesting And against the control.

5:09Ankit Gupta:I mean, that's insane. It means that we have this crazy good world model. And there's this neuroscientist at Stanford named Shaw Druckmann, who basically is of the view that the entire point of the growing neocortex during the great cortical expansion 10 million years ago was to get better and better and better and better world modeling. And having just like my little VLA, which we'll define, of predicting the next action is not as good as having a world model to lean on, either for training purposes or for test time adaptation. Yeah.

5:39Francois Chaubard:What it fundamentally comes down to is we as humans, we think about our intuitive ability to think as coming from some implicit world model we have in our heads encoded by genetics and our ability to learn and whatever else. It seems like models can do surprisingly intelligent things despite not having an explicit world model when it comes to natural language. When they're just talking, it seems like, you know, maybe under the hood, deep inside the weight somewhere, there's some kind of implicit understanding of the world. But there isn't an explicit representation of that. But it seems like in certain domains, especially in robotics and self-driving, as we'll talk about, that sort of breaks down.

6:15Francois Chaubard:And, you know, maybe it would be helpful now to just think a little bit about and sort of define some of the pieces of what makes it challenging in these different domains. And then we can use that to kind of build up to why it's particularly hard in things like self-driving and robotics to get these types of predictive models to work.

6:33Ankit Gupta:Yeah, let's do it. So let's actually like take a step back and just talk about like control, reinforcement learning and define some common terms. So typically in, we teach a course about decision making under uncertainty, which is like the main reinforcement learning course at Stanford. I like to show a specific example of, let's say I have some drone and this is my poor little drone here and it has some mass M and we know that that gravity G is pulling down on it. And it's currently at position T with velocity T, which we will collectively call the state. and to be really clear this is going to be p x py p z t t t and v x v y z v z it's like the six dimensional state vector yep and we have uh some thrust vector u that we control and we're trying to get to some point p star and v star which is v star is typically zero and so you have some platform that i want this thing this drone to land on this is this control problem right and so uh let's say this is like and we'll go through optical or optimal optimal optimal control so how would i actually solve this so the first thing i need to know is my transition function and so this is my state transition function which isst plus one given the previous givenst and my action which which i control is ut and so this is my state transition or dynamics function or a world model this is a world model this is like a very fundamental for

8:17Francois Chaubard:for context you know this this equivalent to transition function you would think about in

8:20Ankit Gupta:rl in general exactly and so uh and then what i'm trying to learn is something called a policy which is like what UT should I emit given some ST? And so this is the ultimate question. What should I do? What action should I take given some state ST? And so the way that we'll solve this, and luckily we have a world model that is perfect. It's just Newtonian physics. Newtonian physics. This is like Newton's second law of motion, which is F equals MA. And so we know that the position PT plus one is going to equal PT plus delta T VT plus one half delta T squared. So everyone's taking high school physics.

9:11Ankit Gupta:And the same thing for the velocity. Delta T A. And then my acceleration is the sum of the forces, which is going to be my UT. I think I divide by the mass and G. And so that's it. And now I have my transition function. Now, how do I get to a policy? And I'm going to apply something called model predictive control or real-time model predictive control, which is like the way that SpaceX lands the rocket on some platform in the ocean. And what you're going to do is you're going to set up your loss function. you're going to minimize sum over all t. You have ut to infinity. And I'm going to minimize my p star minus pt plus v star minus vt.

10:02Ankit Gupta:And usually you add this little lambda ut, which is like how much energy you're exerting. and you can't have infinite thrust. So you typically will have to say UT, U max thrust. That can be achieved. And so this is easily solvable with convex optimization. And so this is convex. This is convex. This is convex. The sum of convex functions is convex. This is a convex constraint. And so I DCP discipline convex programming means that I can put this into CVXPy and it will just give me out my policy, which will be the solution will be the optimal ut plus one all the way to infinity.

10:50Francois Chaubard:So we can solve this in closed form, basically. Because we have this world model of Newtonian physics, we can say at every step exactly how this drone should fly so that it lands on the appropriate thing. Exactly. Under a set of constraints like max thrust. Exactly.

11:04Ankit Gupta:You'll run your log barrier, interior point, whatever, to some solver. on this and it will give me my optimal, then this would be literally the optimal path that this thing can take to get to this state. And that will minimize and I can increase this if I want it to do the least energy path and I make that zero if I want it to be the fastest. And so that's typically the way that you would do what I would call like deterministic

11:37Ankit Gupta:differentiable control. right and why differentiable because i can take the i can form the lagrangian by by taking this minus this constraint and uh and take the gradient of it and i can do monorobins you use the fact that it's differentiable to to do the optimization exactly if this is if this is non-differentiable you cannot do convex optimization and you cannot do sgd uh even if it's non-convex you could still solve and get and get a pretty good solution uh as we do in deep learning but i i you if it's non-differentiable you kind of can't there's nothing you can do so yeah let's have an example

12:10Francois Chaubard:then of how you could make this non-differentiable like well what's a what's a scenario i guess even it's like this drone scenario where it now becomes non-differentiable yeah so i'll put this adversary

12:18Ankit Gupta:named unkit okay and and your job is to you have another drone let's say unkit's drone is to try

12:26Francois Chaubard:to hit me and stop me from getting there now from the position of your drone you don't know what actions I'm going to take.

12:32Ankit Gupta:Right. And so now let's just call this the, uh, this would be now we're definitely not, not deterministic or stochastic, um, and stochastic and non-differentiable. Yeah. And in this case, my state transition, what is ST plus one? It's going to be my say, I mean, now my thrust and what Ankit's going to do. Right. And these, it was all differentiable until this new variable. And I can't like back prop through your brain to say what you're going to do with your little drone controller. Right. It's completely non-differentiable now. And I'm resorting and I have to resort to this awful area called reinforcement learning, which is just super brutal.

13:19Ankit Gupta:And it's sprawling. And there's so many different things. And you'll hear things like when you study initial reinforcement learning called value iteration or policy iteration. And there's DQN or deep Q learning or just Q learning. There's actor critic. There's all this stuff. All of this stuff ultimately comes down to ways to estimate to model this non differentiable stochastic process.

13:49Francois Chaubard:Exactly.

13:50Ankit Gupta:Yeah. And so that's basically the main thing is you're going to start talking about this as a model where I'm going to introduce this psi to say that this is going to be some model that's going to take in these things and then output this. And that we're going to train it over many, many instantiations of this. And that's to get a better and better world model. And then I need to train some policy, A, T, S, T. And then typically you also need a value function. And that is the value of some state and to discern between the value of different states. And like in this case, I don't know what a valid state is, but like, let's just say I was doing like the SpaceX with launching rockets and landing rockets in Florida.

14:35of let's just say that like there's different if i have my launch pad here and i have a whole bunch

14:40Ankit Gupta:of houses here let's just say the path going from here to here i may think that doing this and then coming across here and burning all these houses alive right maybe not highly value so i might say as an example they typically call this like some type of cone here and i might say like it's low value to be here and it's very high value to be to be in this cone or something right

15:07Francois Chaubard:in a sense the value gives you some expectation of future rewards like the sum of future rewards you're getting and so if you're in a bad space you would set the value to zero or negative

15:16Ankit Gupta:negative infinity or something yeah so so we can we can we should introduce our rt as well and so typically like if you're playing go or chess like winning the game uh you can say winning the game is plus one minus one for losing draw zero that's what's done in alpha go in chess we have these heuristics like a a pawn is point was worth one point a rook is worth five et cetera et cetera so you can like already have reward is the difference in in in board board state um and then this yes So it would be the sum of my discount. I should just do T of RT. Yeah. Given. And it's important also to, to, to use this nomenclature V pi.

16:03Ankit Gupta:And the reason why that's important is because what it, what's actually happening here is this is the discounted reward following policy pi. Correct. And that means that when I'm in this state, I will take this action and then I'll end up in this to SC plus one. And then I'll take this action. And it's, it and taking it greedy and so that's the value with respect to pi yeah and so ultimately what

16:22Francois Chaubard:it comes down to is we are trying to still find a new policy pi and along the way we will use gene learning models in various capacities this is standard rl to estimate the value function given the rewards we're receiving right and then where world models come in is a way of incorporating all of those into some sort of joint modeling of the state and action distribution so that we can to make more intelligent policies off of it.

16:48Ankit Gupta:Right. And so your standard kind of setup for this is what I'm always trying to get to at the end of the day is some joint distribution, which would be ST plus one given where I'm at now, where I'm at now. And then this factorizes with chain rule simply to my pi, my policy, AT given ST. And my world model. And I'll give this, this is usually represented with theta. And this is my world model, which would be ST plus one given ST and AT. And so and these are typically learned separately. And like and like you can imagine, in fact, actually, you can actually learn this. This is a video generation model.

Read the full transcript

17:36Ankit Gupta:And I have the frame ST and I predict the next frame ST plus one. Right. And then and we'll get into this. For those of us who kind of saw our diffusion model series, often people these days use video diffusion for exactly this. Yeah. And then what you can do, and this is like the in vogue thing to do since Danijar and the Dreamer paper series from V1 to V4 is do action conditioning later. Like similar to clip where we will inject this like input head, input tail to come into the model to influence and enable the world model to have embodiment. What does that mean? It means that not only can I predict like as a plant or tree on growing on the side of the building, I can like see the world passing by, but I can actually influence it and I can change the world and I can learn that with AT.

18:25Ankit Gupta:And that's far fewer samples to do this post action conditioning if I already have a really good ST to ST plus one world model.

18:34Francois Chaubard:And so here you're saying, you know, what's also in vogue now is jointly training these versus separately training them. Exactly.

18:42Ankit Gupta:So this is called that is called a world action model where some of the issues here is one. There's all these training dynamics. If these things are just disparate training on different sets and things like that. The other issue is plainly obvious. What I have to do to actually do test time planning is I'll have to sample my with model one, invoke theta, and then pass that sampled action into here and then roll it out to ST plus one. And it's very expensive and it's a very not real time. Two major issues and why, like, why can't we just scale up alpha go to like solve all the problems is because it's because of this property.

19:20Ankit Gupta:if I have one invocation to the model and it gives me both. Here's the action I should take. And here's the ST plus one that'll end up much, much cheaper and much, much faster.

19:29Francois Chaubard:Okay. So I think that's a really good segue. I think why don't we now motivate everything we just described through a series of increasingly complex environments. So I'll contend that I think the right set of environments for us to consider is chess followed by Go, followed by self-driving, followed by robotics.

19:48Ankit Gupta:All right. So let's go through a couple examples of problems that we want to apply, uh, reinforcement learning to. So chess is, is a pretty easy one. There's an eight by eight grid. Um, and so typically when you, when you, uh, approach any, uh, RL problem, you're going to look at, uh, star. And so this, this, the size of the state, uh, uh, the number of states I can be in. So if I have these eight here and these eight, so this was eight, 16, 32. So it'd be 32. to the 64 yes quite large quite large then uh my transition function is stochastic and non-differentiable because you can you know what the other player's gonna do so if i'm like in uh playing chess.com at my house i move and then something happens and it comes back and then now you moved and the board has changed so i can't really differentiate through what the other player uh is doing the car line my action space is actually quite small um even though there's 32 pieces and all that stuff there's only 8 possible moves in expectation that you can actually that are legit moves

20:55Francois Chaubard:in any given state there's only 8-ish moves you could do

20:59Ankit Gupta:let's just say in the beginning I can move all my pawns I can move my horses so that's 10 that's like not that much so this is extremely small and then my reward we can use the heuristic based approach or we can just say plus 1, 0 or minus 1 if I lose, plus 1 if I win and so this is very tractable

21:17Francois Chaubard:You say it's tractable even though there's a really big state space here. Why don't we talk about that for just a second? I think this is a really important point. I think when you say it's tractable, you're specifically referring to the action space being small because it affects the kind of like combinatorial expansion here. Should we talk about that for just a second? Yeah. Or maybe we can add go and then kind of contrast the two. Yeah.

21:36Ankit Gupta:So why don't we do that? Because I want to get to the alpha go, the way that they solve this. And you're right. So if I were to do this naively and I just took my ST plus one and I want to do look aheads, what I would do is I would take all of the actions I can take. So there's eight. So I would do action one, action two, action eight. And then each one of these, I need to expand it for all possible states. And so now I need to do cardinality S, which we just said is this huge freaking number. And so I have to do that eight times and I have to do it again. I have to do it again. So just looking forward, one move is quite intractable.

22:18Francois Chaubard:Although at the same time, everyone starts at the same starting position. And while it is a really large space, there isn't an infinity number of potential, there's actually a really small number of game boards, even four moves into the game, as opposed to a game where you could start in any permutation, for example, of initial game state and what a few states

22:42Ankit Gupta:style yeah so this is like definitely over uh um done because there's there's it's it's much much less than this in practice yes but just naively like looking at you know uh what possible game states could be uh as a rough math here but this is roughly the idea and then each one of these leaves i need to invoke my value function right uh which is the value of that state t plus one and so i have to do that all many times and we'll get this while if we go but like this ends up being estimating the leaf node uh because at the end of the day my policy atst i want to pick on the arg max of like the value of the the following action i guess it'd be yeah a exactly yeah the arg max over a of the value of the state of the of the end statest plus n let's say it's like that's the the main goal here um and so for me to do that i need to roll all this out estimate the value and then pick the best one and so this this quickly grows um however and we'll see this without we go which actually actually has an even bigger state space um so i think it's 19 by 19 um i think it's about right now so you have this 19 my 19 grid you can in each one it can be black white or or nothing there so i have three uh so let's do our star again so the cardinality of the state i think is going to be s uh two or three my ternary thing here to the 19 squared i think it's 361 something like that 361 um my transition same issue i don't know uh my action space is going to be 361 let's say so it's a good amount

24:31Francois Chaubard:bigger than chess much bigger but it's still not uh enormous yeah as we'll see in a second yeah

24:38Ankit Gupta:and so basically what they do they call this z which is kind of annoying but let's not r and it's the terminal it's the terminal when they won the game and they basically you know you have your trajectory which is um s zero a zero r zero um then all the way to the end of the game yep s n a n r n and if you won then all of these uh all the moves that black if black won all the moves that black did get plus all the moves that white did were minus one and they just that's how they

25:16Francois Chaubard:create their um their rollouts rollout refers to a taking n steps of play of all players one after another yeah of moves under a specific policy at the particular instantiation of it right so let's

25:33Ankit Gupta:just let's probably under this policy p theta t and we're going to overload t but like this is that instantiation we froze that model we froze that model and we play i think it's like 70 games and we like treat all of those and we're going to subsample a bunch of these state action results to train to update our policy in our world model, our transition model. And what it's actually doing is we take in anst we give it to some theta and it wants to output the probability ofst plus one being played which is our transition function and the value of the current state. And how do we get the value? And so the value of the current state, well, both of them are coming out of the model, but basically the loss function L theta is going to equal, and it's going to be really close to this control problem one, is we have some V theta minus this Z, which we'll just call it R here.

26:45Ankit Gupta:squared and then plus actually it's minus this pi which I'll explain in a second log p theta and I think they everyone includes this but they include it in the paper so I included there as well which is the weight decay yep and so so this is basically what our loss function is and then we'll play a bunch of these games. Let's try to be a little bit organized here. And so this is our setup. This is our architecture. And now once we train this thing, we do an insanely expensive task of test time planning. And so this trend in RL is just called test time planning. And a specific algorithm they use here for this is Monte Carlo Research.

27:37Ankit Gupta:And it's called MCTS. And so this is one of the possible things that you could do. It ends up working extremely well if you have small action spaces.

27:45Francois Chaubard:Yeah, so let's just like very intuitively talk about what MCTS does. A lot of people have heard about Monte Carlo research because AlphaGo was such a big moment. How exactly does that map into our star and value function and policy?

27:58Ankit Gupta:Yep. So I'll take this ST. This will give me 361 numbers that sum to one. and so i'll have some probability of uh of where these things are going to go for the of where my my opponent will play um here so these are like the sets of actions yeah so i'm here so that i have all myst plus ones i have 361 of these things um and then to be clear this is like action one action two all the way to action yeah 61 exactly yeah and the um we have to estimate the value of each one of these and so then we have to invoke the model all 361 times to give me values for each one of these things and then i will select i'll select it based on the the ucb the upper confidence bound which is this equation that is roughly something like um balancing uh my value function of ST plus one, which they're going to, in the literature would be called the Q value because it's actually the difference between a value function and Q value is just that I have the action as well.

29:12Ankit Gupta:So it's BST then AT. So we'll just call that Q value, which is my exploitation term. And then my exploration term will be something like, It's this funky square root of n. So it's the argmax of a of my q. And then I have this, which is the probability of this move being played, which we have from here of s, let's just call it s t plus one. And then I have this term, which is this sum over NSB divided by NSA. And what's the intuition behind this term? So these ends is the visit count during my MCTS process. So this whole tree, I'm going to.

30:10Francois Chaubard:So this tree can get really big, right? It's 361 per thing.

30:13Ankit Gupta:And it's a depth of 30.

30:15Francois Chaubard:So you can't visit every single week though. Exactly.

30:17Ankit Gupta:And so you want to keep track of which state did you end up in and what action did you take when you were in that state. And you want to make sure that you have good exploration, right? And so the way you keep track, the way you ensure that you have good exploration is you want to not just be greedy and always pick the highest value one because that could be very myopic. And so what you'll do is during this MCT process, you'll start this dictionary, which will be all zeros of the visit count of being in this state and taking this action. And then once you go through your first rollout, you'll go here.

30:56Ankit Gupta:All these things will be added to zero. You'll have some probability. We're going to bias it towards the higher probability of places to go. And then we'll expand those trees. and then we will update the counts that we visited this and that will basically reduce the amount of probability that we're going to select it again because this will reduce my exploration term. And if it's highly valued, then we're going to increase the Q on this because this is the expected value of going down this path.

31:29Francois Chaubard:The gist of it is fundamentally like you want to take the optimal-ish path but have enough exploration in this really expensive step you're doing here so that you are making sure you're getting a decent chunk of the other potential leaf nodes you could traverse to in these 30-step rollouts.

31:51Ankit Gupta:And so I'm going to do this MCTS simulation 800 times here. And then for all 800, I have to go through this whole process and I have to invoke the model at least 30 times to get through all here. and so that's you know 27 000 800 times 30 yeah so uh 24 000 uh invocations of the model to to develop this tree and then once i have it per step per step just to do one action into the game a lot of people don't understand that this is like you don't like store this mcts tree you like you throw it away after uh you you make the move um but once it's very expensive to develop this mcts tree And once you have it, the probabilities of traversal are actually extremely useful for training.

32:37Ankit Gupta:And then you end up biasing it and you train it with the MCTS tree, which is like a little bit seems like circular motion or something like that. Like, but you end up treating that as the pie that you'll train in your loss function. So we have the R of did we win or lose? We have the pie of what was the end result of this whole expensive process. And then at test time, we are going to do these 24 ,000 steps, every single move to pick the argmax that satisfies both exploration and exploitation.

33:17Francois Chaubard:in this case you know this still feels somewhat tractable though because the action space is small enough where this like kind of works exactly now like let's say hypothetically maybe we can draw like an imaginary go a game of go where it's like you know let's let's say this game of go was like a thousand by a thousand and so now you have a equals uh you know more or less a million and now this tree we're drawing here that has to take this has cardinal or like you know width I guess one million and there's like S0 through S one million and the number of steps you would have to take here presumably would have to be way more than 800 in order to get any reasonable kind of sampling of this and so you're probably multiplying the test time cost of doing a rollout or of doing a next step prediction astronomically if the game was even let's say you know this is only 100x bigger than the current

34:25Ankit Gupta:game running 50x bigger than the current game everyone was very excited about alpha go and at the time and what was this 2017 2016 uh everyone's very excited about this and the The important thing to pick up is that we did 800 MCTS simulations and to cover 361 possible actions on average. So that gives us about two samples roughly on an expectation for every single action. So here you need like 2 million of them for a similar depth. For a similar depth. And then that's still to do a depth of 30. I would still have to do this times 30. This would be 60 million invocations of the model. So that better be a small model.

35:03Ankit Gupta:Right. That's a lot. So, yeah. So that's to do a single action to be clear. Yeah. So exactly to do one action. So just imagine, uh, so why alpha go, uh, doesn't scale. Yeah.

35:19Ankit Gupta:To me, there's one, uh, the cardinality of the action space must be extremely small. If it's big, sad, uh, two, the, um, I need a perfect, uh, deterministic environment, right? Like this, this doesn't change. The rules of this game don't change. But the rules of the stock market change all the time. The rules to venture change all the time. The real world changes quite often. So homoscedastic.

35:53And real time.

35:54Ankit Gupta:If you saw the movie, the documentary, it's such an amazing documentary. I'd highly recommend it to anyone that watches it. The guy is sitting there for like 60 seconds maybe five minutes waiting for the computer to like decide and and it's kind of like imagine we were driving a car and like you took like 60 seconds to like turn the steering wheel everyone's dead like the whole car is dead right and so like you know uh now let's talk about uh robotics and self-driving car um and why this why that approach kind of can't scale yeah i think the really good

36:27Francois Chaubard:contrast here because intuitively i think in thinking through this exact star layout it actually really changed how I think about the kind of problem space of both of these two so like let's take self-driving car as an example this is one you know many people have started to experience for the first time because we have some self-driving cars that actually work you have waymo and tesla fsd and whatnot that seem like they kind of work so like let's maybe apply your same star framing here um I would contend that the state space of self-driving car is enormous and it's actually not intuitive to me whether it's more or less large than this one right i mean in a sense the chess and alpha ghost state space is already like more than the number of atoms in the universe or something to that effect right but like just to emphasize that here you know you are considering you know surroundings vehicle state yep uh like you know camera like weather I guess the point is like road conditions it's like massive this is massive for all intents purpose it is infinite for all intents purpose it is infinite correct yeah

37:34Ankit Gupta:and so is the space of pixels like you know like what can I put in an image I can take a picture image of anything yes true and so we're able to handle it and the same thing here where we compress from the board state we don't represent the board state we compress it with a com net so they have some deep some some deep com net that actually takes this state and converts it into a latent right and that latent compression is sufficient to kind of like do pattern matching do do some type of like symmetric symmetric uh equivariance kind of things and same thing with this and even better with jpa which we can talk about at the end there which is like basically taking some type of state space and doing all of our optimization in the latent space which stable diffusion did uh that worked extremely well which reduces our state space dramatically because i'm in some latent high

38:26Francois Chaubard:dimensional space so like the key thing there is that yeah despite this state space being effectively infinite we've actually gotten really good at compressing this yeah and we'll talk more about some of the tricks for how we actually do this in practice here but the tldr is you know where there's like 10 years of deep learning work that basically makes us extremely good at compressing that very fast exactly right exactly right t seems to have a similar problem as before right in fact maybe even more extreme there's like infinity other variables around you right

38:55Ankit Gupta:in some ways you'd think that it's this is physics newton's laws laws of motion should apply if i turn the steering wheel like this and i hit the gas i should be able to really easily model this but what is non-differentiable is that i have if i'm going into a circle right it's like the most the biggest issue that that we we faced in when i was doing self-driving car is like you're imposing your will onto maybe driving in india i think you're imposing your will onto the environment and like people just kind of adapt naturally like if you were doing a lot of motion you were going to collide and so that the optimal policy if you were doing strict newtonians here would be like don't move because anything you do you're going to crash yes but it's not true like that then we wouldn't function like cars wouldn't go down the road um and so you have to model the the You have to include other people in the environment and understand the embodiment of how your action will change other people's actions.

39:52Ankit Gupta:YC's next batch is now taking applications. Got a startup in you? Apply at ycombinator.com slash apply. It's never too early and filling out the app will level up your idea. Okay, back to the video.

40:05Francois Chaubard:Now let's talk about the action space. One way to look at the action space is that it seems relatively small. seems like well you know you turn the steering wheel left to right you hit the brake you hit the you hit the gas doesn't seem that big but like how big is it actually like how do we actually represent these action spaces when it comes to a realistic self-driving car scenario yeah i don't

40:24Ankit Gupta:know how they how they do this nowadays um they're doing a whole bunch of like bird's eye view

40:29Francois Chaubard:different things like that that's considered even just like a very simplified but what do you have

40:33Ankit Gupta:you have a steering wheel that you can turn left right you have a brake pad yeah and you have the

40:38Francois Chaubard:gas yeah and so this thing is like 365 degrees yeah so it's like a one to 365 let's say yeah zero to 365 yep and you let's just say you break this up into 10 different uh severities you're already uh even with just this oversimplified model your action space cardinality right is 365 000 so that's like a hundred x bigger than alpha it's in fact it's about the size of the example or in fact a decent amount smaller than the size we said exactly break and cts

41:11Ankit Gupta:and so yeah so 36 000 action space is very large and then even worse unless you're tesla we have a bunch of video of people driving cars we don't have video of like dash cams and like that like you actually don't have again only tesla has this of the action as well and so the things that you have access to that your trajectories are just like s t s t plus one yes s t plus two so there's

41:35Francois Chaubard:you're saying there's a decent number of these that's from like dash cam footage on youtube or something but not really that many either yeah relative and so if you wanted to do a self-driving

41:43Ankit Gupta:car and you didn't want to go spend a million dollars trillion dollars on going collecting all this data then you want to leverage this data somehow and this is going to be really applicable for uh robotics because we have a lot of uh videos of people doing things yeah right especially with ecocentric like we we have those videos but we what we don't have is the actions they take yeah yeah so this is like this is this is a sequence yeah you're showing here unless you're tesla unless you're tesla tesla has this so this is a huge competitive mode of like what do people do in that state and then so you can behavior clone to go from here to here from here to here go here to here etc but even then it's still very very difficult you have to it's not sufficient People think that like, okay, I have this, we have a self-driving car, right?

42:30Ankit Gupta:I mean, the amount of work that they're doing at FSD is like incredible. And it's, it's not generally available. Like you can't, you know, it's not Waymo level yet.

42:39Francois Chaubard:Would this be a good moment to briefly talk about model free versus model based RL? Yeah. I think that's an important distinction that's going to be relevant when we talk about more world models. Yeah.

42:48Ankit Gupta:So this is a perfect point. Um, so model free just means that my, my policy pie, uh, of a T given ST, uh, I have no world model involved. It's literally, it's literally doing what I said. I grab a bunch of these and I go from S to A, S to A. Just predict the next day. That's it. And that's, and this is largely called VLA. Um, you know, this is like giving us pretty good results. It's behavior cloning. It's all the, the, the, the stuff that it's not getting us to Rosie, the robot just yet,

43:18Francois Chaubard:But in many ways, it's the closest thing that just looks like the next token prediction from LLMs that seems to scale pretty well with natural language. I mean, it's not exactly the same thing because there's no action exactly. Picking a token is not exactly the same thing, but it's very analogous to that.

43:32Ankit Gupta:I basically take away the tokenizer head and I give it an action space and I collect a bunch of teleops data, you know, like this as the self-driving car does in Tesla. and I just take in the state, which is some image and or maybe sequence of images and then I'll output some action and that's it. And this is, let's say, model free because I don't have a model for the environment. And then now if I do model based RL, I have not just some pi, but I have also pi psi as well here. And so by including this, I can have a much stronger policy, but it would take a lot more time to perform inference because I have to do this full test time planning.

44:22Francois Chaubard:Just to remind us, that size is referring to this specific transition function, right? It's referring to this. You're saying this is specifically referring to a function ofst plus one givenst and action t. Yes. so it's like your ability to predict the next state you'll be in is the crux of it yep as

44:43Ankit Gupta:opposed to just directly predicting the actions yeah and the main thing that i believe is that this is required for agi this is what the human brain is at least in the way the human brain does yeah and let me go further in saying that like if you look at the um billions of years of evolution basically there's this thing called 10 million 10 million years ago called the great cortical expansion which you see the size of a brain just explode get bigger bigger bigger exponentially up until us and it basically stops and if the entire point of the neocortex is world modeling what happened is we started from vlas this would be like ants fish or whatever and fish yeah right just like very like you know lizard brain whatever i call it and then we develop this neocortex to like you know go from our motor cortex to actually simulate what's going to happen and that makes us just so much smarter.

45:36Ankit Gupta:And then once we get those samples, we can compress it when we sleep or otherwise with this hippocampal, shortwave ripple, whatever you want to call it. And then that helps us develop a better policy. And that marriage between the two not only helps us train on hallucinated examples, but it also allows us to test time plan.

45:57Francois Chaubard:I guess the kind of extreme case then of self-driving car is kind of general robotics. Yes. right so if you're like a humanoid company like figure or pi or whatever again same star setup yep i guess the gist of it is that a is now even bigger yeah right it is like i guess a very simple robot would be yeah how would you how would you put it in the action space like let's take a very

46:21Ankit Gupta:basic one if i take like my six axis uh arm yeah as your standard here that we're actually working on right now in stanford robotics center um you have two degrees of freedom two degrees of freedom two degrees of freedom uh and then you have another two for the end effector right and so that's a simple end effector not even like a not even like it's literally a one axis like you know you can rotate but you have the the the one axis yumi style uh thing so this is eight so you have 16 degrees of freedom yeah and let's just say that you do the 365 to 10 or whatever you know kind of thing i mean it's like 10 to the 16 it's like insane it's like yeah it's an insane number um and so much bigger than self-driving car um and even worse like getting tele-ops data is extremely painful and expensive it's not just like oh we'll just get some people in the philippines we'll give them like some you know things or whatever it's like totally totally doesn't work and nor is there yet something like uh tesla's fleet where there are cars deployed

47:21Francois Chaubard:that people are just using and they're not even necessarily realizing that every time they turn the steering wheel they're providing this this data set for tesla and then even worse you have

47:31Ankit Gupta:this like what's called cross embodiment gap and so if i were to like train this policy on tesla model x and i were to like put it on a tesla model three it wouldn't work no like it totally wouldn't work like all the so much so much of this uh the the way that if i were to break on a model three versus a model x the model x it weighs more it has different dynamics aerodynamics and things like that and so what's actually going to happen is very different like the degradation you have across cross across embodiments is very very very strong and clearly tesla's figured

48:05Francois Chaubard:various ways to get around that i mean they have these that roll out but actually even with tesla's new fsd today they don't roll out in all the cars at the same time probably for more or less that reason and in this case it's even harder now i mean you have bigger differences between embodiments than a model three versus y yeah and you have way bigger action spaces you have to

48:22Ankit Gupta:sell a model yeah uh lane mcintosh i played hockey with at stanford who now runs tesla fsd um i can ask him but i would bet money that they shard the data per model per uh car type yeah i wouldn't be surprised i just because that's what i would do there's no way that like you know i i would trust you know data that was collected on a model x on a model three i just wouldn't know way i would trust it.

48:46Francois Chaubard:Okay. So now that we understand the basic setup here and why the action space problem is so big, why don't we talk a little bit about how world models actually fit into this? You know, maybe first, you know, I guess what didn't work about the naive world models and how do we fix those? And then let's kind of talk about some of the newest world modeling techniques.

49:01Ankit Gupta:Cool. So like in robotics in particular, it's very hard to get these, this kind of trajectories that you want, that you kind of need to train for your VLAs and people spend up, you know, with a whole bunch of tele-ops data it's very expensive very expensive ideally what we would do is take like data like this from someone who is just like puts a camera on them and just like making sushi okay like i want to make a sushi robot um how do i do it give it to all the sushi chefs don't put anything in

49:27Francois Chaubard:their hands and just have them start cutting up sushi making sushi and ideally we would train it in that way you were describing of like somehow we would train a model just on these two yeah and

49:36Ankit Gupta:later of this yes afterwards and so the first real person that um you know went after this was jürgen smith humor please uh so he doesn't yell at us we have to we have to make sure we cite him uh but he has this really cool paper called world models uh very aptly named and it's basically he took these like um open ai gym classic uh games car racing and i think doom as well and then just like trained a model at that time was like an RNN. Um, he had some funky, uh, uh, zero order stuff in there or whatever. But basically the key premise was I can take an environment. I can extract a whole bunch of this type of data off of it.

50:18Ankit Gupta:I think he actually does actually this data, but we'll get into dream or where he does it in this paper in this way. And then, uh, trains a policy on only the synthetic data, the imaginated rollouts. And it actually performs well in the environment. This is the first time in my understanding that that actually happened and it actually works really well. And then...

50:42Francois Chaubard:So the key thing there, you can basically use this, if you have some predictive model of this in that case, and eventually of this, you can use that as basically a synthetic training set to train your policy model and then basically fine tune it on real data later.

50:54Ankit Gupta:Exactly. And which is just like a really powerful idea, especially since in robotics, the limiting step is access to large amounts of state action data. And so now the Dreamer series, so basically this publishes in May of 2018. Danijar Hafner publishes Dreamer 1, I think, in November of 2018. And then now he's been on this rampage for the last seven years publishing these papers. and Dreamer V4, I think, is the capstone of it where he basically does the same thing and he focuses on Minecraft and he trains a world model on this type of data and then injects action conditioning on a very small amount of data to get to this type of world model that has the action conditioning as well and then samples a lot from it and then trains a policy on those synthetic elements uh, imaginated rollouts.

51:53Ankit Gupta:And it's, the policy is so good that it's the first paper to mine diamonds in Minecraft. I'm not a good Minecraft player, but apparently that's extremely difficult. That's like next level difficulty. And it did it all on synthetic data, which is kind of crazy.

52:05Francois Chaubard:And the key on luck there, yeah, use synthetic data specifically on a model trained on just this sort of state transition type of thing. Yes. And this ends up being very convenient because it turns out we as a society have a lot of this.

52:18Ankit Gupta:Exactly. Yeah. All of YouTube, right? He does do a very small amount of data from app for to enable the action conditioning and that get that allows you to do this full uh simulated rollout but yeah it's true so we have we have youtube we have like flicker we have all these data sets online of like you know people doing things we'd like to use it and no one has really gotten that to work and then now that with this um these like video generated generation models we can take that data create a world model out of it add action conditioning post train it with action conditioning for some new task that is we want it to do chopping down wood or uh you know um making sushi or folding my bed or whatever it is only a few amount of examples and then we can train a policy on this in this neural uh simulation yeah and you

53:09Francois Chaubard:know we put out a video um about diffusion models very recently in flow matching i imagine that now ties very closely to this right ultimately the kind of current state-of-the-art best way to do this on basically infinity data that we have available and can keep generating is using say

53:24Ankit Gupta:video diffusion slash flow matching exactly yeah so like if you have your your c dance or your sora or one exactly all those models like basically the idea is now we have them and they're already trained and they're great let's do a small amount of action conditioning on them to get to this uh this world model and then we can sample from it a bunch and then train and this is exactly what wave uh did with gaia and gaia i think they raised 1.5 billion dollars to basically run with this idea for self-driving car um i think a bunch of companies um nvidia uh uh this this paper here uh is basically talking about doing exactly the same this dream zero for robotics um and what i

54:07Francois Chaubard:thought was really cool about this paper is that they do exactly this process where they have this joint model of state transitions and actions. They train it by first instantiating it with the open source one video diffusion model. And then it only takes them about 500 hours of tele-op data, which is basically exactly this, to get it to be pretty good. And they have a lot of clever tricks that allow it to be cross embodiment and work on unseen tasks with relatively small amounts of data. And it really is taking basically the exact concept, I believe, from the Dreamer paper and applying it specifically to these robot embodiments.

54:41Francois Chaubard:And it turns out it actually works actually better than I would have anticipated it working.

54:46Ankit Gupta:Yeah. So I think that this is basically the path to it was the path. I believe it was the path to get humans to be as good as we are genetically over the last 10, 20 million years of evolution. A bigger world model helps for training and for test time planning. And I think it'll be the same thing as true as for robotics.

55:08Francois Chaubard:What's also cool is there's a bunch of applications of this, the things outside of robotics too. I mean, there was a weather planning paper, for example, we were reading, it was GenCast paper, which I think applies a relatively similar concept in terms of how they model, you know, literally the world, the world's weather with something like this.

55:27Ankit Gupta:Yeah, we have to talk about the world model for the world. Yeah, so basically they do this exact same thing where, you know, the key unlocks for this whole thing was getting diffusion to work, in very high dimensional state spaces, like we talked about in the last lecture, and then learning to use that to action condition in the way that he's done. But they did this for the entire world with this exact same diffusion steps, which go from some, and they go back to two time steps, lag of order two, AR2 for the statisticians there. And then basically predict the next state of the world based on those things with this lingive and diffusion rollouts.

56:07Ankit Gupta:my big assertion is that um it was necessary for the human brain to develop world modeling i actually just saw this paper that i wanted to make sure to call out that that was so great uh out of uh university of washington where they say explicitly in the in the abstract each cortical area estimates both latent sensory states and actions and the cortex as a whole predicts the consequences of those actions. That sounds like a world model to me. Yeah.

56:37Francois Chaubard:Right? It's actually describing exactly these two equations here. Exactly. Where we're estimating both the sensory latent states and actions. I mean, I guess it's really the joint model that we showed earlier. Right. Is what he's describing here. It's exactly this equation we're showing.

56:52Ankit Gupta:Exactly right. And so if it works in us, it should work in robotics. And I think that that takes us the rest of the distance.

56:59Francois Chaubard:Why don't we talk briefly about latent world models, especially the JEPA concept? Because I think there's been a number of papers that use JEPA as an element of their, I guess, architecture. Why don't we just briefly introduce JEPA and how it fits into the current landscape of world modeling?

57:15Ankit Gupta:Yeah. In classic RL, you'll have, like, you know, if you do study Q learning, for example, you basically keep this matrix called the Q matrix. Yep. And it's going to be S by A. and so I have this S by A states and actions and each one I need some amount of counts of being in this state action and I take the average value of taking that action in this state and that's my Q value there and it's a little bit more complicated than that there's Bellman equation, all this backup, all this stuff like that so this scales horribly because as the cardinality of my state space gets bigger and the current action space gets bigger stuff I don't have enough time I become less and less sample efficient right in case of like robots or whatever state is like yeah it's

58:04Francois Chaubard:this whole thing we described earlier right it's absolutely massive because it has all of these elements in it couldn't really enumerate a huge grid and so the classic trick I mean since I took

58:11Ankit Gupta:you know C229 with Andrew Wong in 2012 is you do this stick a neural network on it exactly and you basically are just going to compress that state into some lower dimensional state space this is actually predates deep learning. We were doing stuff like this. I think my first paper was basically doing something like this, basically turning like a grid into like a bunch of like pyramids. And like, and the state was how much I'm in pyramid one or pyramid two or whatever. But anyway, the neural network can just do this. And so basically what the, the key idea in JPA, if I have an image one and I have image two and I have image three, I can do my world modeling my world modeling of ST plus 1 given ST and AT in pixel space and have this is let's say at time T, T plus 1 T plus 2 etc etc and I have to actually predict now the full image that's extremely expensive from a computation standpoint and also from like a sample efficiency standpoint What I can do instead is put this through some comnet, some encoder, some encoder, and then I'll get a latent for T and I'll have a latent for T plus one of a latent for Z T plus two.

59:39Ankit Gupta:And then I'll have from this, from ZT, I want to predict Z T plus one hat. and my goal is to make this and this make my loss function will be something very simple like I want to minimize this that's it now this doesn't work this collapses hard and so what happens is basically if it just predicts zero done just output zeros which the model will learn to do and I'm actually incorporating this into my current research right now and so what you need to do is something called SIGREG or this is one technique, VICREG is another, where basically I add this another term that basically says, I want the, over a large enough batch size, I want the distribution of ZT plus one to follow a Gaussian, you know.

1:00:39Francois Chaubard:It's kind of like a normalize it, like a batch norm type of trick.

1:00:43Ankit Gupta:I mean, not in the same case. And if it's zero, it can't be this, right? Because then this is non-zero. And so maybe I think that there's probably this or something like that. But basically this prevents it from modal collapse and it makes it do something good. And this is the most recent paper for the audience is LEWM, LE World Model, which is super, super great. However, to be completely frank, this is self-supervised learning, super great. It doesn't work that well. If you were to not do these techniques and there's, there's a bunch of other techniques that you can do, it will actually outperform much better that are, let's say, for example, if I'm going to do an LLM and you have like, you know, Francois likes sushi, which is definitely true.

1:01:30Ankit Gupta:And I tokenized this into a bunch of different tokens here. And this is token ID 6, 19, 28, whatever. And I look up the encoding into this. And that's going to be E1, E2, E3, et cetera. What you can actually do is have the LLM output. the LLM will take in these things and will output

1:02:05Ankit Gupta:the next token. And so it would be like, let's call it H. This would be the logits coming out of it, T plus one. And what you can do is actually have this be close to E T plus one. And a lot of people are playing with this idea and getting rid of the cross that should be lost entirely. and so if you were to do this it actually is a proxy for the cross-entropy loss and there is no cross-entropy loss and the cross-entropy head is actually very expensive and so this is very cheap and like this lady just grabbing it so people are playing around with this idea um and as a basically as a cheaper proxy for the cross-entropy loss so there's lots of different ideas on basically uh taking this jpa idea to not just pixels but to lns as well yeah interesting yeah So just to define what JPA is, it's joint embedding predictive architecture.

1:02:58Francois Chaubard:I think one of the things I find cool about this JPA idea is it feels like an idea we see over and over in deep learning. There's a version of this idea that's basically the stable diffusion idea. There's a version of this idea that in my company training graph convolutional neural networks to design drugs we use to do latent variable generation, for example. and it's like it's an idea that comes back over and over and then has this yeah various tricks that it actually takes to get it to work in practice okay now we have a pretty good sense for how world models work we have a pretty good sense for what the state of the art looks like if we trust this paper and it seems like these kind of work on robots too this paper is only from the end of uh end of last year this year and it seems like they have various methods that allow you to train on relatively small amounts of data that's tractable and pre-train on diffusion models so are we good or does it all work yeah this is 2016 26 will be the year of

1:03:52Ankit Gupta:the robot we're gonna have rose the robot in your house you know um yeah no i don't think so what

1:03:58Francois Chaubard:are one or two you know because there's lots of open problems remaining what are like a few open problems maybe we can emphasize here that the community can go emphasize working yeah so um

1:04:08Ankit Gupta:I think the first one is that pins doesn't really work. What is pins? Physics informed neural networks. So pins doesn't really work. This is physics informed neural networks. And so basically if like almost all of the self-driving car data looks like this, the car is driving down the road. And let's just say, for example, I have, you know, a house here. and I want to train the model on not driving into the house. And so let's say I put it into a state right here to drive into the house. What's going to happen is because almost all the data looks like this driving down the road, this will just turn magically into like a highway.

1:04:55Ankit Gupta:And I'm just like, ooh, don't worry, you didn't crash at all.

1:04:57Francois Chaubard:It basically needs like a ton of data not to do that, either from simulation for that to not happen.

1:05:02Ankit Gupta:In fact, I actually don't even know if because of the data distribution, there's no data here. There's almost all the data here. And like when you're training a neural network, it has a tendency to collapse if you don't keep the mini batch composition like very even over the, you know, over the class space or whatever you want to call it. But like you'd have to train on, you have to be very careful about your data mixing to make sure you get this right to solve this problem that no one really has. But even then, if you take just a simple thing like this, this is like the economic example. And I have some sine wave.

1:05:40Ankit Gupta:And I want and I have these as my X.

1:05:48Ankit Gupta:And I have these as my Y. So this is complete interpolation. No. That may mess this up. But Y like this. No. We can't get to like machine precision. We can't, what is it? I don't know. What is it? 180 minus 16 or whatever it is. We can't, we, we, the SGD will not get to zero effectively zero. So we'll always have some residual. And for us to be like a really good world model to simulate body interactions, like to, to simulate this, what's going to happen when I do this. And like, let's say that I'm trying to be LeBron James. Like there's like, I saw this one video of, um, Steph Curry dribbling about a basketball on a court.

1:06:25Ankit Gupta:And he just felt that there was a dead spot in the court. and he because he's so good and he knows exactly the physics of what's going to happen if i hit this you know the ball with this force like the ball is going to come back exactly this spot and it just didn't and he knew it wasn't him it was the the court and he found a dead spot in the court like that's how good the the human brain is at world modeling in my opinion i think it's an sgd issue i think it's probably an architecture issue i think sam altman just kind of came and just said that he thinks that there's definitely an architecture that's going to be more performant than the transformer i think he's right um i think the the transformer doesn't do compression uh in the time domain at all it just keeps around everything um so anyway so i think that the getting higher fidelity in the world model is extremely important one i think two seems like test time probably is going to make thing like adaptation exactly test time planning we the how quickly the human brain can you know in times of in sports and things like that when you're playing tennis i think you're a tennis player like how quickly we can adapt to what a player is doing and things like that we're not going to sleep and like retraining we're very quick to adapt to a new environment like the out of distribution prediction exactly really challenging and like one little data point we can like quickly adapt to that new thing and change um i think there's been a lot of papers uh on like basically estimating the friction coefficients and so like those can change over time if you go to a human environment or not for example like this this friction might change and that's important in control no um and so you need to estimate that very quickly and adapt and that these models just kind of don't have a mechanism to do it

1:08:06Francois Chaubard:yeah and then i guess there's like those practical speed elements of these right a lot of these are doing some sort of expensive planning step and we're doing some sort of like uh we're kind of hacking around it with this retraining process and synthetic data but even so like to really get maximum performance right now you'd want to do something that's closer to like the alpha go style rollout and that's extremely slow right the mcts process which can't happen um the other thing

1:08:34Ankit Gupta:that is pretty crazy about the way that the brain works is that like everything is kind of running autonomously and so like you'll you might be like in the middle of saying sentence one and be like oh actually no something else and so like what does happen there it's like type one and type two thinking are happening at the same time in some way and so like there's definitely uh you know some um really cool mix of these like heterogeneous models and like some are overriding others and like taking control of the motor cortex and like commanding the body to do a thing you know okay

1:09:07Francois Chaubard:But on the flip side now, we talked in the past video about the squint test and how we felt that autoregressive LLMs maybe don't pass the squint test. Why don't we reintroduce what the squint test was for a second? And then maybe let's think about whether this passes the squint test, despite all those limitations.

1:09:24Ankit Gupta:Yeah. And the squint test for me, I think, is like this comes from the Yann LeCun. We didn't need flapping wings to achieve flight. And to that, I say, well, we did need two wings. and like if I squint and I look at a bird and I squint and I look at a plane, I'm like, yeah. It's kind of similar. It looks right. Similarly, if I squint and I look at the human brain and I squint and I look at all these world models, we have like this VLA, this action policy, and that they're doing test time planning together and things like that. It's getting really close. It's much, much closer. It seems closer than an autoregressive LLM.

1:09:57Francois Chaubard:Like this concept of a world model of implicitly predicting future states and actions feels intuitively like what our brain is doing. And it seems like there's some, you know, neuroscience evidence.

1:10:08Ankit Gupta:I mean, I'm getting to the conclusion that I think that the brain is the optimizer, not the model. And that the brain emits, like has models that it invokes, but the brain is somehow also the optimizer itself. And so in that way, it doesn't pass the squint. Because like, you know, something magical is happening when you're sleeping. There's no intelligent species that we're aware of that have any amount of intelligence that don't sleep. And so like octopuses, dolphins, all those elephants, they all sleep. There's some reason for that. And that seems like a really, think about like the evolutionary recourse of sleeping.

1:10:43Ankit Gupta:Like you get eaten when you sleep. So like for the benefit of sleeping should be so much better to outperform that. So I think we don't have this idea of awake sleep in our current architecture, but I can imagine I'm like simulating, you know, compressed from the hippocampus, some like experience in the day, I'm like training on more of those examples. Right.

1:11:04Francois Chaubard:You're like collecting a whole bunch of these experience rollouts and then you're updating your policy function.

1:11:10Ankit Gupta:There's got to be something like there's this thing called shortwave ripple where like the hippocampus when you're sleeping like emits these spike trains that are actually reversed from when they actually happen back in through the both both the hemispheres and for like seven times. And then it like stops. So like there's something happening there that's very training something. Yeah. And if you don't sleep, then you don't have long-term memory. And so there's definitely a reason why we're training things that happened into our brain.

1:11:39Francois Chaubard:So where does that put us now? We have all this work happening with world models. How should we think about what's coming ahead in these next few years in the research community?

1:11:47Ankit Gupta:Yeah, I think that we're going to see a lot more of these world models in robotic policies. I think that's going to unlock probably full self-driving would be one of those examples that they can get the real timeness of it. It seems like that's coming. They can probably solve it with more compute to like have parallel things. And you probably don't need it for like most standard things, maybe like, you know, getting out of weird parking jams and like things like that would take us some time. Similar to the Rosie the Robot, which we've always wanted to have a Rosie the Robot to like, you know, clean up my room for me.

1:12:18Ankit Gupta:I think that like this feels like we're getting to good enough that we can pay up for data and compute to get to Rosie the Robot. It does feel like that. It'll be expensive to collect the data and do the Dreamer sequence of going from state to state and then getting the action conditioning to work. But, like, I feel like it should work.

1:12:38Francois Chaubard:Yeah. I mean, what's pretty cool is we see a lot of companies at YC working at every step of this, from the collecting egocentric data, collecting the tele-op data, training their own world models and action models, building new embodiments, and then making ways of adapting those embodiments. And it feels like this is the first year where you see demos where you're like, okay, this actually like kind of is starting to look like it's going somewhere. Yeah. And it seems like a very exciting year.

1:13:04Ankit Gupta:Yeah. So anyway, I think that there are real AI problems to solve still. We talked about pins. We talked about the real time issues. And then on the robotic side, there's real issues. Like it's amazing how effective our epidermis is in terms of we can detect tactile.

1:13:20Francois Chaubard:Oh, epidermis. Yeah, epidermis.

1:13:22Ankit Gupta:Our tactile, we can detect sheer force. we can detect temperature and it's everywhere yeah and so like versus you know like the we get like one little sensor that only does tactile we don't have the the friction component we don't have temperature we don't have all these the feeling we can't estimate coefficient of friction very quickly i can touch something and say oh this is smooth this is rough it doesn't we don't have any of that and if i numb your hands i actually had this experience um just recently if i numb your hands like you actually can't tie your shoes yeah so you can't perform control and And so like, yeah, if you like, you know, if you train enough on enough human data tying your laces, do I think you can do it with no feedback?

1:14:04Ankit Gupta:Maybe, maybe. But like, how much would you need if you did actually have the human like touch? Like, I think it'd be so much easier. Yeah.

1:14:11Francois Chaubard:Well, there's a lot of more research to do then. Yeah. Francois, thanks so much for joining us. Thanks so much for watching, everyone. We'll be back for the next episode of Decoded.

1:14:23Thank you.

From the publisher

Why do even our best AI models need tens of thousands of examples to learn skills that a human picks up in a handful of tries?Solving this problem is one of the great open challenges in modern AI. World models, which give AI an internal simulation of its environment, are one of the most promising paths forward.In this episode of Decoded, YC's Ankit Gupta and Francois Chaubard discuss the intuition and math behind world models, new research, and current applications in self-driving, robotics, and more.


Full Transcript: https://ycrootaccess.substack.com/p/world-models-an-intuitive-introduction

More from Y Combinator Startup Podcast

All 148 episodes
World Models, ExplainedY Combinator Startup Podcast · 1 h 14 min
Listen in VO