World Models and the Future of Spatial AI with Justin Johnson - #775

1 Sep 2026 · 1 h 6 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that AI’s next frontier is “world models”—systems that understand, simulate, reconstruct, and plan through environments—moving beyond language-only scaling. It discusses what “world model” means (no single agreed definition), how POMDPs (agent, hidden state, observations, actions, reward) relate historically and technically, and two implementation paths: explicit 3D (e.g., Gaussian splat representations) vs implicit 3D/pixels-only real-time generation. It also explains why Gaussian splats integrate well with neural networks (smooth, differentiable rendering) and how World Labs builds both explicit and implicit systems.

Guest backgrounds

Justin Johnson is co-founder of World Labs (with Fei-Fei Li) and an associate professor of computer science at the University of Michigan.

Key claims

World models are needed for robots/agents acting in the world or for navigable virtual worlds; “state” is an abstraction relative to the task; Gaussian splats are representations, not the learned world model itself; consistency can come from explicit 3D or from large-scale training.

Notable examples

Marble (Gaussian splat worlds from image/text/video inputs via a 360 panorama intermediate); RTFM (real-time frame model generating pixels without explicit 3D); traffic-light observation vs full atom-level state; behavior cloning in robotics as a non-RL training route.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding World Models

1:09 to 2:52

Justin Johnson explains the concept and importance of world models in AI.

“I'm Sam Charrington, and this is the TwiML AI Podcast.”

The Evolution of World Models

2:52 to 6:15

Discussion on the shifting definitions and implications of world models.

“And I think that's causing part of the confusion, right?”

World Models as Theory Builders

6:15 to 9:40

Examining the philosophical and scientific aspirations of world models.

“Like they all have, like we should probably build all of them as a community.”

POMDPs and Reinforcement Learning

9:40 to 14:00

Exploring the relationship between POMDPs and world models in reinforcement learning.

“Then maybe the video I want is like Einstein giving a lecture explaining like the new theory of quantum gravity or something like that.”

Understanding World Models in Reinforcement Learning

14:00 to 17:33

Learn about the relationship between reinforcement learning and world models, including the concept of POMDPs.

“that tells it something about the world.”

The Nature of State in World Modeling

17:33 to 23:01

Explore the concept of state in world models, including distinctions between ground truth and learned states.

“So then it like, that's the kind of the original technical setting where the term world model was used.”

Exploring 3D World Generation Techniques

23:01 to 28:00

Discover the differences between explicit and implicit 3D representations in world generation, including Gaussian splats.

“And that's why, that's one of the reasons why I think a lot of people are attracted to this, this area right now is that language modeling feels like we kind of have a set of best practices that work pretty well.”

Understanding Gaussian Splatting

28:00 to 29:40

Learn about the distinction between Gaussian splatting and world models in AI.

“is like, how do you draw the line between, you know, that thing and a world?”

Ecosystem of Gaussian Splatting Tools

29:40 to 31:50

Explore the tools and file formats associated with Gaussian splatting.

“And that model happens to output Gaussian splats.”

Comparison with Traditional Graphics

31:50 to 36:20

Understand how Gaussian splats differ from traditional triangle meshes in graphics.

“And an SPZ is a little bit more like a JPEG, where you're going to compress some parts of the data.”
Show all 22 chapters

Consistency in AI Models

36:20 to 38:50

Discover the importance of data consistency in AI model outputs.

“splats and why people got excited about them over the past few years is because it's a graphics representation that integrates with neural networks really cleanly.”

The Marble World Model

38:50 to 42:00

Learn how the Marble model generates 3D Gaussian splats from user inputs.

“training costs, then I think the implicit pixels only approach is very appealing.”

Generating 3D Models from Images

42:00 to 43:50

Learn how models can generate 3D representations from 2D images and text prompts.

“like try to give a plausible completion for everything else in the world that's not visible in the input image, right?”

Training Data for World Models

43:50 to 45:52

Discover the importance of diverse training data for building effective world models.

“And that's kind of the major difference between, you know, having these generations backed by a big world model versus, you know, classical Gaussian splatting just fitting to these thousand images that I happen to have.”

Taxonomies of World Models

45:52 to 47:48

Explore different types of world models and their functionalities in AI applications.

“Yeah, the kind of rawest output from this model is splats in some way.”

Categories of World Models

47:48 to 53:10

Learn about the three categories of world models based on their outputs: planners, renderers, and simulators.

“So there, like I said, like we talked about a couple of times, I think there's a lot of different flavors of things that people are training and calling world models that look pretty different from the outside.”

Unified World Models

53:10 to 56:00

Understand the concept of unified world models that integrate various functionalities.

“like then you can do, you know, measure distances between points and that, you know, it's a representation, but it's also the observation, like.”

Understanding Unified World Models

56:00 to 58:09

Learn about the concept of unified world models and their interrelated capabilities.

“the state is a set of theories about the world.”

Architectural Evolution in AI

58:10 to 1:00:22

Explore the current and future architectural needs for world models in AI.

“So I think like that's kind of where the field is going to get the next couple of years.”

Tokenization and Context in World Models

1:00:23 to 1:02:39

Discuss the implications of tokenization and context length in modeling worlds.

“Like, yeah, in principle, there's problems you need to fit whole book, cram whole books into your context.”

Incorporating Geometry into Deep Learning

1:02:40 to 1:04:22

Understand the role of geometric deep learning in spatial AI applications.

“But I think the lesson we've learned over and over again in deep learning is that you want simple representations and then scale up the model behind the simple representations.”

Resources for Learning About World Models

1:04:23 to 1:04:59

Get recommendations for resources to dive deeper into world models.

“Are there any canonical resources that you like to point folks to who are new to the space and want to dig in more deeply?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Justin Johnson:The race to build more capable AI isn't just about making language models bigger. Increasingly, it's about world models, systems that understand space, predict how environments change, and act in the world around them. Justin Johnson is helping shape this shift. He co-founded World Labs with Fei Fei Li and is an associate professor of computer science at the University of Michigan. I asked him why so many researchers think world models are AI's next frontier.

0:26TwiML AI Podcast Host:There's a shared like low level belief among many researchers in the field that there's something that language models aren't doing, but that there's other kind of models that we should be building. They do other kinds of things. And that's something around understanding the world, generating worlds, simulating worlds, reconstructing worlds, planning actions through worlds. These are all capabilities that we want to build models to have. And why do we care about this is because we want to build systems that are not just stuck in a terminal or stuck as a virtual agent. You want to have visions of AI systems that are going to be robots that are out in the world acting in the world.

0:58TwiML AI Podcast Host:Or we maybe want to build virtual worlds and live in those and simulate interesting things there. So all of these are capabilities that really don't feel like they're falling naturally out of the language modeling paradigm.

1:09Justin Johnson:I'm Sam Charrington, and this is the TwiML AI Podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.

1:32Justin Johnson:Over the past few years, one idea I think that has been espoused is essentially the idea that a world model is like an emergent property of either language models or diffusion. Like if you, you know, diffusion can, you know, illustrate some physical properties of the world or, you know, language models can tell stories that, you know, kind of seem like they know something about the world. and I guess there's a couple of questions emerging in my you know my question here one is like maybe it's asking for a concrete definition a concrete definition of a world model because I think you know early on in those conversations world model was really talking about you know these models having like foundational knowledge of the real world, like the world that we live in.

2:38Justin Johnson:More recently, you know, the world model conversation, I feel like has shifted to talking about being able to create artificial worlds, but have them be self-consistent and navigable and, you know, properties like that. So, you know, A, kind of, I'd love to hear you you know elaborate on the relationship and if you see kind of the same shift in the terminology but also this idea of like you know emergence and you know do we need new things or like if we throw enough you know data compute etc at the models that we have you know does that get us there or more likely why won't that get us there I think there's a lot of interesting questions to

3:20TwiML AI Podcast Host:unpack there one I mean the biggest one is just like let's get it out of the way there isn't a clear definition of world models that everyone in the field agrees on. And I think that's causing part of the confusion, right? Like there is not like a thing where we can say a model that has X property or produces X kind of input and produces X kind of output, like is a world model definitionally. I think we don't, we don't have that crisp definition as a field of what we mean, which leads to the confusion. But to your point, I think there's a couple variants of this that feel like they're like, I think there is some notion of like implicit world knowledge that you dimension that there are other kinds of models that produce certain kinds of inputs and outputs.

3:58TwiML AI Podcast Host:Maybe if a model that is producing text, if it produces the right kind of text, the only way it could have produced this kind of text answer is because it knows something about the real world or because it's modeling some kind of implicit world internally in its neural network weights. Or similarly for video models, right? Like if I am able to generate a video that is super photorealistic and has detailed and like has physics and has water running in very particular ways, like maybe implicitly the model must have been modeling something about a real world in order to generate an output of that kind.

4:29TwiML AI Podcast Host:So I think that's kind of this notion of an implicit world model that you did this other, like maybe you could be doing any number of tasks, but there's certain kinds of answers that you could give that would require implicitly modeling something about the world. But then I think there's like two other threads there that are really interesting. I think the term itself, world model, actually goes back to reinforcement learning literature. And there there's a specific technical definition related to POMDPs that maybe we'll get into later. But there's a specific technical term of world modeling that goes back to reinforcement learning for quite some time.

5:04TwiML AI Podcast Host:And then there's another thread I think that ended up is like, you know, we've been talking a lot about generative models the last couple of years. And an image model is a model that produces images. A video model is a model that produces videos. And then maybe a world model should be a model that produces worlds. And then what does it mean to produce or generate a world? So I think we've got these like three kind of different threads that different people in the community use is around like some notion of implicit world modeling that I can answer really hard problems, but I must be modeling something about a world to answer the word, to get the right answer.

5:34TwiML AI Podcast Host:Or like there's the specific, you know, RL, RL formulation of world model. And then there's the, there's like a generative model that creates or generates worlds.

5:43Justin Johnson:Do you see this implicit world model as, you know, just another definition or like an incorrect set of beliefs? Like when you hear that, do you, you know, do you feel like those are world models? Do you think that's a valid way of thinking about world models? Or do you think, you know, world models have properties that aren't really characteristic of, you know, the current models that we are talking about, you know, vision, language, etc.

6:12TwiML AI Podcast Host:I almost think like, I think they're all valid. I mean, the hard part is that I think all three of the things I just said are very interesting systems. Like they're very interesting models. Like they all have, like we should probably build all of them as a community. But we should probably come up with better terms so that we don't confuse each other by calling different systems by the same term. And I think that's the thing. So I think like maybe in like the academic literature in the last year, like the world model has more coalesced into a particular flavor of a real-time interactive video model.

6:44TwiML AI Podcast Host:And that's like been more commonly what people call world models in the literature the last year or so. But I think all of the other properties that we just talked about are really interesting and useful. And like, especially the implicit world model notion, I think that that sort of can be applied to anything. And like, no matter what kind of beta you're processing, no matter what kind of system you're building, I think if it gets to some level of interesting complexity, then it's going to end up with some notion of implicit world model somewhere in the system. And that's really interesting.

7:12Justin Johnson:One of the ways that that occurs to me is like thinking about it in the context of, you know, the question that we grappled with around or even so grapple with around large language models is like, do they really understand language? Do they, you know, is next token prediction? Is it just faking an understanding of language or is it an understanding of language? And I think, you know, for the most part, maybe we've moved off from that question and said, these things are so amazing. Like, does it really matter? Um, and so from that perspective, like if we apply that to the world model question, like if, you know, those types of models get so good at, you know, generating consistent to some underlying world, uh, results, maybe it doesn't really matter, but it, you know, I still find it an interesting, you know, if, if, if nothing else, like philosophical question.

8:09TwiML AI Podcast Host:I, I agree. There's an interesting philosophical question in there, but at some point it's sort of unanswerable, right? Like as a scientist, you want to be like, at least for me, I like to think about what are things that I can measure about this system? What are like concrete questions that I can ask or falsify ideally about a system? So like, but I think there, but I think the, another thing you're getting at is there is like another notion that I think some people sometimes have when talking about world models, which is pretty different. And I don't think we know how to get there. And that's more like world model as theory builder.

8:40TwiML AI Podcast Host:Right. Because we have this notion that as humans, we're kind of like traversing this really complicated world all around us and having these experiences. But we're not just like letting the photons fall in our retinas and like letting it happen. Like we're constantly building theories in our mind for what's happening. Right. We come up with these great theories about how does gravity work and how does physics work and how does fluid dynamics work. And it's not just that we observe these things. We actually end up with these very compact and very powerful theories about that explain all the mechanisms behind what's what's happening out there.

9:11TwiML AI Podcast Host:And I think part of the part of the like holy grail question around AI is how do we get machines to do that kind of a thing, too? And maybe there's this notion of like, well, a world model should not just be directly thinking about observations, but should be building deep explanatory theories about the world. And I think that one's really, really hard. And I don't know that anyone has a great angle on how to get there. But I think that's kind of a slightly different notion of the question that sometimes is mixed in there.

9:40Justin Johnson:Yeah, you could argue that even before Worlds, it would be great if we can get, for example, a model that can build a deep theory about language, for example, or any other domain that is currently handled by the types of models we deal with currently. you know language you know graphic arts uh yeah interesting interesting but but there there's

10:05TwiML AI Podcast Host:something like you know what if i had this like perfect i mean an lm is kind of like imagine like a perfect lm maybe lm is the wrong example but like suppose you had a perfect video model that could generate any video you asked for um but like maybe what i wanted wasn't actually a video what i wanted as a human as a scientist was i wanted to learn something about the world And like, even if I can generate video of any kind or like of any structure, like I didn't learn what I wanted about the world. Then maybe the video I want is like Einstein giving a lecture explaining like the new theory of quantum gravity or something like that.

10:38TwiML AI Podcast Host:And then it's like not actually the pixels themselves, not the capability to generate video. That was what I really wanted. I wanted to gain some new understanding of the world from this model somehow. And that's a really hard one.

10:49Justin Johnson:Yeah, and I don't think LLMs are a bad example, right? If an LLM had the ability to generate theories, you could argue that it wouldn't hallucinate because it would think more deeply about the relationship between the things that it's generating and would kind of self-correct or could self -correct.

11:12Justin Johnson:So we talked about kind of implicit world models. We talked about generative world models. There's this state machine interpretation that you mentioned, POMDP, partially observable Markov decision processes. Talk a little bit about that and how that history kind of plays into the way you and others are thinking about world models.

11:36TwiML AI Podcast Host:So there's this there's this abstraction called partially observed Markov decision processes or POMDPs that goes back quite a long time. that is a really nice mathematical formalism for thinking about how agents can interact with worlds. So then there's basically like two, you imagine like you basically decompose your system into two parts. One is the world, and that's like everything that happens around you. And then there's an agent, and an agent is something that can take actions in the world. They can move around, they can maybe pick up objects, they can do things in the world or to the world.

12:07TwiML AI Podcast Host:So you sort of partition the whole universe into agent, which moves around and does stuff, and then world, which has stuff done to it by the agent and also maybe evolves in time. So you've got the world and the agent. And then you talk about this formalism of how do the two interact with each other. Then the agent is going to take actions. And in different situations, your vocabulary of actions might be different, right? Like maybe you're a robot and I can actuate my motors in a certain way. Maybe I'm in a video game and I can push buttons on the controller. So in different situations, the agent might have available to them different kinds of actions.

12:42TwiML AI Podcast Host:But whatever the situation is, when an agent makes actions on the world, the world will be changed in some way. And the way we need to know that is we say that the world has a state internal to it. And the state kind of defines everything that makes the world what it is. And the state might be very large and very complex. It might not be understandable. It might be very, very large and high dimensional. But it kind of is the full explanation of what's happening in the world. So then the agent makes actions on the world. And then because the action changes the world, that means the action will cause the state to change or transition sometimes, somehow inside the world.

13:19TwiML AI Podcast Host:But now the state is so big, like it's something very complicated. Maybe it's, you know, you can't directly observe it. So then what the agent gets back are observations. And then observations are somehow some kind of low dimensional projection of the full world state. And again, what that means is different in different contexts. Maybe as a human, the observations we get are the images that we see in our eyes, the sounds that come into our ears, the feelings of touch that we feel on our bodies. Those are all the sensory signals that we get. And those sensory signals tell us something about the world, but they only tell us something very sparse and local about everything that's happening in the world around us.

13:53TwiML AI Podcast Host:So then the PoMDP loop is that, you know, you have an agent does actions to the world that causes the state to transition. Then based on the state, the agent gets an observation that tells it something about the world. And then the agent is going to, this is going to happen in a loop over and over again as the agent tries to do things in the world.

14:09Justin Johnson:That sounds a lot like the setting for reinforcement learning. The agent's operating in the world. It's making observations. There's some reward associated with its actions, et cetera. Is the implication then that you need a reinforcement learning type setup in order to have a robust world model? Or what's like the concrete relationship between POMDPs and world models?

14:37TwiML AI Podcast Host:You got me there. So like the important, you know, technical piece that I left out of the initial definition was the reward, right? So like in sort of the more, like this whole formalism was developed in the context of reinforcement learning. And there it's like, well, the agent has some goal that they're trying to achieve. And how do they get signal about whether they're achieving the goal or not? Then they get some reward signal. So then the idea is like, well, the agent wants to try to take actions that will maximize its reward. And once you go to that level, like that's exactly the reinforcement learning setup.

15:06TwiML AI Podcast Host:And, you know, that's where this formalism comes from. But the reason I chose not to talk about the reward a moment ago is because I think we can take that original abstraction of the PoMDP and then pull it out and use it in other contexts. So I think this notion about thinking about agents and states and observations, that becomes applicable and useful in a lot of other contexts, even outside of reinforcement learning specifically. And reinforcement learning is maybe one setting in which this formalism was originally developed and is super useful, but we can apply it elsewhere as well nowadays.

15:35Give us an example of how it's applied out of that context.

15:39TwiML AI Podcast Host:So one example, one like pretty concrete example might be behavior cloning in robotics, right? So like say you're building a robot and your goal at the end of the day is I want to build this robot that's going to go around the world and do stuff and maybe make my bed for me or clean the kitchen or whatever it is, then once that robot is out there, it's an agent, it's interacting in a world, the world has state, the robot is getting observations about the world. So it's kind of operating on that PoMDP loop once it's trained. But there's a question of like, what was the training signal? And you could have had that thing trained via reinforcement learning, where like every time it got a reward, like every time it took an action, it got a reward, or you could train it via behavior cloning, right?

16:19TwiML AI Podcast Host:Like maybe I had a large supervised data set of like in this data set, when you receive this observation, you should take this action. And then you could train a supervised learning model that would be very different, that wouldn't have an explicit reward signal. You would use like gradient descent and some supervised loss. So then even though you end up with a system at the end of the day that kind of operates in this PoMDP-like loop, you could have had a training objective that was pure supervised learning that didn't require reinforcement learning.

16:44Justin Johnson:So the summary is that there's two parts of this, two parts of PalmDP. One kind of captures the relationship between an agent and the world and the idea that the agent, you know, can observe things that reflect some abstract state, but it can never really know that abstract state. And it needs to operate on its, you know, observations. and the, you know, reward and, you know, training is all about how it translates those observations to actions, but that doesn't necessarily need to be about reward maximization per se. It could be about other things.

17:28TwiML AI Podcast Host:And then the reason why this is connected to world modeling is because this is where the term originally comes from, right? Like the kind of original technical definition of a world model is then, you know, if we're in this PoMDP setting, then a world model is something that inputs the world state, inputs the action the agent takes, and then predicts what's the next world state. So then it like, that's the kind of the original technical setting where the term world model was used.

17:52Justin Johnson:And I don't know how, you know, with that in mind that, you know, this is, you know, in a sense, kind of, you know, broader context or, you know, of historical interest and not necessarily about where we're going with world models today. but I did have a question about state. A lot of times we hear state talked about as, you know, a representation of, you know, almost like the observation is to the state in kind of this formalism of POMDP.

18:24Justin Johnson:Often state is some lower dimensional representation of some other, you know, thing that is the ground truth. is that a distinction that is you know useful for us to explore do you think maybe a little bit

18:41TwiML AI Podcast Host:I think you know there's actually maybe two different notions of state that people talk about sometimes that are getting conflated so that might be useful to try to unpack one is the notion like I think in my mind when I say state I'm thinking about sort of the ground true state of the world that's like the complete description of everything in the world that would

18:56Justin Johnson:be required to answer any question about it this is a another way to come at it or a question that popped up for me earlier that I'll surface now. You know, when you were talking about state earlier in the context of the loop, I was thinking like an extreme example of the relationship between state and observation is like you're sitting at a traffic light. The observation can be boiled down into like a one-dimensional, you know, red, green,

Read the full transcript

19:25TwiML AI Podcast Host:you know, red, green, yellow,

19:26Justin Johnson:or two-dimensional, whatever. um uh the um you know whereas the the state and by the definition of the pom dp is like you know every atom in the world and my question is really when you're talking about world models does the state necessarily consist of every atom in the world? Or is it just a representation? And maybe that's another way of coming at this question.

20:03TwiML AI Podcast Host:I think so. So I think it's all about the question of like, what's the abstraction that makes sense for the problem at hand, right? And like, even physicists will use different representations depending on the question you're asking, right? Like for some problems, I want the wave function of this quantum system. And that's kind of the best version of a state. for other systems, maybe it's macroscopic and quantum effects don't come into play. And then you could think of it as a Newtonian system. And the state is now this collection of particles, what are their masses and what are their positions and momentums?

20:30TwiML AI Podcast Host:Or maybe you're talking about like a chemical system. And then maybe it's just like these solutions with these ions in this concentration. Or maybe you're talking about thermodynamical system. And you think like, you know, it's an idealized gas. And that's enough to describe the state of the system.

20:44Justin Johnson:So it's never every atom in the world.

20:46TwiML AI Podcast Host:No, no, I think it's always I think I think even if we're talking about like ground truth state, it's always like relative to the abstraction, like the abstraction that's necessary for solving the problems you care about. But I think there is another notion, which is a learned state. And that's another interesting thing that people are now talking about is that, you know, the ground truth state, like we're probably not going to have direct access to in the physical world, right? Like even if we think about like every atom or like, you know, an idealized gas, like in a lot of situations, you're just not going to observe that in the world.

21:14TwiML AI Podcast Host:So instead, there's another question of could I have a model which I learn, which is going to learn to take observations and predict some kind of, you know, neural vector. And that learned vector, even though it's something learned by the model, it behaves as if it were a ground truth state, where I can, you know, predict observations that would result from that state. I can imagine how that state would transition in response to actions. But it's kind of a learned internal, you know, vector representation from a neural network. And it's not like a ground truth state from physics, but it kind of behaves like that ground truth state.

21:47Justin Johnson:Is that analogous to the idea of model-based and model-free RL?

21:52TwiML AI Podcast Host:A little bit, yeah. So if you're doing model-based RL, then you have some explicit model of the state. And that's where you have a world model, right? So then I have an explicit part of my RL system that in a model-based approach, I have a component that tries to predict the next state based on the current state in the action. Or in a model-free approach, then I don't have that explicit notion. Maybe I'm just getting the observations, feeding the observations directly to the model, and it outputs actions. and there's no explicit modeling of state. But, you know, maybe then going back to implicit world models that we talked about before, you know, maybe that neural network, if it's big and complicated enough, is doing something like state modeling inside its neural network weights if it's able to predict really good actions in the right circumstances.

22:31Justin Johnson:I think what we're doing here is illustrating the ambiguity of the term state. It can either be, you know, whatever the representation or abstraction of the world is, or in some settings, it could be a surrogate for that, that the agent creates and predicts and predicts against in choosing actions.

22:56TwiML AI Podcast Host:And I think the jury is still out on which of these is going to be the right approach, but that's, that's exciting, right? And that's why, that's one of the reasons why I think a lot of people are attracted to this, this area right now is that language modeling feels like we kind of have a set of best practices that work pretty well. And, you know, there's things to be explored, but there's a recipe that we have that works pretty well. But if you want to talk about like, how do we build, how do we model environments? How do we model worlds? How do we model agents that interact with worlds? There's a lot of kind of basic, you know, technical questions or like basic engineering approaches that, you know, you could do it this way or you could do it that way.

23:30TwiML AI Podcast Host:And we don't really know what's going to be the best way. And that's a really exciting place to be working in and doing research.

23:34Justin Johnson:So let's switch gears a little bit and talk a little bit about the graphical side of things. You know, as you mentioned, when we talk about world models, a lot of that conversation is focused on the generation of graphical worlds. And one of the kind of foundational technologies in that space is this idea of a Gaussian splat. So talk a little bit about, you know, maybe, you know, broadly the way folks approach, you know, this world generation and, you know, the role of Gaussian splats, how that has evolved over the past few years as well.

24:14TwiML AI Podcast Host:I think there's actually a top level fork there that we should call out even before we get down that road. And that's like, you know, are you doing explicit 3D or are you doing implicit 3D? And the Gaussian splats are an example of an explicit 3D approach where I'm going to generate when I generate my world or model my world, I'm going to have some explicit 3D representation of that world. And then I'm going to do things with that explicit 3D representation of the world. And one of those things might be render it down to pixels that you see. There's a different approach, which is sort of implicit.

24:45TwiML AI Podcast Host:I'm going to have a model that's directly generating pixels. Like it's a video model. It generates pixels. It generates observations directly. And it never goes through this intermediate, explicit 3D representational bottleneck. And I think both angles have developed a lot over the last couple of years. So then, you know, that's an interesting fork in the road, even talking at that point.

25:06Justin Johnson:It is. And do you, so World Labs, for example, takes the, well, is it even fair to say that you take the Gaussian splat approach? Like in Marvel and some of the work that you, you know, publish and demonstrate, you use that approach, but do you feel that that is foundational? Or is it just kind of what you have, you know, shown thus far, but you're open to other ideas?

25:33TwiML AI Podcast Host:Well, I think we've actually done both already, right? So it's true. We put out a product called Marble, which takes the Gaussian splatting approach. And there, like what Marble does is lets a user input an image or a sequence of images or a text prompt and then generates a world where that world is represented as a set of Gaussian splats, which is an explicit 3D representation. And then once you have that Gaussian splat representation, then you can render it to view nice images of the world. Or you could imagine compositing objects into it or export it to a graphics engine or something like this.

26:05TwiML AI Podcast Host:But then we put out a research blog post late last year called RTFM, which takes a very different approach called the real-time frame model. And this contrasts with our Marvel approach in that it kind of goes more this straight-to-video, pixels-only direction. There is no explicit representation of the 3D world. You have a model running in real-time that generates images in real-time in response to user inputs. So the user asks to move around and the user sees an image of the world moving. Like you say move left, you see a video of yourself moving left and this all happens in real time. But under the hood, there was no explicit 3D representation.

26:45TwiML AI Podcast Host:It's just frames being generated from a model in real time. So WorldLabs actually has been approaching both directions. We have Marble, which is our explicit Gaussian splat-based world model. And RTFM, we have as a more implicit real-time frame generating world model.

27:00Justin Johnson:You know, and this raises a question for me that goes back to kind of our, you know, very foundational conversation about what is a world model.

27:13Justin Johnson:I'm thinking about like Meta's initial Gaussian splat, you know, work and demos and the tools that I think have popularized Gaussian splats before, you know, they became kind of foundational to this world model idea. You know, they allowed you to, you know, take a bunch of pictures and they would like stitch those pictures together, um, using the term stitch very loosely to create a, essentially a navigable 3d model. And even, you know, prior to, I don't think it used Gaussian splats, but like Apple's AR kit and those kinds of, uh, things produced like these modestly navigable 3D, I'm trying not to call them worlds because the question is like, how do you draw the line between, you know, that thing and a world?

28:10Justin Johnson:I want to say like static versus dynamic, but that doesn't seem right.

28:14TwiML AI Podcast Host:No, I know what you mean. There is a really interesting distinction here that I think a lot of people get confused about with Gaussian splatting, with Gaussian splatting in particular, right? because a Gaussian splat is actually just a particular, it's just a representation. Like a Gaussian splat is just like, it's basically a 3D point cloud. I have a bunch of points in space and each of those points has a position and a color and maybe a couple other properties attached to it.

28:36Justin Johnson:And if your point cloud is large enough, you can move around in it from the perspective of a point.

28:43TwiML AI Podcast Host:But the interesting question is like, where did that Gaussian splat point cloud come from? And there, there's like two big divides that people get confused because where Gaussian splatting originated as a technology is around reconstruction. So there the idea is like, I'm going to take a lot of views of a space, like hundreds or even thousands of views that cover this one space in extraordinary detail. Then I'm going to fit a Gaussian, like a 3D point cloud Gaussian splat representation that explains the images that I saw.

29:11Justin Johnson:Well, put like that, the distinction is obvious. Like it's not just world, it's world model.

29:16TwiML AI Podcast Host:Exactly. There's no world model in there, right? There was no big, powerful model in here that learned on a ton of data. These kind of optimization-based reconstruction approaches to Gaussian splatting are like, I've got a thousand images, the whole universe is like these thousand images, and I'm fitting a Gaussian splat to match these images. There's no generalizable knowledge here. And that's actually very different from what we're doing in Marble. In Marble, we have trained a large, powerful model that's been trained on a lot of data of various kinds. And that model happens to output Gaussian splats.

29:46TwiML AI Podcast Host:So that that the marble world model is the model that knows how to model worlds. It inputs images, it inputs text and it outputs Gaussian splats. And that's very different from this situation where I'm just going to kind of dumbly fit, you know, a point cloud to this thousand images and there's no notion of a model that learned a ton of data.

30:03Justin Johnson:And we touched on this before we started rolling, but I always thought of Gaussian Splat as kind of, you know, a broad idea or process, you know, kind of the technical mathematical process of creating these, you know, rendered point clouds. But it sounds like that has evolved quite a bit. And now, like, there are standard file formats and other things. Talk a little bit about kind of that ecosystem and the tooling that is used around Gaussian splats.

30:38TwiML AI Podcast Host:Yeah, so their Gaussian splats are a pretty particular thing most of the time. It's a collection of points. Each point has a position in 3D space, which is three coordinates, X, Y, Z. It has an opacity, which is a number between zero and one that tells you how opaque it is or how transparent it is. then you've got usually a color which is like an RGB value which is again three numbers and then you'll often have some notion of what are called spherical harmonics and that tells you what does it look what color is it when you look at it from different positions right because you want to model this notion that maybe this point is one color if I look at it from the bottom and a different color if I look at it from the top and that helps you model reflections or right because then if I have a shiny surface like a mirror or a glossy surface then if I look at it from one angle, I kind of see a highlight from a light that bounces from the top.

31:28TwiML AI Podcast Host:If I look at it from a different angle, I see a different color. And these spherical harmonics are a particular concrete way to capture this notion of view-dependent color. So then, you know, a Gaussian splat is this set of points. Each one has an XYZ position, has an opacity, has a color, and also has these spherical harmonics that tell you how the color varies as you look at it from different angles. And then there's different ways to pack that data into bytes in a file, right so then um ply is a pretty popular file format for gaussian splats which is um you know relatively easy to read and write but pretty inefficient and then there's compressed representations like spz um that you know compress some of that data and use lower precision for some parts of it um that let you store something that looks pretty similar um you know looks pretty good but takes a lot there's a lot smaller file size and then you should think that like you know like Like a PLY is kind of like a GIF or a PNG, where it's something that's like pretty uncompressed.

32:21TwiML AI Podcast Host:And an SPZ is a little bit more like a JPEG, where you're going to compress some parts of the data. You're going to throw away some parts of the data. But we think those are parts that you're not going to notice as a human that you throw away some of these parts of the data.

32:33Justin Johnson:Relative to kind of the way we think about like traditional 2D and 3D images, the Gaussian splats are not based on a grid. and are they sparse in general relative to the dimensionality of the space we're representing?

32:52TwiML AI Podcast Host:Yeah, I mean, they're not based on a grid. Any of these points can live anywhere in space, but they typically, I mean, one of the, oh, right, the other important thing I forgot, like the silly me for not having my notes in front of me, but like a Gaussian splat also has a size, right? It's not an infinitesimal point. A Gaussian splat has a size. It's not just a point in space. It has kind of a radius. and, you know, also a covariance matrix because it might not be a sphere, it might be an ellipsoid. And the covariance matrix and the radius kind of, well, really just the covariance matrix kind of tells you like how big is it and how like stretched or squashed is it along different dimensions.

33:27TwiML AI Podcast Host:So then like part of the, you know, part of the original pitch of Gaussian splats is maybe they could be pretty sparse because, or maybe, you know, maybe they could be fairly sparse and you end up with like really big Gaussian splats to cover like a big part of the wall. but in practice that usually ends up in pretty low quality so usually you want the splats to be fairly dense over the geometry that you want to cover that ends up generating nicer images in general

33:51Justin Johnson:That being the case is a part of the technology or what makes Gaussian splats work like that the renderer is able to focus on you know, what's in kind of the user viewpoint versus things that are extraneous to it? Or, you know, maybe the broader question is like, you know, what makes Gaussian splats so interesting? Like a, you know, quick refresher on that relative to the way we've approached this before.

34:24TwiML AI Podcast Host:I think the contrast with Gaussian splats you should be thinking about is triangle meshes. So triangle meshes are kind of your standard representation in computer graphics. And that's saying that we're going to represent the whole world as like little triangles, basically um and everything is like made up of little triangles and and like pretty much any game you've ever played like any vfx shot you've ever seen any computer graph any computer generated imagery you've ever seen um they pretty much model the world as lots of little triangles um and that works that's worked amazing for computer graphics for decades um but the problem is that triangles don't fit with neural networks very well um because if you if you the important part is differentiability you want to be able to differentiate through this uh through this representation past gradients.

35:03TwiML AI Podcast Host:And in particular, that means that you want your representation to have the property that if I change the input a little bit, the output also changes a little bit. And that's not the case with a triangle, because if I've got a triangle here and I move it a little bit, all of a sudden something that became that was invisible now becomes visible. So now I have a sharp change, a sharp discontinuous change in the image that I'm going to see as a function of the parameters. So, you know, Gaussian splats don't have that property because everything is smooth. Everything is partially transparent. So like the image that you see is a continuous continuously changes as you vary any of the parameters of the Gaussian splats infinitesimally.

35:39TwiML AI Podcast Host:So what that means is they integrate with neural networks really well. So neural networks are all gradient based learners, right? I'm going to have an objective function, I'm going to minimize that objective function via gradient descent. In order to do that, I need to be able to pass gradient signal through whatever representation I'm using. And Gaussian splats like pass gradients really, really well so that means they can be plugged into gradient-based optimizers and either like directly optimize against a set of images which happens in the reconstruction case or be plugged into the output of a neural network which is more what we do in the marble case so there the setup is like i've got a neural network it spits out gaussians those gaussians get attached to some loss function then i can back propagate my loss all the way into the parameters of the neural network so that's kind of the biggest delta um that's kind of the big innovation of gaussian splats and why people got excited about them over the past few years is because it's a graphics representation that integrates with neural networks really cleanly.

36:29Justin Johnson:Kind of tying that to some of the core properties of world models, namely consistency, like is it that ability to backpropagate that or how do we get that consistency? Are we, you know, is that coming from modeling? Is it coming from, you know, scale? Like where does that, is it an architecture thing? Where does that come from?

36:59TwiML AI Podcast Host:I think it can come from many places. And that's actually a really interesting jumping off point because there's this property of the world around us that it's consistent, right? If I look at you, you kind of look pretty similar from moment to moment. Or if I, you know, walk to a different room and come back, like you're still going to be here. And that's this notion of consistency. And there's different, and Gaussian splats are kind of consistent by construction, right? Because I've got this explicit 3D representation of the world. So if I look at it and look away, then look back, like it's all there in 3D.

37:27TwiML AI Podcast Host:So it's going to look the same. But you could also get consistency via large scale data and large scale training and large scale compute. And that kind of leans into the more implicit, you know, representations that we've done in RTFM and in other places. So there the idea is, what if I'm going to have a model, there is no Gaussian splats, There is no 3D. There are no explicit points. I just have a model that's spitting out RGB pixel values of the world. But if that model is really, really smart and really powerful and having been trained on a lot of data and maybe with the right expressive architecture, even though it's not mathematically guaranteed, there's nothing mathematically guaranteeing it to be consistent.

38:02TwiML AI Podcast Host:You know, it still ends up being consistent. And I think that's, I think it's actually not that one is better than the other. They're just different points on the technology curve, right? If you have a relatively, you know, low compute budget, and you want things to run embedded, like you don't want to train giant models, then Gaussian splats are appealing because they're consistent by construction. But if you can scale up and use a lot of data and use a lot of compute, then I actually think that the implicit 3D route is going to be the thing that scales up to infinity. So it's more a question of like, it's more of an engineering question of what are the design constraints of the problem facing me right now, and less a philosophical divide for me, right?

38:37TwiML AI Podcast Host:If you want cheap and consistent by construction, Gaussian splats are very appealing. If you want, you know, something that scales to infinity and has, you know, can rely on infinite data, infinite compute, but you're willing to pay that infinite computed inference time and you're willing to pay the big training costs, then I think the implicit pixels only approach is very appealing. And that's exactly why we've done both at World Labs.

38:58Justin Johnson:I think the direction I was trying to go with that question was focusing more on the model and the generation perspective and the idea that, you know, we're describing a world, generating on the fly, the user can navigate through this world, turn away, come back, and we're generating consistent, well, frames in the case of RTFM or splats in the case of Marble. And the question is like, I think part of what you're saying is that consistency is kind of orthogonal to whether we're talking about pixel generation or splat generation. And so the next part of that question is, you know, that consistency seems like a really important part of all this.

39:46Justin Johnson:Like, where does it come from?

39:48TwiML AI Podcast Host:I mean, I think it basically comes from your data ultimately, right? Like, no matter what representation you're using, if it's raw pixels or it's Gaussian splats, like at the end of the day, you have to have a neural network in the system that's been taught to create consistent data. And like Gaussian splats kind of make it easier for a neural network to produce consistent outputs because they're more consistent by construction. But ultimately, it has to come from data of like views of the world that are consistent in the way you want your model to learn.

40:17Justin Johnson:I think then the question is, or the opportunity to like dig into Marble and talk a little bit about kind of the recipe and, you know, how that creates a world model, how that creates the Marble world models.

40:30TwiML AI Podcast Host:I mean, we actually haven't talked explicitly about the Marble architecture, but at a high level, it lets you input as a user different kinds of things. You can input a text prompt. You can input an image. You can input multiple images. You can input a video. Then from that, we generate a 3D Gaussian splat world. And one intermediate step in that is generating a 360 panorama image. And this is where, like, given those inputs, you kind of have one part of a model that generates a 360 panorama view of the world you're about to generate. And that's a big, powerful generative model. that needs to take whatever that user input was and map it first into a 360 pano.

41:13TwiML AI Podcast Host:And then from there, lift it up into a full 3D Gaussian splat world.

41:17Justin Johnson:Is the implication then that the user's input is either a single point or single image

41:28Justin Johnson:or I guess I'm thinking about like if the user's providing a spatially diverse set of images Is it just that the sphere is much bigger?

41:41TwiML AI Podcast Host:Oh, no. So basically what we're doing is like, like the kind of marquee use case is almost single image. And that's probably what works best in marble today. So there it's like, I'm going to input a single image. And then for the stuff in the world that I can see in that image, my generated world should match what I see in my input image. And then the model will try to complete what it, like try to give a plausible completion for everything else in the world that's not visible in the input image, right? So if I'm like maybe taking a picture of a blackboard in the front of a classroom, then the model should know that behind the blackboard is going to be all these chairs where the people sit, where the students sit.

42:17TwiML AI Podcast Host:And then the model should be able to take a picture of a blackboard and then complete that and like generate the chairs behind the blackboard and then lift all that into 3D that you can navigate.

42:24Justin Johnson:And is the user also providing text prompts?

42:27TwiML AI Podcast Host:They can. Yeah, the user can prompt this directly from text. That's kind of optional. If you want to generate your world purely from text, we can do that. If you want to provide text as an auxiliary input to give some extra guidance to what's happening in your input image, that can work too. But we kind of wanted to take this approach with Marvel of maximum flexibility for users that no matter what kind of signal you got, no matter what kind of edit you do, no matter where you want to take it after it's generated, we want to give you a lot of pathways to use this thing and not try to like, you know, guide everyone down one rigid pathway for how to generate things or where to use it after you generate it.

43:02Justin Johnson:I think where my questions around or my assumption of multi-image was coming from, like I've seen some marble worlds that were kind of your classic, you know, room and you spin around the room and it was like a ton of detail. super impressive uh but then i've seen other ones where they're more like kind of these video game immersive worlds where you can like navigate you know through uh you know fantasy kind of village kind of situation and i think are are those all kind of generated from a single image generally and potentially some text conditioning yeah i mean that that's the beauty of the of having a world

43:45TwiML AI Podcast Host:model in the loop there is that you have a single powerful large generative model that's been trained on a ton of data both real world data and fantastical data and photorealistic stuff and non-photorealistic stuff so you've got this big powerful model in the back end that no matter what kind of image or prompt you bring whether that image is of you know a room in your house or you know a fantastical like video game kind of environment um you've got a model that knows all of these different kinds of worlds and can generate them as it's as is necessary for the task at hand. And that's kind of the major difference between, you know, having these generations backed by a big world model versus, you know, classical Gaussian splatting just fitting to these thousand images that I happen to have.

44:25Justin Johnson:And what can you say about the data sets that are used to create these models and like the training approach and recipe?

44:31TwiML AI Podcast Host:Yeah, I mean, you need to train a lot of data. And one thing that's really important is being able to train on a variety of different kinds of data, right? Because ultimately the world itself is 3D and like has a lot of, you know, 3D structure. But there's not a lot of, you know, explicit 3D data for you to learn on. That's a pretty rare form of data. But there's a lot of images out there. There's a lot of videos out there. And images and videos are both, you know, 2D projections of a 3D world. So even if you want a model that at the end of the day is going to produce 3D, it's still very powerful for it to learn on large quantities of image and video data, because those you can get in very large quantities.

45:09TwiML AI Podcast Host:And then when you've got three explicit 3D data, that's very powerful to learn from, but you don't want to be bottlenecked only on learning from 3D data.

45:16Justin Johnson:I'm assuming that you're trying to train on whatever data you have natively, as opposed to like you take an image and project it into 3D or something that sounds super expensive and noisy.

45:29TwiML AI Podcast Host:Yeah. I mean, one of the lessons we learned from deep learning over the past decade is that you want big models trained on a lot of data that are trained end to end. So if you've got a, like, what is your task? Is your task to like generate Gaussian splat worlds today? or is it to generate really powerful 3D consistent frames the next day? Like just think about what is the task that you want to solve and how can I marshal very large quantities of data to let a model learn how to solve that task in a very general way. Is the model's output directly splats or is there some,

45:59Justin Johnson:I think this is kind of what you were just saying, like is it producing some kind of normalized 3D representation, whatever that is, and you can convert that to pixels or splats or you just convert your, it's just spitting out splats?

46:13TwiML AI Podcast Host:Yeah, the kind of rawest output from this model is splats in some way. And then once you've got splats, that's kind of our lowest common denominator 3D format. So once you've got splats, you can render them to an image and that can give you images or videos from these worlds. We can also fit a mesh representation to the world. And then in some contexts, like you want to import this into a game engine or a VFX engine, And sometimes those don't work so well with splats today. And it's useful to have a more classical 3D triangle mesh of the world. So the splat is kind of our most rawest output format from the marble models.

46:49TwiML AI Podcast Host:And then you can go from that to images or videos or meshes. But again, that's in pretty stark contrast to the RTFM model where there's no splats. It just directly spits out pixels.

47:00Justin Johnson:We alluded earlier to a blog post that you recently wrote or your team recently. published. And I thought what was really interesting about that was like, you know, as kind of an analyst, anytime I see a taxonomy that tries to kind of pick apart the distinctions in the space and talk about those, I'm interested. And that is what you try to do, at least in, you know, a particular dimension of world models. I think we've introduced like three or four other taxonomies in this conversation so far. But, you know, talk a little bit about the way you kind of divided up the, you know, that particular dimension of world models.

47:48TwiML AI Podcast Host:So there, like I said, like we talked about a couple of times, I think there's a lot of different flavors of things that people are training and calling world models that look pretty different from the outside. And when we thought about this really hard and realized that there is a framing where they all actually are kind of like different views onto the same thing. And there we realized we could ground this back in this PoMDP formalism that we talked about before. Because in this PoMDP, remember there are three things that are moving around the system. You've got the agent in the world, then you've got the agent producing actions, you've got the world transitioning states, and then you've got the agent receiving observations.

48:24TwiML AI Podcast Host:And we realized if you think about it that way, pretty much everything that people are training today that are called world models are usually outputting one of three things in that loop. Either you're building a model that outputs actions, you're building a model that outputs states, or you're building a model that outputs observations. And once we realized that, it was kind of an aha moment that, oh, it's not that these people are all building totally different things. They're just focusing on different parts of this fundamental PoMDP loop. And they're all connected together, trying to model the world and understand how worlds can respond and evolve and change over time.

49:01TwiML AI Podcast Host:But different people are focusing on different parts of that loop for different applications.

49:05Justin Johnson:Got it. So something like a genie that's producing a stream of like pixels, that would be observations in this model. And something like a robotic system that's producing actions for the robot to take in the world, that's more focused on that as an output.

49:25TwiML AI Podcast Host:Exactly. so then you know then we kind of like taxonomize these into like these three different categories of world models then based on what they're outputting so like you said like if you're outputting the observation like genie or like rtfm then we're calling that a rendering world model or just a renderer right because it's producing a final observation that can be consumed by a by a person or maybe by an agent um so those are kind of what we're terming a rendering a renderer as a world model um then the other the other easy the other the other cool one are all these people training robotics policies, right?

49:58TwiML AI Podcast Host:Like people are training models that, you know, the model operates a robot body and then the robot body does something cool in the world. But then what is that model doing? That model is kind of on the other side of the PO-MDP loop, right? It's receiving observations from the real world. And now the model needs to take actions to, you know, try to make a change in the world. And then they've got a world model that is inputting real world observations and outputting actions to be made in the real world. Whereas the, and that we're calling a planner, world model as planner because it's planning out a sequence of actions to take in the world and that's almost like exactly dual to the world model as renderer because the renderer is sort of receiving actions from a probably from a human user and then outputting observations of what a world might look like under those sequence of actions so the renderer and the planner are kind of like almost perfect duels to each other and then the third one is what about the state like that's a tricky one that we keep coming back to um so that we're calling world model as simulator right because an important another important property of world models is that maybe they should do some kind of simulation of the state um in some contexts you only care about the action of the observation but in other contexts you might want to know something about the state of the world under consideration and under and maybe you know understand or simulate how that state might evolve in response to actions in some explicit or semi-explicit way so that's the the world model is simulator and then we kind of realized that you know planner simulator renderer when you kind of think and break it down to these terms, pretty much all the world models that people are training these days actually can be bucketed into one of these three camps usually.

51:29Justin Johnson:I thought the simulator was interesting in that I typically think of simulation as like this external tool that we're using to, you know, develop models as opposed to this framing of the model itself, you know, being fundamentally a simulator.

51:48TwiML AI Podcast Host:And there, I think like, Another interesting thing is these start to blend together. And I think Marble is actually an example of something that is somewhat straddling the boundary between world model as renderer and world model as simulator. Right. Because Marble ends like when you as a user interact with Marble, right, you're seeing pixels on the screen. Right. And those pixels, you know, so in that sense, you're seeing an observation. But that observation did not come directly out of the neural network. That observation came out of a set of Gaussian splats. and the Gaussian splats are kind of this state representation, this explicit state representation that we have of the world.

52:23TwiML AI Podcast Host:And then once you have that explicit state representation, you can do other things with it other than rendering, right? Because it's an explicit 3D representation, you can measure the distance between two points or I can manipulate the state explicitly by pulling in another object asset and putting it into that world. So at least the version of Marble we have today, like the kind of the end-to-end experience as a user is you're seeing observations and you're seeing pixels on the screen. so it's kind of a renderer, but the model itself is outputting this explicit or semi-explicit world state that you can then do for other things.

52:55Justin Johnson:That feels like a much more nuanced distinction than renderer and planner. Like if I think about a renderer that as opposed to spitting out 2D frames with spitting out somehow, you know, 3D directly, like then you can do, you know, measure distances between points and that, you know, it's a representation, but it's also the observation, like.

53:21TwiML AI Podcast Host:Exactly. And that's the other kind of point that we wanted to make here is that while there is this taxonomy, it's not very rigid. And I think sticking too rigidly to it would be a, would be a, would do a disservice to all of us. Right. So any, like the real world is messy and all of our taxonomies break down and, but it's a, it's a useful framework to think about. But in reality, I think the world models that are going to carry us at the end of the day are actually going to blend all of these aspects together right we we ultimately want to have one combined system

53:46Justin Johnson:that could do all these things yeah but i think i'm also asking for more elaboration on like simulator and and you know beyond the beyond the marble example are there other examples that kind of capture this idea of model a simulator yeah i think there's a couple examples

54:09TwiML AI Podcast Host:I actually don't think anyone's nailed that one right now, which is why it's, but I think there's a couple of flavors of future systems I can imagine here. One is maybe like the marble explicit state future version. Not to say that this is actually what we're going to do, but one could do. Also not to say that we're not doing it either, right? But you could imagine a version of this that like treats something like Gaussian splats as a world state, but actually has a model evolve that state over time. So you could have then a model that like inputs an image, outputs a Gaussian splat world, and then a user is going to take some action, like pick up the bottle or move something around.

54:47TwiML AI Podcast Host:And then you'll have a model that will then go in and update the Gaussian splat world with a powerful neural network model. So that would be then a model that is, you know, working with this explicit or semi-explicit world state and then actually being able to evolve that world state over time in response to actions. And I don't think anyone's really built a system like that, but they could. And the other would be kind of the implicit world state. and maybe we can build systems that maybe they're not working on Gaussian splats directly, but they have some kind of vector, implicit vector representation of a world state.

55:16TwiML AI Podcast Host:And it behaves in the way that an explicit state would, where you can render observations from it. You can have networks that predict how that state evolves in response to actions, predict how that state is going to evolve in time, and sort of have a neural network analog of that explicit world state. And I think people are working on that, but that feels like it's still a pretty open research question as to what's the exact right recipe to get that to work.

55:40Justin Johnson:Yeah, the other thing that jumps out at me in thinking about the simulator is that like maybe a perfect expression of a simulator is this theory builder that we talked about. Like it takes these abstract ideas and kind of boils them down into, you know, in this case, the state is a set of theories about the world.

56:03TwiML AI Podcast Host:Yeah, I mean, then you got to talk about like states and metastates and theories and metastates their theories and push it up a level of abstraction, right? Like you could say like maybe the theory builder is the one who's writing the laws of physics in the states. So like the theory kind of encapsulates the kinds of worlds that may exist and how those kinds of worlds are allowed to evolve. And then the state is like one particular instantiation of the theory that tells you a particular world. But that's like, that's that I don't think anyone has any clue how to do.

56:27Justin Johnson:Yeah. Yeah. Okay. We're getting a bit too abstract there. You talk a little bit about this idea of a unified world model, I think we captured that a little bit is just the idea that it's kind of a leaky abstraction and you're expecting kind of crossover in real world products.

56:51TwiML AI Podcast Host:Exactly. What I think we're going to get to as a field is models that can do all of these jointly and they're all going to benefit from each other. And why is that? It's because they're all kind of asking similar questions. Like they want to understand what are the kinds of worlds that could exist? How could those worlds respond to action? What do those worlds look like? What kind of observations arise from those worlds? How do they evolve in time? What can you do with them? These are all kind of fundamental questions that all connect to each other, right? If I get better at understanding how the world state evolves, probably I also get better at anticipating what it's going to look like from different angles as a renderer.

57:25TwiML AI Podcast Host:And if I'm really good at rendering, like imagining how the implicit world state is going to evolve, probably that's pretty good for planning, right? If I know, if I can implicitly, you know, evolve my world state, probably I can also plan against that state and know how to take actions to affect the world. So I think it's pretty, I don't think we're there yet, but I think over the next couple of years, we'll start to see models that combine more and more of these capabilities into one powerful unified model. And then like what output you want at one moment in time is less a function of having specialized models that are renders or simulators or planners, but more about, you know, is it today I want to drive a robot?

57:58TwiML AI Podcast Host:So it's going to operate in planner mode. And tomorrow I want to, you know, drive a virtual video game. So it's operating in renderer mode. Or tomorrow or the next day I want to, you know, simulate possible counterfactuals in a world. And now I want it to operate in simulator mode. So I think like that's kind of where the field is going to get the next couple of years.

58:14Justin Johnson:That makes me think of like this idea of like the output layer in a neural network is like it could be classification or it could be, you know, regression or something else. But like the core representation is, you know, for the most part the same. and you know we're just kind of using it in different ways.

58:33TwiML AI Podcast Host:Yeah exactly I think that's what we're going to get we're going to have like these giant unified world models that have maybe different input heads different output heads that know how to input and output different kinds of things but ultimately like all the compute all the parameters of this model are going to be this shared trunk that is the that is this world model that knows how to simulate any kind of a thing implicitly in its weights then it can surface that world knowledge as actions or as states or as as observations as needed for the application.

58:57Justin Johnson:I think the next question I wanted to get at was, and really kind of a closing question, is like, do current architectures get us there? Like, and you just said, you know, kind of no, like the architecture evolves, but it's kind of consistent with the way we think about, you know, today's models, transformers, et cetera. Do, you know, do you foresee kind of an architectural step function being required to, you know, fulfill the potential of world models?

59:29TwiML AI Podcast Host:Maybe yes, but not as big a one. Probably not too big. I think transformers are really, really powerful. Yeah, I get this question a lot. Like, does that mean transformers are dead? Do we need new architectures? Like, no, transformers are great. Transformers are super powerful. They scale up really well. They can work on all the different kinds of data. All different kinds of data. So transformers are very powerful. They could use them for all the different kinds of problems. I think there is a loss function question about what is the right loss function for training these kinds of generative models.

59:57TwiML AI Podcast Host:And that's more up in the air. Like, do we want to train these things, you know, as a generative model? If it's a generative model, do you train it via diffusion or rectified flow? Do you train it as a discrete auto-aggression or something else? So there is a kind of a loss function question that's really applicable to any kind of generative model that I think is probably going to evolve and has already. Like we used to do GANs, now it's diffusion. Like that's a kind of a change in loss function more than it is a change in architecture. but you know another one area where i do think we need to see some evolution architecturally is how do we deal with really really long contacts and really really really lots of tokens so you know this is something that comes up in llms already like llms are transformers they work on tokens but you can do a lot of problems you know even maybe maybe this is less true now in the agentic era but a couple years ago like it was hard to imagine situations that an llm needs to operate on million or 10 million tokens.

1:00:48TwiML AI Podcast Host:Because that's like whole books. Like, yeah, in principle, there's problems you need to fit whole book, cram whole books into your context. But, you know, maybe those are not, you can do a lot with language without needing that big of a context. But nowadays, but if you want to talk about worlds, like, you know, I need to generate, you know, model maybe lots and lots of high dimensional images or lots and lots of space in 3D. Like it's very quick, very easy to get situations where you want hundreds of thousands or millions or tens of millions of tokens of context for different world modeling problems.

1:01:16TwiML AI Podcast Host:So I do think we need to see some evolution in, you know, how do we adapt transformers to work at really, really long context lengths to, because that becomes, that becomes kind of a nice to have in LLMs. That becomes the everyday problem for any scaled up world model.

1:01:32Justin Johnson:That begs the question, how should we think about both tokens and contexts in world models? Like, does the context length limit the size of the world? or, you know, are we like paging out sections of the world? And so it's not a hard constraint. And like, does token correspond to a splat or some other construct? Like, how do these things relate?

1:01:54TwiML AI Podcast Host:I think that's where there's a lot of, you know, different things happening. And that's where I do see a lot of evolution. So that's like, in some architectures, a token will be maybe a little chunk of a video, you know, maybe a 16 by 16 spatial, 16 by 16 pixels and like four frames in time. So some architectures, a token is literally like a little patch of a video. in some architectures that token might be a little bundle of splats somewhere in the world in some architectures that token might be might be a little chunk of 3d space maybe i've carved up my 3d space into voxels and now each token corresponds to some chunk of 3d space or maybe those tokens are something abstract and latent like maybe i've got a latent world state that is just a number like maybe i've just allocated you know my world state has a thousand tokens what do they mean it's opaque model figure it out so i think that's where we're seeing a lot of evolution and a lot of different approaches taking doing different things

1:02:46Justin Johnson:architecturally also in this architectural thread do you see a role for kind of these ideas around geometric deep learning or you know different models that incorporate symmetry you know some um it has been a variety of work around trying to incorporate geometry into uh deep learning and And at least from a particular lens, it seems like if we're talking about spatial information, there may be a role for that.

1:03:15TwiML AI Podcast Host:Maybe. But I think the lesson we've learned over and over again in deep learning is that you want simple representations and then scale up the model behind the simple representations. So, you know, when you and then also like choose a representation that's adapted for the task at hand. Right. Like if you want to if the thing you actually want out is video frames or images, like just do that. You don't need to bottleneck yourself through an explicit 3D representation. If you do want a 3D representation, be it a splat or a mesh, find a simple way to interface that architecture, that 3D representation with the neural network.

1:03:46TwiML AI Podcast Host:And the more you put in symmetries and complicated representations, often the harder it'll be to optimize and the more assumptions you're making. So the less it's going to work at scale, right? Like you put in some notion of, you know, symmetry. Well, people are usually kind of symmetric. We have two legs and two arms, but not everybody has two legs and two arms. So those hard assumptions of symmetry maybe get you maybe 80 % or 90 % of the way there, but they're going to break down eventually. And then you'd like to be in a position where your architecture or your model is expressive enough to handle the cases where those shortcuts are going to break down.

1:04:22Justin Johnson:Where should folks go to learn more about this? Are there any canonical resources that you like to point folks to who are new to the space and want to dig in more deeply?

1:04:33TwiML AI Podcast Host:Yeah, so you can definitely try out Marble. That's our product at marble.worldlabs.ai. You can sign up and play around with that today. I wish I could recommend a better lecture series or book or something on world models, but I just, I don't think anyone's written down

1:04:47Justin Johnson:anything super awesome yet,

1:04:48TwiML AI Podcast Host:which is partially what we tried to solve with our blog posts. But I think someone could go a lot deeper and unpack a lot of these ideas in a better way. So I, you know, if someone, if I'm missing something, I'd love to, for your viewers or listeners to let me know.

1:05:02Justin Johnson:Justin, thank you so much for taking the time to share with us a bit about what you've been up to you and the team at world labs very cool stuff

1:05:10TwiML AI Podcast Host:yeah thanks so much for having me this was a lot of fun

From the publisher

In this episode, Justin Johnson, co-founder of World Labs, joins us to discuss world models and the emerging field of spatial AI. We explore why many researchers see capabilities beyond language as an important frontier for AI, and what it means to build models that can understand, generate, and simulate the environments around them.

Justin explains the different approaches to world modeling, including explicit 3D representations and generative models, and why there is still no established recipe for building these systems. We also discuss World Labs’ Marble system, which can generate navigable 3D worlds from images and other inputs, the challenges of evaluating world models, and the role of simulation, planning, and action. Finally, Justin shares his vision for models that bring these capabilities together, supporting everything from interactive virtual environments to agents and robots that can operate in the physical world.

🗒️ Full show notes: ⁠https://twimlai.com/go/775.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
World Models and the Future of Spatial AI with Justin Johnson - #775The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 6 min
Listen in VO