#303 Fei-Fei Li: Spatial Intelligence, World Models & the Future of AI

23 Nov 2025 · 1 h 1 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Eye On A.I. #303 - Fei-Fei Li: Spatial Intelligence, World Models & the Future of AI

Episode Overview In this episode of *Eye on A.I.*, host Craig S. Smith interviews Fei-Fei Li, a distinguished pioneer in artificial intelligence and computer vision. The discussion centers around the concept of spatial intelligence, the development of world models, and their implications for the future of AI. The conversation explores how AI can evolve from understanding text to reasoning about three-dimensional environments.

Key Themes

  • Spatial Intelligence: The ability of machines to perceive, reason, and interact within 3D spaces, which goes beyond traditional language processing.
  • World Models: Models that allow AI to generate and understand complex 3D environments, exemplified by Fei-Fei's startup, World Labs, and their product, Marble.
  • Multimodal Learning: The significance of integrating various input modalities (videos, images, text) for richer understanding and interaction with the world.
  • Continuous Learning: The challenges and potential of creating AI that learns continuously from interactions, much like humans do.

Detailed Highlights

The Evolution of AI

  • Fei-Fei Li expresses that advancements in AI are at a point where technology needs to integrate spatial comprehension, moving beyond simple image recognition and language processing.
  • Emphasis is placed on the need for AI systems to learn from direct experiences in the world.

Marble and Spatial Intelligence

  • Marble is introduced as a pioneering model capable of generating consistent 3D spaces from internal representations.
  • Fei-Fei discusses the importance of explicit and implicit representations in world modeling, suggesting both are necessary for comprehensive understanding.
  • The model allows users to interact and edit 3D environments, providing a hands-on approach to spatial intelligence.

Multimodal Inputs

  • Fei-Fei emphasizes the need for AI to learn from various inputs, such as text, images, videos, and even 3D layouts, reflecting how humans learn through diverse experiences.
  • Integration of multimodal learning is seen as crucial for advancing AI capabilities.

Continuous Learning and Memory

  • Current models, including Marble, primarily rely on offline learning, but there is an openness to exploring continuous learning mechanisms that allow AI to adapt and evolve based on new experiences.
  • The conversation touches on the importance of creating architectures that facilitate ongoing learning and adaptation.

Future of AI and Creative Reasoning

  • Fei-Fei discusses the potential for AI to reach new heights in creativity and scientific reasoning, with the ability to deduce complex concepts like physical laws.
  • The limitations of current AI architectures in abstract reasoning are acknowledged, suggesting that breakthroughs in model architecture are necessary for further advancement.

The Role of Collaboration and Multiverse Experiences

  • The potential for AI to facilitate global collaboration through teleoperated robots and immersive digital experiences is highlighted.
  • Fei-Fei envisions a future where the digital world enhances educational and creative processes, allowing for interactive learning experiences.

Key Takeaways

  • Spatial Intelligence is Essential: Understanding 3D environments is crucial for building AI systems that can effectively interact with the physical world.
  • Multimodal Learning is a Priority: Incorporating various forms of input will enhance AI's learning and adaptability.
  • Continuous Learning is the Future: Developing AI that learns from ongoing experiences will bridge the gap between human-like intelligence and machine learning.
  • Architectural Innovation is Necessary: New AI architectures will be vital for unlocking advanced reasoning and understanding capabilities.

Conclusion The episode concludes with Fei-Fei Li emphasizing the importance of spatial intelligence in the quest for advanced AI. She articulates a vision where AI systems can not only navigate the world but also creatively engage with it, suggesting a future rich with possibilities for education, creativity, and innovation.

Additional Resources

  • Follow Craig Smith on X: [@craigss](https://x.com/craigss)
  • Follow Eye on A.I. on X: [@EyeOn_AI](https://x.com/EyeOn_AI)
  • Explore AGNTCY: Visit [AGNTCY.org](https://agntcy.org/) for more on multi-agent software development.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The spatial intelligence work I've been thinking about in the past few years is truly a continuation of my entire career's focus in computer vision and visual intelligence. Why did I emphasize on spatial is because we've come to a point in our technology that the level of sophistication and profound capabilities of this technology is no longer at the level of staring at an image or even simple understanding, simple videos. It is deeply, deeply perceptual, spatial, and also connects to robotics, connects to embodied AI, as well as ambient AI. Welcome to the podcast. In this episode, I have the honor of talking again to Fei-Fei Li, a pioneer in artificial intelligence and computer vision.

0:51I had Fei-Fei on the podcast a few years ago, and I invite you all to go listen to that episode. I'll put a link at the end of this video. We'll explore her insights on world models and the importance of spatial intelligence, crucial elements for creating AI that truly understands and interacts with the world around us. While large language models are amazing, much, if not most, of human knowledge is not captured in text. And to reach a more general artificial intelligence, models need to experience the world firsthand, or at least through video. We talk about her startup World Labs and their first product, Marble, which generates incredible complex 3D spaces from the model's internal representations of the world.

1:44Plus, she tells us her guilty pleasure watching her favorite TV show on airplanes. and you probably won't be surprised what that show is. Stay tuned for an enlightening conversation. I'm Fei-Fei Li. I'll be joining the podcast of Eye on AI and discussing spatial intelligence and world models. Come and join us. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting.

2:51Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. Collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next generation AI infrastructure together. Agency is dropping code, specs, and services, no strings attached. Visit agency.org to contribute. That's A-G-N-T-C-Y dot O-R-G. I wanted to talk not so much about Marble, your new model, which is amazing, that generates a sort of consistent and persistent 3D worlds that the viewer can move through.

3:56I want to talk more about why you're focusing on world models on spatial intelligence, why that is necessary to go beyond learning in language, and how your approach differs from the Kuhn's. So can you talk first of all about, did the world model work emerge from your work on ambient intelligence, or has this been a parallel track? The spatial intelligence work I've been thinking about in the past few years is truly a continuation of my entire career's focus in computer vision and visual intelligence, right? And why did I emphasize on spatial is because we've come to a point in our technology that the level of sophistication and profound capabilities of this technology is no longer at the level of staring at an image or even simple understanding, simple videos.

5:05It is deeply, deeply perceptual, spatial, and also connects to robotics, connects to embodied AI, as well as ambient AI. So from this point of view, it's a continuation, really, of just my entire career in computer vision and AI. And the importance of spatial intelligence, and I've talked about this, I've had this podcast for a while, human knowledge, language models learn from human knowledge encoded in text. but that's a very finite subset of human knowledge and as you pointed out many people pointed out humans learn a lot from interacting in the world without language so it's important we're going to move beyond the current LLMs as amazing as they are that we develop models that have more direct experience of the world, learn more directly from the world.

6:13Your approach has been, well, certainly with Marble, to take the internal representations of the world that the model learns and create an external visual reality with those. Lacoon's approach is to create internal representations from direct experience or video input that allows a model to learn physical laws of motion and things like that. Is there a parallel between that? Are the two approaches complementary or are they overlapping? So yeah, so first of all, I actually would not pit myself against Young or vice versa, because I think we are on this continuum of intellectual

7:17approaches to spatial intelligence and world modeling. So you probably read my recent long essay that I called Manifesto of Spatial Intelligence, and there I was very clear. I actually think both implicit representation and eventually some level of explicit representation, especially in the output layer, are both probably needed if we are considering a universal, omnipotent world model eventually. And they both play roles. For example, our current world model in WarLabs, Marble, does explicitly output 3D representation. But within the model, there is actually also implicit representation, you know, along with the explicit output.

8:13So it is, eventually, I think we need both, to be honest. And also, in terms of the input modality, yes, learning from videos is very important. You know, the whole world is a continuous input of lots of frames. But the whole world to intelligent or just to animals is not just watching passively. It's also, you know, an embodied experience of movements, of interaction, of tactile experience, of sound, of, you know, smell and also just forces, you know, physical forces and temperature and all this, right? So I think it's deeply multimodal. And if you even look at, of course, Marble as a model is just the first step.

9:12But in our tech piece that we released just a few days ago, we were pretty clear that we believe in multimodality as a learning paradigm, as well as an input paradigm. So, yeah, so I think, you know, there's a lot of intellectual discussions on this. It also shows the excitement in the early days of this model. So I would not say that we have finished exploring exactly the model architecture, the representation and all that. In your world model, the inputs are primarily videos, and then the model builds an internal representation of the world? Not quite. If you experience Marble, our world model, our inputs are generally pretty multimodal.

10:14You can do a pure text. You can do one to multiple images. You can do videos. You can also input coarse 3D layouts, you know, like boxes or voxels. So it's multimodal, and I think we're going to deepen that as we go. And is the ambition beyond, you know, it's a fantastic product with many applications, is the ambition to build a system, and when I said the input is video, but to build a system that can learn from direct experience, whether it's through video or some other modality, but it's not learning through a secondary medium like text. Yes, I think world model is about learning about the world, and the world is very multi-model.

11:13And whether it's machines or animals, we are multi-sensory as well. So the learning is through sensing, and the sensing has different modalities. So text is one form. Again, this is where we depart from, you know, animals, because most animals don't learn through sophisticated language, but humans do. But, you know, our today's AI's world model will learn from a lot of language input as well as other modality. But it's not squeezed through only language. Yeah, and one of the limitations of LLMs is that the models are, the model parameters are fixed after training. So the models don't learn continuously.

12:07There's a certain amount of learning at test time inference. Is that something that you're tackling with world models? because presumably a world model is learning as it encounters new environments. Yeah, the continuous learning paradigm is definitely a very, very important one, especially for, you know, living beings. That's how we do it. And even there, there is, you know, online learning versus offline learning. So in the current form of our world model, we are still more in the batch or offline learning mode. But we're definitely open to, you know, continuous, especially eventually online learning modality.

13:02Yeah. And how would that, I mean, it wouldn't be a very different architecture. It would simply be a matter of engineering. Is that right? Well, I would be open-minded. You know, yeah. Yes, I think it's going to be a mixture of both. You know, obviously, good engineering, good, you know, fine-tuning online, you know, learning can already happen, but maybe there will be new architecture. Yeah, yeah. Can you talk about the real-time frame model, which underlies Marble and your work with world models? Yeah, so you're referring to a tech blog that we also put out a couple of weeks ago that specifically double clicks on our real-time frame model.

13:56So WorldLab is a very research-heavy organization. We also care about product, but so much of our work right now in this stage is model first. So we're definitely looking at how we can push forward spatial intelligence. And this particular line of work, which definitely connects to marble, is really focusing on how we can achieve frame-based generation with as much geometric consistency and permanency as we can. Because some of the early work in frame-based generation, you lose that permanency as you move forward. But in this particular case, we try to really balance and also do it in a compute efficient way during inference, which in this case we achieve through a single H100 in inference time.

15:00We don't really know. Some other frame-based models, they don't tell us how many chips they use during their inference time. So we have a hypothesis. It's quite a number. But we don't know that information to compare ourselves. Yeah, and in your manifesto, I think you called it, you talk about the need for a universal task function. A universal what function? Universal task function. Oh, yes, yeah. Analogous to the next prediction, next token prediction in language models. Does RTFM does have a prediction element to it? What are you talking about when you say universal task function beyond what the predictive element of RTFM?

16:03Well, so one of the biggest breakthrough of Gen.AI is really this discovery of the objective function of next token prediction, right? Because it really is such a beautiful formulation because language is in this kind of sequential, you can tokenize language in this kind of sequential representation. and your learning function of next token prediction is precisely what you need during inference, which is as you generate language, whether humans generate or computers generate, it really is putting token after token forward. So having an objective function that is 100 % aligned with the actual eventual task that is supposed to achieve or carry out is just great because it makes the optimization just on target.

17:12In computer vision or in world modeling, it is not as simple because if you look at our relationship with language, it really is just to speak it or to generate it. There's no language in nature that you're staring at it. I mean, eventually you learn to read, but that's because it's already generated. So, but there is, you know, your relationship with the language is deeply generative. It literally is a generated thing by humans. But the relationship with the world is so much more multimodal, I would say, right? There is a world out there for you to observe, for you to interpret, for you to reason with it, for you to eventually interact with it.

18:08But there is also a mind's eye that can formulate different versions of reality as well as imagination and also allow you to also generate the story, the imaginal world. or so it's much more complex. So what is the task that defines or the objective function that defines a universal function that we can use that is as powerful as next token prediction? It's actually a really profound question because is it, let's say, is it 3D reconstruction? You know, just some people actually would argue that the universal task for world modeling could be that it's just be able to 3D reconstruct the world. Because if that is the objective function and if we achieve that, a lot of things will fall through, just fall naturally out of it.

19:16But that is, you could also argue, I don't think so, because most animal brain don't necessarily do precise 3D reconstruction. You know, like, it's not clear a tiger or a person actually can reconstruct the world, yet we're such powerful, visually intelligent, spatially intelligent beings. So maybe that's not necessarily the right task. Then is it the next frame prediction as if it's the next token prediction? Well, there's some power to that, right? First of all, there's a lot of training data for this. Second is that in order to predict the next frame, you have to learn the structure of the world because worlds are not white noise.

20:11So there's a lot of structure that connects one frame to the other. And if you do it well, maybe that is the right universal task or objective function. But maybe it is also unsatisfying because you treat the world as 2D. And the world is not 2D. It's so you do you really force the representation, collapse it in a very unsatisfying way. And also, even if you managed, you can say if you managed to do it perfectly, 3D is implicit. That's true. But it's also very wasteful because with the 3D structure, there's actually a lot more information that you don't have to lose in the way that frame-based prediction might.

21:07So there is still a lot of exploration on that. RTFM. And I have to ask you, was that a joke, naming it RTFM? It was a brilliant play with the, I did not invent this. One of our researchers is really brilliant in naming, right? Like, as you know, you know, I don't know how much we're allowed to say those words. I'll say it. Yeah, read the fucking manual. That's the acronym, yeah. Every computer scientist knows that. So we find that really fun to just play with that name. Yeah, but RTFM is predicting the next frame, right? That's, that's, and. With 3D consistency. Yes, and that's what's interesting about the internal representation that, that the model learns.

22:06I mean, sitting here, I'm looking at my computer screen. I know what the other side of the computer screen looks like, even though I can't see it. And there's an internal representation in my mind of that. And your model does that. I mean, that's why you can move around objects, even though it's a 3D representation on a 2D screen, but you can move around, see the other side of things. So the model has an internal representation of the 3D object even if its current view can't see the backside of something. And that really interests me. That's essentially what you're talking about when you say spatial intelligence, understanding that 3D world.

23:07So does that learning include things like the physical laws of nature? I mean, does it understand that you can't walk through a solid object? or that I think on one of the podcasts I saw you on, somebody was talking about using this for people with fear of heights. So you could look over the edge of something. But if you create a physical representation, I mean an explicit representation of a cliff, let's say, and you move the point of view of the agent or the viewer over that cliff, will it know that it's no longer standing on solid ground or will it float in space? Yeah, so what you are describing is both physical and semantic, to be honest.

24:18Like, you know, of course, falling off a cliff is very much dependent on the law of gravity and all that. But the fact that you, you know, going through a wall, it's very much material-based and semantics-based, right? Solid object versus non-solid object. So RTFM as a current model right now is not focusing on the physics yet. But most of the physics, to be honest, coming out of the Gen AI models are statistics. If you look at, for example, you know, a lot of these Gen video models where you do see water running and trees moving, that is not based on a Newtonian law of forces and masses. It's really based on I've seen plenty of movements of water and leaves in this particular way, and I'm just going to follow that statistical pattern.

25:24So we have to be a little careful. Right now, World Labs is still focusing on generating and exploring static worlds. We are going to explore dynamic. and there a lot of that will be you know statistical learning um i don't think any of today's ai language ai or pixel ai has the capability to abstract at a different level and deduce physics like at the level of a newtonian law out of out of it everything we've seen is statistics-based, statistical-based physics and dynamic, physical and dynamical learning. We could, on the other hand, put these worlds into physics-based, physics engines. And these engines have the laws of physics.

26:28And they eventually, these physics engines, game engines, and also world generation, eventually they're going to combine into neural engines. I don't even know. Maybe we should call them neural spatial engines or something like that. I think that we are moving towards that direction, but it's still early days. Yeah. Yeah. And I didn't mean to pit you against Jan. What I was driving at is it seems like you're focused on explicit representations coming out of an abstract internal representation. Jan is just focused on the internal representation and learning, and that's where the learning takes place.

27:18And it just seemed to me that they would marry beautifully. That's a possibility. Like I said, like you already repeated, we are, you know, We are exploring both. The explicit output is actually a very deliberate approach because we want to be useful for people. We want to be useful for people who are creating, who are simulating, who are designing. And if you look at today's industry, whether you're creating VFX effects or you're creating games or you're designing interiors or you're simulating for robots or autonomous vehicles or industry digital twins or whatever you're doing, it's very 3D, right?

28:06So there's entire industry after industry that's very much 3D in the workflow of 3D. And we want to be absolutely useful for people and businesses to use these models. And the reason I brought up continuous learning, right now, the model understands, you know, depth of field and other sort of spatial properties. What is the learning that is implicit in the model or that the model has that allows it to generate that explicit representation? I mean, because the ultimate goal is to build a model that presumably can learn over time. I mean, I'm sure everybody, you and everybody else, I mean, I imagine it having a model that maybe it's on a robot or maybe it's attached to a, it's getting data from a video camera that moves around in the world.

29:30but that eventually learns not only the scene that it's seeing, but understands the physicality of the space. And eventually then you marry that with language and you've got a really powerful intelligence. Is that something that you think about? And that would require continual learning, that it's not a finite data set that you're feeding the model, that the model is continuously learning from its interaction in the world. Yeah, definitely, especially as one comes close to a use case, especially if the use case requires continuous learning. You know, there are many ways of continuous learning, right?

30:20Like in language model, we see that taking to context itself is continuous learning, right? like having the context as memory. So it helps inference. And then, of course, there is other methods of online learning, of fine tuning. So continuous learning is a term that can encompass multiple ways of doing it. And I think in spatial intelligence, especially like you said, some of these use cases, whether it's robots in personal or in customized situations, or or artists with particular styles, creators with particular styles, these will all eventually push the technology to be able to be more responsive in whichever time horizon that the use case requires it to be.

31:15Some are real-time, some are maybe just more segmented in the time horizon point of view. So it depends. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux Foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows.

32:24collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's That's A-G-N-T-C-Y dot O-R-G. Yeah. And, you know, you've come a tremendous way since you're running a dry cleaning shop, as I recall, in New Jersey. And in a very short period of time. And you're moving very quickly now. do you have any sense of where you'll be in five years with this technology? Will you have, you know, some physics engine built into the model, or will you have, you know, the ability to learn in longer time frames and build up richer internal representations that the model can begin to understand the physical world.

33:53Yes. I actually think, you know, as a scientist, it's very hard for me to give a prediction in terms of precise time because some part of the technology has moved much faster than I thought, some are much slower than I thought. But I think that is very much a very good goal. And also, five years is kind of a fair guesstimate. I don't know if we'll be faster, but it's a better guesstimate than 50 years, in my mind, or five months. Can you talk a little bit about why you think spatial intelligence is the next frontier? I mean, we've already said that human knowledge is contained in text is only a subset of all human knowledge.

34:51And while it's very rich, you can't expect an AI model to understand the world simply through text. Can you talk about why that's important and how Marble and World Labs relates to that larger goal? Yeah, so I think fundamentally technology should help people. In the meantime, understanding the science of intelligence itself is the most fascinating, audacious, ambitious scientific quest I can think of myself, and it's a quest of 21st century. Both of these, whether you're fascinated by the curiosity of the science or you're motivated by using technology to help people, it points to one fact is that so much of our intelligence and so much of our intelligence at work goes beyond language, right?

35:56It's, I kind of lightheartedly say that you cannot use language to put down the fire. In my manifesto, I described examples, whether it's the spatial reasoning deduction of DNA double helix structure or a first responder putting down fire in a rapidly evolving situation with a team of coworkers. a lot of this goes beyond language so it is obvious both in terms of use case point of view as well as in terms of a scientific quest that we should try our best to unlock how to develop spatial intelligence technology that you know takes us to the different to the next level so So that's really the 30 ,000 feet view that describes how I'm motivated by this dual purpose of scientific discovery, as well as making useful tools for people.

37:09And then we can dive deeper into, for example, being useful, like we alluded to a little bit earlier, whether you're talking about creativity or simulation or design or immersive experiences or education or healthcare or, you know, manufacturing. There's just so much that you could do with spatial intelligence. I actually, I'm very excited that so many people who are thinking about education and immersive learning and immersive experiences are telling me that Marble, our first release of the model, is inspiring them to think how to use it for immersive experiences, you know, to make learning more interactive and fun.

38:01And that's just so natural because, you know, pre-verbal children learn entirely through immersive experience. And even today as grown-ups, our life, right, like so much of our life is immersed in this world and involves, Of course it involves speaking and writing and reading, but it involves doing and interacting and enjoying and all that. So it's so natural. Yeah. I mean, one of the things that has struck everybody is that this marble, I'm not sure it's marble or the itself or the RTFM, the generation of next frame in a moving through a space, is that it's running on one H100 GPU. And I've heard you in other talks refer to experiencing the multiverse, which, you know, everyone was very excited about until they realized how much computation it was going to take and how expensive it is.

39:29do you really think that this is a step toward creating worlds for education, for example, because you've been able to reduce the computational load? Not just that. First of all, I really believe at the inference front, we will be speeding up, we will be more efficient, and we'll be also better, bigger, higher quality, um longer you know um experiences all this i that is the trend of technology i i really do things are moving fast i also do believe in multiverse a multiverse experience you know humanity um as far as we know the entire history of humanity our world experience is in one world And it's literally physically this earth, right?

40:26Okay, a handful of human beings have gone on to the moon, but that's about it. And that is the one shared 3D space. And we build civilization. We live our lives. We do everything in it. But with the digital revolution, digital explosion, We are moving part of our lives in digital world. And there is also a lot of crossover. This is not, I don't want to paint a dystopian picture of we have abandoned our physical world. nor am I going to paint a utopian, total hyperbolic utopian world of everybody wear a headset and never even enjoy and look at our beautiful real world. And that's the fullest of life.

Read the full transcript

41:19I would reject both notions, but both pragmatically speaking, as well as just projecting to the exciting future. The digital world is boundless. It's unlimited. And it gives us so much more dimension and experiences that our physical world will not allow us. For example, we already talked about learning. I would love to have learned chemistry in a much more interactive and immersive way. because I remember my college chemistry class has so much to do with arranging molecules, understanding the parities and the asymmetry of the molecular structure. Man, I wish I could just experience that in an immersive way.

42:12I would love for creators to, you know, creators, I meet these creators, I realized in their mind's eye, they simultaneously, every given moment, they have so many ways to tell their story. There's so much in their head, but they're rate limited by tools. You know, if you use a real engine, it's going to take weeks, hours after hours to express one world in your head. And whether you're making the fantastical musical production or you are, you know, designing the bedroom for your newborn child. so many of these moments if we allow people to use the digital universe as much as the physical world to experiment to to iterate to communicate to to create it's just so much more fun and also we also are digital age is also helping us to break the physical boundaries of of labor right Like I can tell, I mean, we're looking at robots that can be teleoperated.

43:31I can also totally imagine creators collaborate across the globe through embodied robotic arms or whatever form factor and through the digital space so that they can both work in the physical world as well as work in the digital world. And movie, right? Movies will completely change. Right now it's passive experience as much as it's beautiful, but we're going to change the way we got entertained. So all this requires multiverse. And in teleportation or the teleoperated robots, there's a lot of talk of mining rare earths on asteroids and stuff. You don't need to physically be there if you can operate a robot remotely that's in those spaces.

44:35And what you're talking about is, again, creating explicit representations of 3D space that people can experience. How much in your models does the model itself understand the spaces that it's internalizing before it explicitly projects them? I mean, and that's, again, one thing that I think about more than the productization or practical uses of this is working toward an AI that truly understands the world. doesn't just have a representation of 3D space that really understands not only the laws of physics, but what it's seeing and maybe the values or the usefulness of what it sees and ways that it could manipulate the physical world.

45:52How much of that understanding do you think is there and what needs to happen for the models to really understand the world. Yeah, okay. Great question. So this word understanding is a very profound word. You know, when AI understands something, it just fundamentally is different from how humans understand it, partially because we're very different beings, right? humans as a level of consciousness and self-awareness in an embodied body. For example, when we understand something, when we understand my friend is really happy, it's not just an abstract understanding my friend is happy. You actually have chemical reactions in your body that releases happy hormones or whatever chemicals that your heartbeat might increase, your mood might.

46:57So that level of understanding is very different from an abstract AI agent that has the capability of correctly assign meaning and link meaning to each other. For example, in Marble, our model product, you can actually go to advanced mode of world generation. And in this advanced mode, it allows you to edit. You can take a preview of the world and say, I don't like this couch being pink. Change it to blue. And it changes it to blue. does it understand at the level of blue couch and the word change it does because without that understanding you wouldn't change the pink the couch to the blue couch does it understand it that that the same way you and i understand it including everything about this couch couch including every useful or not even useful information of the couch?

48:10Does it have a memory of the couch? Does it take the concept of couch to afford us, to many other things? No, it doesn't. As a model, it's limited to allowing you to do whatever it's necessary the model needs to do, which is to create a space that has the couch that looks blue. So I think, so the answer to your question is, I do think AI understands, but let's not mistaken that understanding from a anthropomorphic human level understanding. Yeah, and in fact, whatever understanding there is, when you tell it to swap a red couch for a blue couch, that's semantic. That's semantic understanding. It's not understanding at the level of, you know, light and hitting the retina and knowing and having a concept of blueness as opposed to redness.

49:21You know, I saw your talk with Peter Diamandis and Eric Schmidt in Saudi Arabia. And I think you were wonderful in that, much more grounded than some of the questions or even Eric's views. but one of the things that struck me is there was a brief discussion about the potential for ai to be creative or to to help in uh scientific research and the the analogy that is given uh you know if uh if there had been ai and you know before some of them like einstein's relativity theory, or I can't remember the other one that was referenced, uh, could AI reason to that discovery? Uh, what's missing for, for AI to, uh, to be creative.

50:30I'm not talking about creative in the arts. I'm talking about, uh, in the scientific reasoning. It seems to me that that should be within reach. I mean, I'm, again, with no time frame, but that seems just intuitively something that should be possible. That's a great question.

50:58You know, I would think we're closer to AI deducing double helix structure that AI formulating special relativity. Partially because, I mean, we already have seen a lot of great work in protein folding, partially because deducing double helix structure is still the representation, is more grounded in space and geometry. Whereas the formulation of special relativity is on the abstract layer that is not just expressed in unlimited amount of words. That's just not how it works, right? Everything we see in physics from Newtonian law to quantum, we abstract it to a causal level that the relationship of the world, the concept, whether it's mass or it's force, or is abstracted at a level that is just no longer pure statistical pattern generation.

52:34Like language can be very statistical. 3D worlds or 2D world, doesn't matter, can be very statistical. Dynamics can be very statistical. But the causal abstraction of forces and masses and magnetism and all this are not purely statistical. It is very profoundly causal and abstract. So I'm more in the pontification mode than answering your question mode. I think Eric and I on stage were saying, we've got plenty of celestial body data, movement data in the world now. Just aggregate all the satellite data and all that. Give it to today's AI. Can it come up with Newtonian law of motion? See, yeah.

53:37And actually, I found that really, that was the other example I was trying to think of. Relativity. Yeah, there's no physical system that you can observe, at least in day-to-day life, that would lead to that deduction or inference. But with the data of the movement of celestial objects, I'm not sure. Just again, intuitively, I'm a journalist. It seems that maybe not today's AI systems, but that an artificial intelligence could deduce the laws of motion. If given the data and given the time to think, why do you think it would not be able to deduce those laws? When we say those laws are deduced, Newton has to abstract concepts like forces and mass and acceleration and fundamental constants.

55:02those are at an abstract level that I have not yet to seen today's AI can take that whatever vast amount of data and abstract representation or variables or relationships at that level it's uh there isn't much evidence yet I'm you know I don't know everything that's going on in AI. So I'm happy to be proven wrong. I just haven't heard any work that has done that level of abstraction and in the architecture of a transformer model. I don't see where that abstraction can come from yet. So that's why I'm questioning that. Yeah, yeah. It just from my sense, you know, you're building these internal, abstract internal representations.

56:03And as that knowledge accrues, I just, I don't see why an AI model wouldn't be able to apply the, you know, rules of knowledge. knowledge. I'm not saying AI shouldn't or shouldn't try, but it probably takes more progress in our fundamental architecture of our algorithms. Yeah. Yeah. And that was something I wanted to ask. I mean, you use transformers in your models. And I've been talking to people about post-transformer architectures. you know we're certainly not at a breakthrough but you know you we started talking about this this predictive function universal task function as you called it is is do you have a an expectation that there will be that kind of a breakthrough, a new architecture that will unlock some of these capabilities?

57:19I do. I do think we will have architectural breakthroughs. I do not think Transformer is the last invention of AI. and you know in the grand scheme of things humanity hasn't been around for that long compared to the entire history of the universe we know of but in this short history of thousands of years we have never stopped innovating so I do now think Transformer is the last algorithm architecture of AI In one of the conversations or something you wrote. I saw you were talking about the history of your research and that at one time you felt if you could get an AI system to label or caption images, that would have been the pinnacle of your career.

58:21And, of course, you've blown past that. What do you imagine as the crowning achievement of your future career today? I do think unlocking spatial intelligence, creating a model that really connects perception to reasoning, spatial reasoning, seeing to doing, including planning. and imagining to creation would be incredible, a model that can do all three of that. I have a couple of silly questions I'm going to ask you. What's your favorite food or one of your favorite foods? I'm just curious. I mean, you were born and spent some of your early years in China. Is it Chinese food? Is it, you know?

59:19So I'm married to an Italian, and he's a phenomenal cook, in addition of being a phenomenal AI scientist. So I love Italian. And in our family, when we are being nonpartisan, we love Japanese food. But these days, yeah. And then what's your guilty pleasure? Do you read romance novels on the plane or binge watch, you know? Yeah, that's a great question. First of all, I don't have too much leisure time because what I do is so much fun. But I do have a guilty pleasure. I don't think I've ever told anybody. if I'm really tired in a plane, I watch Big Ben Theory. I love that show. I graduated from Caltech.

1:00:22I was a physics major. Everything of that show, I just identify so much with the people there. I don't binge watch because it depends on the duration of the flight, but that's my totally if i'm totally exhausted i watch that yeah yeah and you get the jokes right because i can tell you i love every single jokes and every single character the nerds in the show

From the publisher

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents. 

Visit https://agntcy.org/ and add your support.


How will AI evolve once it can understand and reason about the 3D world, not just text on a screen?

In this episode of Eye on AI, host Craig Smith speaks with Fei Fei Li about the rise of spatial intelligence and the world models that could transform how machines perceive, imagine, and interact with reality.

We explore how spatial intelligence goes beyond language to connect perception, action, and reasoning in physical environments. You will hear how models like Marble build consistent and persistent 3D spaces, why multimodal inputs matter, and what it takes to create digital worlds that are useful for robotics, simulation, design, and creative workflows. Fei Fei also explains the challenges of long term memory, continuous learning, and the search for training objectives that mirror the role next token prediction plays in language models.

Learn how spatial reasoning unlocks new possibilities in robotics and telepresence, why classical physics engines still matter, and how future AI systems may merge perception, planning, and imagination. You will also hear Fei Fei's perspective on the limits of current architectures, why true understanding is different from human understanding, and how world models could shape the next generation of intelligent systems.


Stay Updated:
Craig Smith on X: https://x.com/craigss 
Eye on A.I. on X: https://x.com/EyeOn_AI

More from Eye On A.I.

All 266 episodes
#303 Fei-Fei Li: Spatial Intelligence, World Models & the Future of AIEye On A.I. · 1 h 1 min
Listen in VO