In short
Robotic “foundation models” for general-purpose robot control—what the buzzword means, where the training data comes from, and why Physical Intelligence emphasizes cross-embodiment learning over humanoid-only approaches. Sergey argues that broad, diverse embodied data (from many robot platforms and tasks) creates a “foundation” that can be fine-tuned with smaller, higher-quality datasets or improved via reinforcement learning.
Guest backgrounds
Sergey Levine is a UC Berkeley professor and co-founder of Physical Intelligence. His work focuses on reinforcement learning algorithms for optimal decision-making and robotics.
Key claims
“Foundation” implies broad common-sense physical understanding from diverse, possibly low-quality data; downstream skills need less high-quality data once the foundation exists. Simulation is useful for edge cases but not a substitute for real-world diversity. Teleoperation is a starting point, but scalable supervision should shift toward instructions and autonomous experience. World models and VLAs aren’t a strict dichotomy; they use different abstractions as needed.
Notable examples
RTX project (data from ~30 labs; generalist model ~50% more successful than each lab’s specialized model). Box-assembly failures in the real world (extra boxes, torn items, misplaced phones) show why generality beats narrow specialization. Mobile robots trained with only 3% mobile data still generalized to home kitchen cleaning/put-away.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Foundation Models
0:46 to 2:41
Discussion on the meaning and implications of foundation models in robotics.
“It orchestrates hundreds of smaller submodels purpose-built to understand the nuances of voice, like tone, timing, stress, and intent.”
Sergey Levine's Background and Work
2:57 to 6:32
Sergey Levine shares insights into his work on reinforcement learning and robotics.
“Like it might be just data harvested from the web.”
Challenges in Robotics and Simulation
6:33 to 10:11
Discussion on the challenges of using simulation in robotics and the importance of real-world data.
“And to build that model, I mean, to collect that data, there are various strategies.”
Data Collection Strategies for Robots
10:12 to 14:01
Exploration of various methods for collecting data to improve robotic models.
“Although couldn't you use simulation, as I mentioned with the Boston Dynamics robot, to get over the hump and then start collecting real, real life data?”
Reinforcement Learning and Robotic Learning
14:01 to 17:38
Learn how reinforcement learning improves robot performance through collective learning.
“already gives the robot a lot of learning signal and can improve a policy without direct teleoperation.”
The RTX Project and Generalist Models
17:39 to 19:45
Discover how the RTX project showcased the power of generalist models in robotics.
“So this was a project that we did in 2023.”
Cross Embodiment Learning in Robotics
19:46 to 21:44
Understand how cross embodiment learning enhances robot adaptability across different tasks.
“Now, I should say that generalizing to an entirely new morphology that was not seen in the data, that is still kind of at the bleeding edge of current research.”
Teleoperation and Imitation Learning
21:45 to 26:32
Explore how teleoperation data can enhance robotic learning and decision-making.
“Basically, instead of using the data to answer the question of how should I do this, use the data to answer the question of what is possible.”
World Models and Predictive Ability
26:33 to 28:00
Examine the role of world models in developing generalized robotic capabilities.
“And I think it's where we'll see a lot of future developments.”
Understanding World Models in AI
28:00 to 29:46
Learn about the concept of world models and their significance in AI predictions.
“Typically, what people mean when they say world models is some kind of predictive model that operates at the level of raw observations.”
Show all 22 chapters
The Role of Abstraction in Robotics
29:46 to 31:03
Explore how different abstractions impact robotic behavior and decision-making.
“models view, but I also suspect that the dichotomy between language models and what people today call world models is not as large as some people might see.”
Models and Robot Communication
31:03 to 33:43
Discover how AI models communicate with robots and the challenges of on-device processing.
“So the way that our models work is they actually perform inferences at different levels of abstraction.”
Training Robots: Techniques and Challenges
33:43 to 38:30
Understand different training techniques for robots and their applications in diverse environments.
“Like these are things that are very important.”
Vision Language Action Models Explained
38:30 to 41:46
Gain insights into how vision language action models integrate visual inputs with robotic actions.
“shape of the stage, you're just worried about the robot's body.”
The Future of Versatile Robotics
41:46 to 42:06
Discuss the potential of robots to perform multiple tasks using advanced models.
“Well, and I should be very upfront about this.”
The Challenge of Generalization in Robotics
42:06 to 43:34
Learn about the complexities of teaching robots to perform various tasks in real-world situations.
“So roughly like half a second of joint angles to hit.”
The Importance of Versatility in Robots
43:34 to 45:54
Discover why having generalist robots is crucial for functioning in unpredictable environments.
“that rather than trying to build this generalized model, just train it to do one thing.”
Future of Robotics and Humanoids
45:54 to 47:28
Explore the potential and challenges of humanoid robots in various applications.
“do one thing, and I think that's a lesson that's been learned time and time again in robotics.”
The Need for AI Foundation Models in Robotics
47:28 to 49:38
Understand the importance of foundational AI systems for advancing robotic innovation.
“I don't know, but, but I think that one, like one vision of this that's very appealing to me is that, you know, with robots, since we get to build them, we can build them however we want.”
Evaluating the Reality of Humanoid Robotics
49:38 to 51:45
Discuss the hype surrounding humanoid robots and the challenges they face.
“And on the humanoid question, I was asked to sort of write a brief for somebody late last year or mid last year about humanoids.”
Building Continuous Learning Robots
51:45 to 55:40
Learn about the vision for creating robots that continuously improve through experiences.
“So I have two things I can say about this.”
The Importance of Creative Thinking in Science
56:02 to 57:38
Explore how creative thinking can enhance scientific and engineering work.
“you know is kind of silly or or a waste of time but you enjoy it nonetheless i remember you were you were a big video game guy.”
Transcript
Automatic transcript. May contain errors.0:00A lot of startups these days say they're building foundation models for robots. What does that actually mean for a non-technical listener? I think every time that people try to take robots out of the factory and into open world environments, they very quickly realize that in the real world, there's a huge range of things that can happen. But since then, there's been this explosion of humanoids, and everyone's talking about humanoids. I mean, did I get it wrong? Do you think that humanoids are much closer to being in the world? Most AI is just speech-to-text, plus a language model. It's full for reading transcripts, not understanding conversations.
0:38Velma from Modulate, an AI built on ensemble listening model architecture, specializes in audio analysis. It orchestrates hundreds of smaller submodels purpose-built to understand the nuances of voice, like tone, timing, stress, and intent. perfect for fraud defense, deep fake detection, agent attrition prevention, or customer service moderation. Check out the live Velma preview at preview.modulate.ai. That's preview.modulate.ai. To see how the model breaks down audio, providing timestamped, explainable signals. Stop transcribing, start listening with Modulate.ai. My name is Sergey Levin. I'm one of the founders of Physical Intelligence.
1:37I'm also a professor at UC Berkeley. And what I work on these days is algorithms for reinforcement learning for optimal decision making, as well as applications of robotics. And something that I've been very interested in lately, in particular, is robotic foundation models. These are general purpose models that control any robot in principle to perform any task. And I think we've seen some pretty dramatic transformations in the last few years in the capabilities of these kind of generalist robotic systems, where we can use very diverse data sources for many different robotic platforms performing a wide range of different tasks and acquire a kind of general physical understanding from these data sets that then make it much more feasible to rapidly acquire effective and robust and highly generalizable robotic skills.
2:26So this is something that I've been very interested in the last few years. I think it's an area where we see a lot of progress. Yeah. And a lot of startups these days say they're building foundation models for robots. What does that actually mean for a non-technical listener? Yeah, this is a, it's actually a surprisingly nuanced question because after the success of ChadGPT, you know, the term foundation model became obviously very much a buzzword. So, you know, in some cases, it's almost synonymous to saying like, you know, I have a good model. It's foundational. But I think that insofar as there's a consistent definition, it's something like this, that the principle behind language models, vision language models, things like this, is that you can use very large and diverse data sources that are not necessarily of extremely high quality.
3:15Like it might be just data harvested from the web. and you get a model that digests all this data and acquires a kind of broad and general understanding of how the world works. And this kind of understanding is not enough to be an expert. It's not enough to be extremely proficient, but it gives you that kind of basis of common sense. And that's why the term foundation model was coined by Percy Lang and his colleagues at Stanford for precisely this reason, because this kind of broad basis of knowledge gives you a foundation on top of which you can then put other things. And the key thing about the foundation is that in order to be useful, it needs to be broad.
3:50Insofar as there's a deep technical insight here, the insight is that if you need such a large amount of data, it's almost impossible to get in a single domain. But if you are willing to use data from any different sources, maybe all of the text data on the web, all of the image data you can get your hands on, or in our case, data from all the robots that we've seen, that could be big enough. And that can give you that foundation. And on top of that foundation, now you can build up individual skills. In the case of language models, you can fine tune them for expert level computer programming. In the case of our robots, we can fine tune them for, you know, like assembling things or making coffee or cleaning the kitchen with very high quality data, but a much more limited amount of it.
4:32Because with that, once you can put it on top of that foundation, you don't need a huge amount of data of very high quality for the downstream tasks. So that's the really important thing. Now, to come back to your question, you asked, well, there are many startups, many organizations that are building foundation models. I think one of the most important things about a robotic foundation model is to answer the question, where does all that really broad and diverse data come from that can establish that foundation? This is a place where the answer is very, very delicate. it. The strategy that I think will be most successful here is to not be too picky.
5:15In the same way that language models are trained on all the text data that can be mined from the web, a robotic foundation model should be trained on all of the embodied data that we can get our hands on. And it's one of those things where once you cross a certain threshold of scale, it actually becomes easier to incorporate other data sources. So if you want really, really high quality data of like one particular high-end humanoid robot. That's pretty challenging because now you're very constrained. You have to have that system. You have to get good data for it. You have to figure out how to teleoperate it, put it in the right environments and so on.
5:47But if you're willing to pull in everything, then you can pull in lots of robot data from many different kinds of robots. Some of them might be good. Some of them might be bad. Once your model understands diverse physical embodiments, you can also start adding in data from humans because to the model, the human One body will look like yet another robot body. So this kind of diversity actually makes it easier to include other data sources. So no, no, no, no, go ahead. I was going to bring this back to your question for the final conclusion, which is now to actually answer what you asked me. I think a big difference between how physical intelligence is approaching this and how most other research labs approach this question is that we are not being very picky about which robots we use for this.
6:32We're bringing in everything and trying to build this very broad foundation. Yeah. And to build that model, I mean, to collect that data, there are various strategies. One is simulation. And as I recall from our past conversations, you're not a great fan of simulation. but I just saw 60 minutes a little while back and they had, I'm trying to think, which was Boston Dynamics in a Hyundai factory and they were showing the simulation of sort of an endless army of Atlas robots performing tasks. uh why are you not a fan of a simulation that's that's one and then two you're using uh vision language uh action models is that right uh and can you talk about how that differs from other strategies like uh world models or or uh
7:46other kinds of foundation models, you know? Yeah, for sure. So in regard to simulation, what I would say is this, that simulation is a very appealing tool for kind of very easily acquiring lots of data of a robot doing all sorts of different things. But it's not a very appealing tool for getting experience of very diverse environments and very diverse objects. So if we look at kind of the domains in AI where simulation has been successful and the domains where to struggle to get adoption, computer vision is an area where simulation has actually been used very little despite a lot of attempts. Why?
8:27Well, it's not actually because rendering images is hard. In fact, like computer graphics is very, very advanced so we can render very realistic images. It's that getting real images is so much easier, right? Like you can just take a camera and go and photograph stuff and get lots of real images. And I think that in robotics, a kind of a mental trap that people sometimes fall into is that they say, well, maybe in robotics, it's hard to get data. It's not as easy as taking a camera and going out and taking pictures. And I think this is a little bit of a mistake because actually, if you're serious about building general purpose robots, they'll go out into the world and do lots of things.
9:00The, the kind of boundary condition is in your favor, meaning that the better you get at at building generalist robots, the more robots there are and the more data there should be coming in. There's a little bit of like an initial activation energy problem where you have to get over that hump to get enough systems out there. But that's like a transient period. Once you get over that, then you have lots of robots out there and lots of data coming in. So it actually to me makes a lot more sense to pay a little bit more of that upfront cost to kind of like force it over that threshold and then get lots of real world data.
9:31That doesn't mean that we shouldn't use simulation at all. It just means that we shouldn't worry so much about how hard it is to get robot data. For robot data, we should treat that as the industrial problem that it is, get robots out there, get the data coming in. And then simulation can be very useful for addressing other edge cases. For example, you can simulate, you know, this is what the autonomous driving folks do all the time. You can simulate cases that you don't want to experience in the real world. Like you can simulate a car collision, but you don't want to experience a car collision in your car.
9:56So there's a lot of these kinds of edge cases that you might want to take care of with simulation, But I don't think we should think of it as a substitute for real experience because in other areas where diversity has been critical, like computer vision and natural language processing, real data has been essential to get there. And further, once you have that real data, it's actually easier to incorporate other data sources. Yeah. Although couldn't you use simulation, as I mentioned with the Boston Dynamics robot, to get over the hump and then start collecting real, real life data? Yeah. Yeah, that's a very good question.
10:30I think you could. I think in practice, in robotic manipulation, we found it to be a lot easier to do that with real data. Part of it has to do with the fact that simulation is very good for simulating the robot. It's much less good for like simulating everything else because the robot you only have to model once where if you, whereas if you want to have like thousands of scenes, thousands of objects, then you have to go and, and, and model and simulate each one of those. I don't think that's impossible. I think it is a tractable problem potentially, but it's just more costly than just going out and getting lots of real stuff.
11:02Yeah. And, you know, there's been this explosion of humanoids. Yours are not necessarily humanoid. I mean, you're kind of platform agnostic. Is that right? But on the humanoids, Joanna Stern at the Wall Street Journal last year had NEO, invited NEO into her home, and it didn't perform very well. And apparently NEO is available for home use, but it comes along with a teleoperator that spends time in your house, like walking around doing something. And that's to collect that data to train the robot. I mean, they need a lot of these deployments to collect enough data. That doesn't seem like a very scalable solution.
11:55So how are you guys approaching it? Is it, are you doing any teleoperation or, you know? Yeah. So this is a very good question. And in fact, in some ways, this is like kind of the big question in modern robotic foundation models, which is what kind of data can you use? Now, I think it makes a lot of sense to start off building the initial foundation with teleoperation data. But as you pointed out, there are major challenges with this, like, you know, not the least of which is that if you actually want to collect data in deployment scenarios, whether it's somebody's home or a business or a factory or a warehouse or whatever, like that's an additional kind of inconvenience that you have to deal with.
12:38And it's a barrier to scale. I think that the right way to proceed with this is to think of it as a mixture of different data sources where as your model gets better, it should be able to leverage more accessible and more scalable data sources. So initially, when the model is not very good, We need data from teleoperation from humans that basically illustrates like this, how the robot should act. But once the model gets better and the robot can be deployed with at least some degree of autonomy, then we can handle a more accessible source of supervision. One more accessible source of supervision is instructions.
13:14This is something that we actually found, not entirely intentionally, like we were just kind of like, you know, trying out a few things in some of our research projects, but we found that we could actually get improvement in our policies by supervising the robot essentially through language. And this only started happening once the model became powerful enough that the low-level skills were already pretty good. Then you could correct the robot and say like, oh, maybe it's cleaning up the kitchen, it messed up. You could say like, oh, you needed to pick up the plate and make sure you put the plate in the sink.
13:43And the way the model works internally is it's very similar to how these modern reasoning models, LLM's work, where there are internal thoughts that are generated and then the final action is chosen based on those thoughts. So essentially this kind of language feedback supervises the internal thoughts rather than the low-level actions. But once the low-level actions are good enough, actually supervising the internal thoughts already gives the robot a lot of learning signal and can improve a policy without direct teleoperation. But again, this only emerges once the base model is strong enough. The other thing we can do is leverage autonomous experience where we can improve the system through reinforcement learning.
14:19So there was a research project that we actually published just a few months ago that describes a reinforcement learning system that we built on top of our foundation model. And again, it's the same story that you need the foundation model to be strong enough so that from there it can improve with autonomous experience and reinforcement. And one time, as I recall, you had, I didn't see it, but you were setting up ranks of robotic arms and they were going to be sort of informing each other. So any arm that learns something, then that learning would be transferred to all of the robots. Is that kind of scale useful?
15:11I mean, are you doing that with, yeah. Yeah, that's exactly right. So the The big benefit of having robotic learning rather than human learning here is that you can have this fleet effect and you can do collective learning very effectively. So all of our robots share all the experience. And again, it comes back to this idea that the stronger the base foundation model is, the more readily it can incorporate experience from diverse robotic platforms. So the experiment that you're referring to, this was done at this point, maybe about eight years ago, the Google Arm Farm project. Like there, every single robot platform at every single robot station was as close as possible to each other.
15:52They were virtually identical. And that worked very well with like the state of the art learning technology of, you know, 2017, 2018. But these days when you have a foundation model that accommodate very diverse platforms, very diverse tasks, very diverse environments, the collective learning stuff becomes much easier because now we can pull in data from many of our own robots that are deployed. we can also work with other companies that have their own robotic platforms. And in fact, initially when we started this, we thought that we would need to do something very special to accommodate this.
16:24I mean, we need to somehow tell the model, like here's the morphology you're controlling. Here are the details on this robot platform. It actually turns out that very little cleverness is required in that respect. Like the model can figure out from looking through the camera at the camera image, what kind of robot is dealing with. A lot of careful engineering is still needed to make sure that training is efficient, to sure the model is set up in the right way, but handling the cross embodiment aspect of this turns out to actually be pretty straightforward. Yeah. And across form factors as well, is that right?
16:57Not only tasks. That's right. Other companies we've worked with have actually adapted our models for controlling multi-fingered hands, humanoid, that sort of thing. In fact, some of them can use them for mobile robots, things like agricultural equipment that we wouldn't conventionally think of as robots in the usual sense. Yeah. Wasn't there a project that you were involved in that brought together data from all around the world, from all various robotics labs? And what was the outcome of that, or is that ongoing? Yeah. So what you're referring to, I think, is the RTX project, which in many ways was actually part of the impetus for starting physical intelligence.
17:39So this was a project that we did in 2023. And myself and many of my colleagues that worked on this then went on to found physical intelligence. In the RTX project, this was very much an academic research project, but what we did is we contacted academic research labs, about 30 labs in total. And we asked them to basically send us the data from their robotic manipulation experiments. And we limited this to single arm robot manipulators with parallel jaw grouper, just to pick kind of the most common form factor. And then what we did is we trained one model across all of these different data sets.
18:14And we sent back that model to some of the labs that had donated data and asked them to essentially evaluate it in comparison to whatever they were developing on their own robot for their own application. So each lab was doing a different research project with a different robot and a different task, and they had their own methods that they were developing. And we just said, like, whatever is the best you've got, just measure that against our generalist model. And what we found is that the generalist model on average was about 50 % more successful than whatever each individual lab was developed.
18:43And that's really, really exciting because this is kind of paralleling a lot of the development that we've seen in language models. With language models, the big result scientifically, it wasn't actually ChatGPT. The scientific result that was so exciting is that the generalist model, the generalist language model, they outperform specialized models for machine translation, sentiment analysis, you know, all these NLP tasks that typically would require very specialized data sets and very specialized models could be done better with this more general. But what we saw with RTX was an early hint that something like that was actually happening in robotics.
19:18And I think that's actually really important. Yeah. But in RTX, you're using exclusively robotic arms. Does that data inform other form factors? Or do you need, if you're training humanoids, do you need to collect data across various different humanoid platforms? Yeah. So the current models we have at Physical Intelligence are trained on many more robots, and they do vary in morphology. Now, I should say that generalizing to an entirely new morphology that was not seen in the data, that is still kind of at the bleeding edge of current research. So we're not really that concerned with that. We're concerned with the case where someone has a robot, they have some data from that robot, but what they want is to benefit from transfer from other robotic platforms.
20:06You know, here's an anecdote that I can tell you about that maybe underscores this point. In the first year or so of the company, this was in 2024, we worked almost entirely with static arms. So these are not mobile robots. These were arms that were attached to a table and they were doing manipulation tasks. And then in early 2025, we decided that we wanted to start experimenting with mobile robots. So these are basic arms on a wheeled base. And we had very few mobile robots, so we could only collect a little bit of data with them. And our first kind of publicly released research project on this, which came in April, 2025, used a training set of which only 3 % of the data was collected on mobile robots.
20:52So 97 % was from these statically mounted arms, but we could get the mobile robots to actually generalize very broadly. They could go into a home that was never seen in the training data, clean up the kitchen, you know, put away the dishes, that sort of thing. And most of the knowledge in these models came from not the mobile robot, but these static arms bolted to a table. So that kind of underscores the power of this kind of cross embodiment learning where we could use lower cost, more accessible platforms to get the bulk of your data and then adapt them to a downstream morphology. Yeah. Uh, that's, that's fascinating.
21:23Uh, although I imagine when you move to bipedal, bipedal, uh, for legged robots, uh, you're, you're going to, that data, it'll help in the robot arm manipulation, but not necessarily in mobility. So, and this is a kind of an uninformed question, but when you see teleoperation, uh what is happening there and what kinds of things are robots learning is is that uh uh imitation learning is is that uh something else so um the standard kind of default way to use teleoperation data is imitation learning but this is i think a place where there's a lot of room for improvement in current research and we've started studying this a little bit there's other research groups that are studying this, basically what you would like to do with data ideally is not just copy it blindly, but actually understand dynamically which parts of what is being demonstrated are good and which parts are not so good.
22:35Basically, instead of using the data to answer the question of how should I do this, use the data to answer the question of what is possible. And then among the things that are possible, pick like the best things. Uh, and that's basically where a lot of the ideas in offline reinforcement learning can come in. Roughly speaking, the way that this works is instead of supervising the model to produce the same actions that are in the data, what you do is you supervise the model to predict the outcomes. So you train the model so that it can predict, like, if I see this and I do this, will that be good or will that be bad?
23:06And if you can do a really good job predicting those outcomes, then you can tell the model, okay, now do whatever will lead to the good outcome. And that's actually potentially a very powerful tool because now you can bring in heterogeneous data of different quality. And now the variety of data quality actually becomes a blessing rather than a curse. Because if you see lots of good things and lots of bad things, then you can figure out how to distinguish good from bad and do better at test time. So this is kind of where a lot of the current research is situated. I see. How does, I mean, does teleoperation, is that using a VLA a model in the background.
23:43Is that collecting data for the vision language action model? That's right. So, so our models are based on vision language action models. And this is, this has kind of become essentially like a de facto standard in robotic learning research. This is something that many of the folks on the team here pioneered back in like the early 2020s, but now it's basically what everybody uses. VLAs are kind of an interesting thing because initially, like the early, what I refer to as first generation VLAs, they were trained in a very straightforward way. Basically, vision language models are models that answer questions, and they can also take in an image.
24:21So they answer visual questions. Early VLAs were trained by basically taking this visual question answering paradigm and simply turning robotic control into like a visual question. So in robotic control, the question is the prompt, like pick up the socks, and the answer is the numerical value of the actions. So that's like a fairly straightforward, naive way to cram robotic control into a format that vision language models can understand. But there's a lot more that we can do than just that. And there's broadly speaking two big buckets where there is room for improvement with VLA's. One is dexterity and the other one is reasoning.
24:55So dexterity means go beyond treating actions as an answer to a visual question and actually develop a model design that handles dexterity first and foremost. So control is not a discrete thing. It's not, it's not an answer to a question. It's a continuous thing. It's a trajectory. So you can use models that are very well adapted to high dimensional, continuous dynamical systems. Diffusion models are really good for this. So incorporating diffusion models into, uh, vision language models can give you these kind of much more dextrose VLMs. The second thing is the knowledge that is learned by language models and visual language models from the web.
25:35A lot of that knowledge is semantic. it's not physical. So a big way to improve VLA's is to better hook into that semantic knowledge. Essentially, when the robot doesn't know what to do, what it should do, much like a person, is it should pause and think. And that thinking maybe taps into more semantic knowledge that is not yet fully grounded in the physical world, but can lead to reasonable inferences. So maybe it's trying to open a drawer to take out a knife to cut a vegetable, and the drawer isn't opening. Now, maybe the robot experience is not enough to inform what to do, but there's a reasonable semantic everything you say, well, why isn't this thing opening?
26:12Maybe I should try a different one. Like there's kind of this common sense inference you can make. And that common sense, like if you ask like chat GPT, it can make that common sense inference that it will tell you something. And the trick then is to digest that inference into a format that the motor control component of this can understand. That's basically a thinking process. So that's where I think there's room for a lot of improvements for these models. And we've seen a little bit of that in some of our work on chain of thought. And I think it's where we'll see a lot of future developments.
Read the full transcript
26:38Yeah. You know, I had Fei-Fei Li on recently talking about world models. And in the past, I've had Yan Le Koon on several times talking about world models. Where do world models fit into the stock? I mean, are you, because world models, as I understand it, certainly from Yan Le Koon's point of view, It's building an internal representation of the world that then can be used to predict futures or reason through problems. If you don't have that, it seems that it would be much harder to build a foundation, robot foundation model that can generalize and operate in the wild. I mean, right now in your lab, in a lot of the industrial deployments, they're very controlled environments and the robots training is very task specific.
27:47But if you want a more generalized model, where do world models fit in? Yeah, it's a really interesting question. I think that some folks tend to present world, like, I guess we should nail down what we mean by world models. Typically, what people mean when they say world models is some kind of predictive model that operates at the level of raw observations. It doesn't mean that it predicts raw observations. It may be that it is, you know, like Jan LeCun, for example, I know he advocates for essentially a latent space world model, which predicts a sufficient statistic of observations. But roughly speaking, it's something that predicts something about your future observations.
28:25It's a very reasonable idea, but I think that something that we should keep in mind is that for human behavior, like prediction definitely plays a role. But there are also things that we do that are not grounded entirely in prediction. There's a place where prediction is easy and there's a place where prediction is difficult. And the abstractions that we use are really critical to intelligent behavior. So to give you an example, if I want to figure out how to get from where I am now in San Francisco to New York City, maybe I'm going to imagine something about it. Maybe I imagine how I get my car keys and I get in my car.
29:01I might imagine that I'm taking an airplane. But the further out that I think about this, the more abstract that imagination becomes. I'm not imagining exactly what my seat in the airplane is going to look like. So abstractions are really key to actual effective world model. And I think at some level, what language models do, what visual language models do, and what video prediction models do and other kinds of world models is not actually that different. They're just operating with different abstractions. And I think for a real capable embodied intelligence system, like a robot, we'll need many different abstractions.
29:36And I suspect that once we figure out how to use abstractions in general, like that general part of the question is actually the important one. And then we'll use kind of the right kind of thing at each level and that'll be fine. So I guess what I would say is that like, definitely I'm very sympathetic to the world models view, but I also suspect that the dichotomy between language models and what people today call world models is not as large as some people might see. Yeah. So, so with a good world model, there isn't anything that a robot couldn't do with being trained through with a vision language action model.
30:11It's just, yeah, go ahead. I suspect the key is to have a system that correctly uses the right abstraction for the jaw. So if you are, you know, folding a piece of clothing, probably, you know, that's clearly not a semantic task. You're not thinking in words about like, oh, this fold is here, this fold is there. But you're probably also not imagining literally how all the particles of clothing move. At some level, especially if you're good at the skill, you're mostly kind of using your muscle memory a little bit, but it's, it's reactive. It's using perception, but it's, it's not as model based as just like imagining exactly how every particle will move.
30:48So the really cool thing about a proficient human motor skills is that they blend prediction and this kind of model free behavior and semantic reasoning. It all kind of comes together with the right tool for the job at each level. And I think that kind of blending is really, really important. So do your models actually imagine possible futures before acting? Right. So the way that our models work is they actually perform inferences at different levels of abstraction. I think there's still a lot of research to be done to figure out exactly how that should be done and what abstraction should be used and where.
31:20Like, you know, right now, I would say that our choice of abstraction is quite naive, but you can reason semantically about higher level things. Like if you're cleaning a kitchen, do you pick up the plate or do you pick up the towel first? You can figure out where things go spatially and then you can produce the actions. So there's already a little bit of that multiple abstractions, but I think there's a lot more research that needs to be done to get the abstractions to emerge automatically and to automatically figure out the right abstraction for any given stage of task. Yeah. The, the, the, the, so, so world models in your view are not necessary for a generalized robots that can operate in messy environments.
32:00I guess what I would say is not that it's necessary or not necessary, but that it's not as much of a dichotomy as I think that some folks tend to, tend to present. Like, I think that the view that I would disagree with is that there is like model free policies, reinforcement learning, VLAs and world models. And these are like totally separate things. I actually don't think they're separate things. And I think that, uh, the right answer will be a model that can do all of those things together that can use world model like prediction when it's necessary and can do the model free stuff when that's more appropriate and smartly decides what kind of abstraction is the right obstruction to use for a given problem.
32:32Yeah. Your models, the AI brain, so to speak, is not on the robots. It's communicating with them remotely. But when you send robots out into the world, you really can't afford, as with cars, autonomous vehicles, you can't afford for them to be waiting for instructions from the cloud. So how much of the models that you're building can be on device? Yeah, that's a really interesting question. So, you know, so far we're still very much in like the research and development phase. We haven't had to worry about this very much. And you're completely right that currently the models actually live on the cloud.
33:25They've actually run through an inference API that looks very similar to what someone might imagine using for an LLM. But in the long run, I think you're right that there needs to be an on-device component that is reliable, that is not vulnerable to connectivity issues. generally, I think that the way to move towards this, which I think is already reflected in current models that we and others have been developing, is to have a system that performs multiple types of inferences in parallel and at different levels, where the highest levels are maybe more appropriate to offload to a remote inference server, and the lowest levels, the ones that are really doing motor control and closing the loop very tightly in perception, function run locally.
34:09Now, the good news is that the lowest levels are probably also going to be the smallest ones in terms of the number of parameters, because they're not, you know, you can sort of think of these as like, uh, instinctual reactions, reflexes, that sort of thing. Like these are things that are very important. They need to be very fast, but they also are not as cognitively demanding and not as complex. So they can run locally. Now the trick of course is figuring out all that communication so that you still preserve the benefit of end-to-end training. So this whole thing is trained together to active concern, but at inference time can be partitioned in this way with different size components running in different places.
34:44And it's kind of cool to imagine like, you know, you have good internet connectivity, you get, you get a lot of intelligence. Your internet connectivity starts to degrade. Okay. Maybe the robot gets a little dumber. Like it has to, you know, uh, maybe pull down some nice inferences from the cloud, keep them locally and do some stuff. And then, okay, now let's, let's stop and think again. Let's ask the cloud to think some more. So right now you're not concerned about shrinking models to fit on device. The first step is to get a generalized model and then you'll do the distillation or whatever is required to get it on device.
35:21Is that fair? So that's right. But even separately from any inference concerns, we do actually, we've sort of, I guess, conversion evolution in some sense, we already have models that have this kind of multi-scale property where the lowest level motor control components are already smaller. So while we're not so worried yet about running things on device, the natural trajectory of that technology is leading to a place that makes that actually reasonably straightforward. Yeah. So I have some questions about generalization, but before I get to that, I get very confused about all the different kinds of ways that people are training robots.
36:04So there's vision language action models, but there is one shot, two shot, few shot, vision language action, world models. Can you sort of give me a little primer on the different ways that people are training robots and why you've settled on VLAs and RL? Yeah. Yeah. Let me think about how to best describe this. So I think it's, it's often hard to get like a complete picture of the robotic learning world. Just by looking at the kind of results that people present, there's kind of one general truism of the robotics demo, which is robotics demos can be set up in such a way that shows something really cool, but doesn't actually provide like a general solution to a problem.
37:00because if you want to just like stage a robot demo, you kind of make things work in that one setting. So because of that, it's like a little hard to figure out what are like the big clusters of major effective techniques. But if I were to like very soberly look at the current robotic learning environment, I would say that there are actually like two big things that work very well. One thing which we've been discussing is vision language action models. And more generally this idea of like learning manipulational skills and things like that from data. Typically it's with imitational learning, but it can also be with reinforcement learning insofar as it can leverage that data.
37:31And oftentimes actually the recipe is like use imitation learning to initialize and then maybe like fine tune it with RL. And then the other big cluster is sim to real transfer using simulation to learn motor skills in a sufficiently randomized way and then run them in the real world. And when you see demos of like, you know, robots doing like, you know, dancing and acrobatics and all that stuff, that's typically doing sim to real. And it's kind of a funny, unstable equilibrium of the current research world that these two types of techniques are very different, and they are also used to attack very different domains.
38:04So the vision language action models, they're basically currently the dominant paradigm for robotic manipulation. Problems where robots need to interact with diverse environments and diverse options. The simporeal stuff is the method of choice for highly acrobatic and athletic movements, typically for human points. And these are situations where the task is physically extremely demanding, but the diversity of the environment is very low. So if you put the robot on stage and it's dancing, you're not really worried about the shape of the stage, you're just worried about the robot's body. So this maybe tells us actually a lot about the strengths and weaknesses of these two types of approaches, that the simptorial stuff is great for really understanding the physics of the robot, but not so great for generalization.
38:48The vision language action models are great for generalization, but of course, because they're dealing with the physical world, the real world data, they can't run these like giant RL loops that practice for billions of trials and really overfit to the particular body of the robot. Now, of course, there's a lot more to do in the future, but this hopefully gives you some sense for the layout. And VLA, just walk me through the process of where the vision comes in, where the language comes in to get to the action. Yeah, so maybe one way I could describe this is to start with visual language models.
39:22These are also sometimes called multimodal LLMs. So if you use like Gemini or ChatGPT and you upload an image and you ask some question, that's basically a VLM. And the way that these models work today is you start with a language model. A language model is very, very simple. It's a transformer that takes in text and predicts future text. And to get these things to process images, what we do is we train a vision encoder, basically another little piece of neural network that takes an image and puts it into the same space as the language tokens. That's sometimes called a vision encoder. So now you're going to feed images to this thing the same way that you feed language, because this little like virtual visual cortex kind of takes the images and turns them into, you know, semantic like looking thingies that the language model knows how to process.
40:11So now with VLAs, well, with first generation VLAs, as I mentioned, people basically just took exactly this model and just changed the output to literally like output numbers in text that represent actions. The second generation VLAs, which is what everyone's basically using today. They take inspiration from how VLMs add a visual cortex to, to the language model. And they also had a kind of a virtual motor cortex, a specialized little piece of circuitry whose job is to take the outputs from the language model backbone and decode them into continuous actions. And this is typically done with diffusion.
40:44Basically the same kind of technology that's used to generate images and videos is now used to generate trajectories of robot joints. That makes a lot of sense because trajectories of robot joints are a continuous spatial object the same way that images are continuous spatial objects. So it makes sense that the same technology would be applicable there. And I think this is really cool because almost like what we're, what we're doing is like building a brain piece by piece. Like there is this, the language model backbone, I guess kind of like a prefrontal cortex almost. There's the little visual cortex part that encodes images and now there's a little motor cortex part.
41:14Anyone familiar with biology would be very offended at this point because it's backwards, right? Evolutionarily, the motor cortex comes first, then vision, then prefrontal, but here it's the other way around. Yeah. And then, uh, to get the continuous action, it's outputting a stream of numbers that are telling the actuators how far to move or right. And control theory comes in at the end. Is that right? You're outputting. Well, and I should be very upfront about this. The amount of control theory in this is like pretty minimal. So currently the way that most of these systems work, including ours, is that the model outputs a short trajectory of future target joint angles.
42:06So roughly like half a second of joint angles to hit. And the actual actuation, the motor commands to reach those joint angles are computed with a very simple impedance controller. This is kind of like, you know, the most basic type of controller you can put on a robot car, basically. Yeah. You know, past robots were often built to do one thing really well, particularly, I mean, certainly before, you know, they were, AI was involved. How realistic is it that one robot with one of these models controlling it can do many different things? Yeah. And or and I had a conversation with a guy, interesting guy, maybe you run across him.
43:00Mike LeBlanc, I think his name is, he's got a startup called Foundation and he's doing humanoids for the military. and they're training them to do one thing, like put an explosive charge on a door, which is a very dangerous thing for a soldier to do. So they drop it from the Humvee, it walks, slaps this thing on the door, and comes back, hopefully in one piece. And that made a lot of sense. that rather than trying to build this generalized model, just train it to do one thing. But how much generalization is developing from yours and other robots in the field? So here's maybe how I can try to answer this question.
43:57I think every time that people try to take robots out of the factory and into open world environments, they very quickly realize that in the real world, there's a huge range of things that can happen. And it's very, very hard to get a system that is highly specialized and still robust enough to work in the real world. This lesson was probably learned very, like, I don't know if it's for the first time, but an early instance of this lesson was actually autonomous driving. Like in the very early days of autonomous driving, some people thought that, well, okay, driving is complicated. There's lots of stuff that can happen.
44:33But if we just like an instrument the road and like prevent other things from getting into like the autonomous car lane. Maybe then things will work. Like we'll install like magnetic sensors and all that stuff. And that really never took off because essentially the gap between a closed world and an open world is, is enormous. And you can't, it's like, you can't be just like a little open world. Like as soon as you're, you're out in the wild, immediately stuff can happen. Maybe it happens rarely, but that's not, that doesn't save you. Like even if it happens rarely, you have to deal with it. So here's an example from our work.
45:03Um, we had a project on using our systems to assemble boxes. So you take like a flattened cardboard box and you have to like build it up, fold it. And it's like a little bit like an origami problem. Basically you think that's pretty structured. Like it's one thing, just build a box. But sometimes you grab the boxes off the pile and you get two boxes instead of one. So you have to put one in the back and maybe there's someone, something is torn a little bit, so you have to discard it because it's torn and maybe someone left like their, uh, phone on the table. So you have to put the phone away. So there's just so many things that happen.
45:31that even like once you cross that boundary from the factory to the real world, that you can't actually have a narrow specialist. That even if what you want is to do one thing, you have to handle all these other things that can happen all around. And the generalist model, the one that could handle a wide variety of tasks actually becomes a better specialist because it can deal with all that weird stuff that arises. So I think that generality is really essential, even if you really want to do one thing, and I think that's a lesson that's been learned time and time again in robotics. And, and are you going to, uh, move into humanoids or have you?
46:03Um, yeah, that's a good question. So the main reason why we've stayed away from humanoids ourselves for the, uh, the data collection that we're doing here is actually, it's not ideological. It's just because humanoids are expensive, complicated and teleoperating humanoids is harder. So we've basically stuck to robotic platforms that can give us the most data, the most data diversity. But I think humanoids are really cool. Like we've worked with other companies that have humanoids, the model runs on humanoids. And I think that, you know, it's something that I'm sure we'll want to explore more in the future.
46:33My own take when it comes to robot morphology though, is that I think that, you know, this is maybe like a little bit, uh, idealistic, but I, I actually really hope that robots will kind of end up being a little bit like personal computers where there's like general software and the form factor of the device can be very different for different jobs. So when you get a computer, maybe you get a laptop, maybe your computer is your cell phone, basically, maybe it's a big desktop, maybe it's a big server machine. It's like whatever is the right for your job. And I think robots will be the same way that maybe like, if you live in a small apartment in New York, maybe you have like a little home robot that kind of attaches to the ceiling and pivots around and cleans up the thing.
47:09And maybe if you, you know, if you live on a farm or something, you have a big mobile robot, you know, tractor with a bunch of arms attached that can drive around and do stuff. And maybe there'll be one software stack that can accommodate these things with different applications and prompt engineering people do for particular domains. And there'll be this heterogeneity of platforms. like physical intelligence everywhere yeah well how does how do your models work on humanoids because your your models have not been trained on operating legs and walking yeah so typically right now when we work with folks that have humanoid robots uh they have their own software stack that takes care of balance and all that and all those things and then the model drives basically manipulation behaviors yeah uh and and what you were just saying do you think we'll end up with an iphone of robots that everyone uses and uh and maybe there are specialized robot bodies for for different uh tasks uh and would there be do you think there will be a dominant model Yeah.
48:18Yeah. I don't know, but, but I think that one, like one vision of this that's very appealing to me is that, you know, with robots, since we get to build them, we can build them however we want. And I, and I think that's one of the advantages of robots is that they can be designed not to replace people, but to do the, to do things differently from people. And so you could design form factors that are suitable for particular domains. They can be much smaller, much bigger. They can have, you know, five arms, seven arms, one arm, whatever is like appropriate for the cost point that you have, whatever is appropriate for your application.
48:52And people can be pretty creative about it. It's just the thing that's been preventing folks from experimenting with these things is that if you want a robot to do anything, you have to solve the intelligence problem. And if there isn't a ready-made off-the-shelf thing that at least gives you like a rough prototype, you just can't get started. So what I really hope is that good robotic foundation models will provide that kind of middle layer, you know, in a computer, that layer is basically taken up by the operating system. The operating system is the thing on top of which you put applications.
49:19So there's an intelligence layer that's pretty general on top of which somebody can essentially prompt engineer their tasks. And then they can experiment with all the aspects of that application, really nail it correctly, design the right form factor. Then we'll see a lot of experimentation, a lot of creativity. And then I think we'll actually see what robots should really look like instead of just thinking of them as like a metal version of a person. Yeah. Yeah. And on the humanoid question, I was asked to sort of write a brief for somebody late last year or mid last year about humanoids. And I wrote this very bearish thing about how, you know, all of the different problems from, you know, dexterity in the hand to models that can deal with unpredictable environments and things like that.
50:11But since then, there's been this explosion of humanoids. And everyone's talking about humanoids. I mean, did I get it wrong? do you think that humanoids are much closer to being in the world than, than I, than I thought? It's a question. Um, I mean, I think it's, it, it, it, I think there's a good reason to be excited about humanoids and like, in the sense of like, just emotionally, this really kind of captures the imagination. Uh, it also makes it easier for people to think about because like, we know what people can do. So if we kind of build a robot that does what people do, that like kind of makes sense.
50:46But I do think that it's a somewhat limited view to restrict ourselves just to that. Yeah. My sense is that humanoids will have a place. Like I think there will be humanoids that do something that people find useful. But I think there will also be a place for all sorts of other things. And, you know, as with any new technology, oftentimes what's most important to figure out the right way for it to exist in the world is the tools and machinery and structures to allow people to experiment, to prototype, to try out all sorts of different things. And I think the trouble with robots is that just like the barrier to entry for serious open world robotic systems is extremely high.
51:23Like you have to solve open research problems just to get your prototype out the door. And I think that's kind of what's actually limiting a lot of creativity. Yeah. And for example, with Boston Dynamics, but in the past they were pure control theory, but now they've they have an ai partner and they're they're using ai models to control the robots and as i said there's you know there's been a lot of talk of them being used in factories i think hyundai is is uh deploying them um and i understand the the argument that well it's the human form factor the world is built for that and if we can have robots operate in in that form factor with that form factor there's it can they can presumably eventually do everything that humans can do uh how much of this do you think is hype uh it's very hard for a layman when you when you watch videos or, as I said, the 60 minutes piece, to know how real it is and how much is our sort of fascinating demos that in two years people say, oh yeah, yeah, we had that pilot, but it didn't work out.
52:48So I have two things I can say about this. The first is, I think in regard to demos, I mean, I think you're right that for, for a demo it's like, there's a big difference between setting up a demo and setting up something that works. And you know, one thing that I've struggled with in my career, I think quite a bit is that, you know, I work on robotic learning and robotic learning that the purpose of it is to get systems that can generalize and work in open world environments. But it's very hard to illustrate. Like it's, it's, if you, if you show somebody a demo in a particular setting and the robot does something cool, like, yeah, it's obviously cool stuff going on.
53:21If you show doing something fairly simple, but in like a hundred different environments, well, then each of the videos of that is just a robot doing something simple. So the fact that it can do it in all these different settings is harder to convey. And that's kind of a science communication challenge that I think we have to be cognizant of when we look at the videos. But I think there's a deeper, more technical thing I want to say about humanoids and about robots in general, which is that conventionally when somebody thinks about building a really cool robotic system, they naturally start with building the physical robot.
53:55So if you start from the premise, I'm going to build one robot, it makes sense like if it's going to do something general, it should be a very general body. And you should really get that right because once you've committed to like that form factor, like you're kind of stuck. So I think there's actually a technical point here, which is that the old way of thinking about robotic software as something that drives one robot kind of naturally leads to that. Because if you want one very general robot, then you kind of have to like have it have everything. But if you accept the premise that we're going to have robotic foundation models that can drive lots of different robots to do lots of different things now, kind of on unchains your thinking.
54:31And now it's okay to experiment with different form factors, experiment with different applications, different ways of approaching things. It's just that you need this kind of general AI system to be able to do that. And without it, it actually kind of makes sense that you would go into this, you know, single ideal, perfect body plan and maybe a humanoid is a good choice for that. Yeah. And you've made tremendous progress since I first met you. What's on the horizon? What are you guys working? So for me, one of the things that I really want to figure out over the next year, or maybe the next few years, both in my academic work and in the work here at physical intelligence, is how to go from the foundation, which I think at this point, we have a pretty good understanding of, to something that is a true data flywheel, a true continual learning system.
55:17where the robot experiences more and more tasks and the more it experiences them, the better it gets. And that involves, I think a lot of autonomous learning, a lot of reinforcement learning, a lot of learning from other supervision signals, like these language feedback that I mentioned before. And turning that into a continuous cycle where the more the system is deployed, the more capable it gets. And I think there's some very deep technical challenges there that'll stress kind of all aspects of this. So that's what I'm most excited. Yeah. Okay. uh sergey i'm gonna ask one last question i asked feyfei lee this and uh david ha and some others what's your guilty pleasure do you have one oh uh hmm something you do to relax that that you you know is kind of silly or or a waste of time but you enjoy it nonetheless i remember you were you were a big video game guy.
56:13Yeah, I actually, back when I was in college, I actually thought that my career would be making video games. I think these days it's probably science fiction. I'm a big science fiction nerd. I made sure to put like a little epigraph with quotes from science fiction stories that I like into a lot of the papers that we've been putting out. I like that very much because I think it's just, it's very good, I think, especially for somebody working in science or engineering. And this is maybe a bit of advice that I'd have for folks listening to this too, is that it's good to not feel too constrained.
56:43And sometimes the rigor of doing serious engineering and scientific work forces them to forces them to very constrained thinking. So it's very important, I think, to try to get out of that mindset sometimes, not to get too carried away with it, but it's good to like, let yourself think things that are maybe not, not good to do in polite circles in science and engineering to think about things that are more out there that are more fanciful, more creative or more improbable, let's say. Yeah. Is there a favorite work of science fiction that you've read that you would advise people to read? When I was younger, I didn't read very much Robert Heinlein.
57:19So I think I only discovered Robert Heinlein's work in the last few years. I've been really getting into trouble. Yeah, that's wonderful. Classical American sci-fi. That's right. This very optimistic aspect of American culture that I think is very refreshing especially in today's day mage yeah okay sergey well i hope we have a chance to talk again and i'll see you at the conferences and this has been fascinating most ai is just speech to text plus a language model it's full for reading transcripts not understanding conversations velma from modulate an ai built on ensemble listening model architecture specializes in audio analysis.
58:00It orchestrates hundreds of smaller submodels purpose-built to understand the nuances of voice like tone, timing, stress, and intent. Perfect for fraud defense, deep fake detection, agent attrition prevention, or customer service moderation. Check out the live Velma preview at preview.modulate.ai. That's preview.modulate.ai to see how the model breaks down audio, providing timestamped, explainable signals. Stop transcribing. Start listening with modulate.ai.
From the publisher
What does it actually mean to build a foundation model for robots?
In this episode of Eye on AI, Craig Smith sits down with Sergey Levine, co-founder of Physical Intelligence and professor at UC Berkeley, to explore a fundamentally different approach to building robots, one inspired not by programming a single perfect machine, but by training AI on the broadest and most diverse data possible so robots can learn, adapt, and operate in the unpredictable real world.
Sergey explains why the secret to general-purpose robots isn't perfecting one single machine, but training on massive, diverse data from all kinds of robots and even humans. The more variety the model sees, the better it gets. Just like ChatGPT learned from all the text on the internet, robotic foundation models learn from every robot that has ever moved, grabbed, or interacted with the real world.
We also get into the big humanoid robot debate. Are they the future, or is it mostly hype? Sergey gives an honest and technical take on why the form factor conversation is changing now that foundation models exist, and why that actually opens the door for more creativity, not less.
Finally, Sergey shares what he's most excited about next, building a true data flywheel where robots get smarter the more they are deployed, creating a continuous learning cycle that could change everything.
Subscribe for more conversations with the people building the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Introduction: What Are Foundation Models for Robots?
(01:44) Meet Sergey Levine: Physical Intelligence and UC Berkeley
(02:51) Breaking Down Foundation Models for Non-Technical People
(06:46) Why Real World Data Beats Simulation
(15:00) Building a Broad Robotics Foundation From Scratch
(24:00) The Open World Problem in Robotics
(40:00) Generalist vs Specialist Robots: Which Wins?
(47:00) Humanoid Robots: Real Innovation or Just Hype?
(55:10) The Future: Continuous Learning and the Data Flywheel
(56:23) Guilty Pleasure: Sci Fi and Thinking Beyond the Limits




