In short
Physical Intelligence builds robotic foundation models that learn end-to-end from pixels/images and language to actions, aiming for robots that can perform any task across new environments. They argue the classical robotics pipeline (perception → planning → control) is fundamentally wrong for real-world generalization, and that reinforcement learning from real experience is what enables reliable deployment.
Guests
Karol Hausman and Tobi Springenberg of Physical Intelligence. They focus on robotic foundation models, scaling data, generalization, performance, and deployment; they also publish results and open-source model artifacts.
Key claims
Robots can be trained like vision-language models plus an “action expert,” using mostly robotics data and initial human teleoperation demonstrations. Generalization is driven by diverse data; performance plateaus with more of the same data, so RL from deployed experience (PiStar 0.6) is needed to escape the plateau. Deployment reliability is the main bottleneck.
Notable examples
PiStar 0.6 improved policy throughput 2x on box building, making coffee with an espresso machine, and folding laundry; robots served coffee for 13 hours and folded laundry for 4 hours. RL handled real-world long-tail failures like stuck cardboard sheets in box building and corrected espresso “tamping” behavior after 34–50 human correction episodes.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Physical Intelligence's Mission
1:30 to 3:00
Carol and Toby explain their focus on building foundation models for robotics and the significant progress they've made.
“Carl, Toby, thank you so much for joining us here today.”
The Shift from Vertical Integration to Foundation Models
3:00 to 4:30
The guests discuss the advantages of focusing on intelligence rather than building specific robotic products.
“there are companies that are building fully vertically integrated robotic products right now.”
The Hardware and Intelligence Bottleneck
4:30 to 6:00
The discussion delves into the hardware advancements and the ongoing challenges of robot intelligence.
“would just target this problem head-on focus on the intelligence and if we can do that that would lead to many different vertical products.”
Challenges in Generalization and Performance
6:00 to 7:30
Carol and Toby outline the challenges of generalization and performance in robotic models.
“So even with simple robots, we are not yet at the level of a human operator.”
Progress Towards Deployment of Robots
7:30 to 9:00
The guests share their experiences and timelines about deploying robots in real-world scenarios.
“the lighting is different than what you've seen in the past and so on.”
The Open Sourcing and Testing of Models
9:00 to 10:30
The discussion highlights the open-source approach and its impact on testing and application of models.
“If you could limit that, where do you think generalization and performance will need to be before we can deploy these kind of robots?”
Current Technical Architecture of Foundation Models
10:30 to 12:00
Carol and Toby explain the technical architecture behind their robotic foundation models.
“So kind of similarly to how with large language models, you know, you train this model, you kind of cook it in-house, you try to make the best job possible.”
Transformers in Robotics: Architecture and Data Scaling
14:00 to 15:40
Learn about the transformer model architecture and data scaling in robotics.
“And so broadly, it's like a transformer model that is a fairly large model, up to like some billion parameters at this point, that we use, that we pre-train on our robotics data and on internet data.”
Historical Evolution of Robotics: From Rules to Learning
15:40 to 19:00
Explore the historical evolution in robotics and how problem-solving shifted from rule-based to learning-based approaches.
“But I think the foundation of like the data and how we bring it into the model will probably stay like this.”
End-to-End Learning in Robotics: Overcoming Challenges
19:00 to 23:00
Discover the concept of end-to-end learning in robotics and its associated challenges with data requirements.
“You can just add some action components on top of it and have a common world understanding and connect it to how to actually perform things in the world.”
Show all 28 chapters
Integrating Reinforcement Learning in Robotics
23:00 to 27:20
Learn how reinforcement learning can enhance robot performance and the significance of real data collection.
“are kind of baked in or because we are focused on the text problem, on problems like math and coding.”
Real-World Applications of Reinforcement Learning
27:20 to 28:00
Understand the practical applications of reinforcement learning in real-world robotics tasks.
“you're hill climbing on your reward signal.”
Generalization vs. Performance in Robotics
28:00 to 29:22
Learn how generalization and performance can coexist in robotic tasks.
“and I kind of nail it down in a way where I can solve it from many different positions and I can deal with all the long tail of failures that I will encounter, right?”
Real-World Challenges in Box Building Task
29:22 to 30:33
Discover the complexities faced by robots when building boxes in real scenarios.
“that we've actually looked at for this release where there were failure modes that we saw that if you had just done a simulation of it, you might not have seen it.”
Success of Reinforcement Learning in Real Deployment
30:33 to 32:36
Examine the success of RL in real-world tasks like box building and coffee making.
“And then our kind of method can kind of figure out that, oh, actually, what I need to do is I need to separate this.”
The Bottleneck of Reliability in Robotics
32:36 to 34:21
Understand the reliability issues that hinder robotic models from being deployed.
“So the headline figures were we increased throughput of the policies by over 2x on these three tasks.”
Innovations in Robotic Capabilities and Deployment
34:21 to 36:29
Explore how robotics can innovate both in capability and customer deployment.
“Because if they break every other trial, they're not really deployable.”
Continual Learning in Robotics: Progress and Potential
36:29 to 38:55
Investigate the potential for continual learning in robotics and its current stage.
“We haven't so far figured out how to learn from your own experience.”
Generalization and Task Transfer in Robotics
38:55 to 42:00
Delve into the relationship between task learning and generalization in robots.
“I would say we're at the very beginning of this, right?”
Improving Model Performance Through Data Feedback
42:00 to 45:00
Learn how data feedback loops enhance model training for robotics.
“That as you deploy these models, the data comes back, the models get better, you can deploy them more, then the models get better, you can deploy them more and so on.”
Bootstrapping and Deploying Models for Robotics
45:00 to 47:50
Discover the challenges and strategies in bootstrapping and deploying robotic models.
“And this gets better with more data and more tasks.”
Challenges of Generalization in Robotics
47:50 to 51:40
Explore the complexities of achieving generalization in robotic applications.
Comparative Analysis: Robotics vs. Self-Driving
51:40 to 55:20
Understand the differences and similarities between robotics and self-driving technology.
“No, you guys have a grand, grand vision.”
Surprising Advances in Video Models
55:20 to 56:04
Hear about the unexpected trajectory of video model advancements in AI.
“that maybe the problem is not that much harder and it might be actually easier.”
The Surprising Progress of General Intelligence Models
56:04 to 57:06
Explore the unexpected advancements in models that seem to exhibit general intelligence.
“And like every little advance that I see, winning IMO math challenges or applying it to finding new stuff in science.”
Challenges in Machine Learning Approaches
57:06 to 58:26
Discuss the historical challenges in machine learning and the shift towards multitask learning.
Learning from Experience: The Evolution of Intelligence
58:26 to 59:30
Examine how intelligent species learn and the implications for artificial intelligence.
“Do you think it's like an accordion where we go from one framework to the other framework?”
Parallels Between Child Development and AI Learning
59:30 to 1:00:42
Analyze the similarities between child development and the learning process in AI.
“and it's kind of interesting to you know how similarly to how we learn you would think that if there was a way to pre-bake all of the intelligence, the evolution would have figured this out.”
Transcript
Automatic transcript. May contain errors.0:00Just like the fact that this whole thing works, it's kind of mind-blowing. Yeah. Right? Like you build this like loosely brain-inspired thing that has very general purpose learning algorithm. You feed it data and it somehow gets it and gets it way better than anything we've ever had before. And this applies to robots and it applies to vision and language and sound and all kinds of other things. And like, I think if you stop for a second and just think about it, how it works and that it works, it's just like absolutely mind blowing.
0:49In this episode, we sit down with Carol and Toby of Physical Intelligence, a company building foundation models for robotics. Carol and Toby explain why the classical approach of breaking robotics down into perception, planning, and control was fundamentally wrong, and how end-to-end learning with reinforcement learning is finally making deployment possible. You'll hear how they achieved robust, real-world performance, getting robots to make coffee for 13 hours straight, and how these models generalize across radically different tasks, from surgical robots to drone flying, in ways that we don't fully understand.
1:19We also talk about the technical insights behind Pi Star 0.6, which is physical intelligence's newest model that learns from experience using reinforcement learning. Enjoy the show. Carl, Toby, thank you so much for joining us here today. Thank you for having us. Excited to talk everything physical intelligence, general robotics, etc. Maybe before we get into it, just for our audience, can you share a little bit about what physical intelligence is and the mission that you're after? Yeah. So at Physical Intelligence, we are building robotic foundation models. These are models that in principle should be able to have any robot do any task.
1:56And over the past one and a half years or so, we created the right building blocks that show how these models could scale. So we've shown that they're able to control many different robotic form factors, many different types of robots. We've also shown that they're able to generalize so you can bring it to completely new environments and what it takes for them to generalize. And this last release that we just had called PyStar06 that we also wanted to tell you more about shows how we can get them to good performance so that they're starting to become deployable. And this is really important to us because we want to see this technology actually deployed in the real world, but also because we don't have the benefit of having the free data on the internet.
2:45There is no data of robot actions. So we need to create the datasets ourselves. So we are after the problem of physical intelligence, after the problem of creating foundation models for robots. And we've made quite a lot of progress. Wonderful. And can I ask why the decision to build foundation models as opposed to, you know, there are companies that are building fully vertically integrated robotic products right now. You know, the Sunday launch last month is in the back of my head. You can buy a cute little robot helper for your household. There's companies working on cooking robots. There's obviously the humanoid companies.
3:18Why build a foundation model versus build a robot yourselves? Yeah. So I think if you look at the history of robotics, it's very, very clear to me, and I think to many roboticists, that we've been always bottlenecked on intelligence. We've had robots that are capable of doing incredible things, whether it's in the home or in industrial settings. We've seen robots more than a decade ago that if teleoperated, they can clean the entire house. and the really important caveat is if teleoperated so if there is a human mind behind it it's clear that the hardware is capable of doing lots of different things and for a very long time robotics companies have been structured the way you described where you kind of think of creating a specific robot that's designed to do just a single task or a single application and instead what we thought would be would really help the field is to focus on the bottleneck on the intelligence so we created a company to focus on that bottleneck because we think that this is if we if we address that bottleneck we can actually make robots happen and if you do it any other way you're basically not making as much progress on the bottleneck as you could be so we thought we would just target this problem head-on focus on the intelligence and if we can do that that would lead to many different vertical products.
4:38It will lead to robots being deployed in the home, in industrial settings, basically anywhere. Can I just pressure that, test that a little bit? So on the hardware side, like I've seen the latest videos, for example, of the Optimus hand. It's like, it's exquisite. It's a piece of art. And I hadn't seen the videos of people, you know, tele-operated robots cleaning houses 10 years ago. But I'm wondering if there's a set of tasks that's maybe now just on the cusp of becoming possible? For example, cooking or being able to peel and dice an onion that you couldn't have done with hardware prior to where we currently are.
5:14So how much of a why now do you think hardware is or isn't? So there's a lot of progress in hardware, especially in humanoid hardware. Like dextrous hands, for instance, as you mentioned. I think they're much better now than they were even a few years ago. but it still doesn't address the bottleneck. We could have had robots operating, you know, chopping vegetables or doing cooking even with simple grippers before. The problem is that we don't have the intelligence to operate these robots. And the more complex the hardware is, it doesn't really resolve that bottleneck, right? Like it allows you to do more potentially but you're still bottlenecked by the fundamental challenge of robots not being intelligent enough.
5:58I see. So hardware may raise the ceiling on what you're able to do, but the capability floor, we're not even there yet. That's right. So even with simple robots, we are not yet at the level of a human operator. So the limit being the intelligence layer, what's the limit to developing the intelligence? Is that collecting data? Is it doing it cheaply? Because you've broken down the problem. We're going to keep asking you why, why, why, and just drill down further. So what's the next layer of the, okay, what's the bottleneck for solving intelligence? generalization? It's a good question. So we thought about it in terms of three factors.
6:34We refer to them as capability, generalization, and performance. With capability, our idea was that we want to get to the point where as long as you can collect data for something, for a task or for a robot, you should have a model that should be able to replicate that, to automate that task. This is something that we've gotten to fairly quickly. This was our PyZero release around a year ago or so, showing that it's basically possible, that if you can collect data for any task, for any robot, you should be able to automate it, and the model should be able to learn it. The next challenge is around generalization, and this is still an open challenge.
7:13So we wanted to get to the point where the robots can just work zero shot, and you can just bring them to a new home, for instance, and they should know how to operate in that home. And this is a really, really difficult problem, right? Like if you put a robot in a new home, it needs to understand where different items are, the counters look different, the lighting is different than what you've seen in the past and so on. And I wouldn't say that this problem is solved, but I think we start to get a handle on how to solve it and how it scales. And the only answer to generalization that we know in machine learning is through diversity of data.
7:46So if you see a lot of different diverse data sets, you should be able to generalize to a setting that is similar to the one you've seen. And this is something that we've seen with our Py05 release in April of this year, that we got to the point where we can bring a robot to a new home that it's never been to before, and it's able to operate in that home. It's not perfect yet, but at least it has some kind of common sense on how to go about simple tasks like cleaning up the kitchen and things like that. And then the last challenge that is also not fully solved yet is performance. So how can we get these models to the point where the performance is good enough so you can actually deploy them?
8:23And deployments here are really, really important because, as I mentioned before, we also need to gather data. And I think that is going to be the most scalable way of collecting data because you'll have robots out there in the world doing economically valuable tasks. And that way, the cost of the data collection is basically negative. And the more broadly you can deploy this technology, the more data you'll be getting. and I think in the limit that will be the biggest source of the data you can imagine, much bigger than internet data for instance. And how far away do you think we are from generalization or from a performance level that maybe it's a controlled environment, maybe it's a general environment in homes or offices, but not the whole world?
9:04If you could limit that, where do you think generalization and performance will need to be before we can deploy these kind of robots? I think we are actually fairly close to deploying these robots. We started deploying them ourselves already. We thought this was something that was going to take something like five years to get to the point where the technology is actually ready to deploy a robot in a commercial setting and have it do something valuable. But we've done it, I think, two months ago or something like that. so I think we're now getting to that threshold that that the models are useful enough they're perform performant enough and they can do enough of variety of tasks to be actually useful so that's a really really exciting moment we I think we just crossed that threshold I think it's still to be determined how wide is the aperture of where we can deploy there are some tasks where the failure can be really catastrophic.
10:00Maybe these are not the best tests to deploy just yet. There are some tests that require a ton of generalization, like deploying in the homes, or that have privacy concerns or safety concerns and so on. Maybe these are not the best places to deploy just yet, but I think that the aperture is growing. As we collect more data, as these models get better, we can deploy them in more and more settings. So I think we're starting to get Where is the current aperture that you're deploying right now? So we're actually, this is a really difficult question to answer, because with these foundation models, sometimes you don't fully know.
10:36So kind of similarly to how with large language models, you know, you train this model, you kind of cook it in-house, you try to make the best job possible. And then at the very end, you get this artifact, and you can't really predict how good the artifact is going to be. you kind of have to test it. And that's where we are with these models as well. So for instance, we open source them so that we are not the only ones testing it and we're not the bottlenecking knowing what their capabilities are. And by open sourcing them, we see them being applied to actually many more applications that we could have imagined.
11:11Things like driving or surgical robots or agriculture and places like that. so I don't have a very good estimate of what the aperture is I think it's wider than what I had expected and I think it's also will be it will be growing over time the more data these models get the more mature they get I think the aperture will continue to grow I would add maybe like on the performance level like as you said the aperture is probably wider the starting point is wider than we thought but at the same time of course if you actually want each of those starting points that you start out for each of those applications to be at a level where people would want to use this as a day-to-day driving their businesses.
11:56That's probably still quite a bit of hill climbing to do in terms of performance. So with this release that we're going to talk about a little bit in a bit, I guess, the Pystyro 6, we've made progress on learning from experience data, getting that back and making the models better when they are deployed. It's still for a lot of things that I can naively imagine that there'd be lots of scenarios where there's a really, really long tale of things that can go wrong or that you can encounter that we don't yet have a great grasp on how to completely solve, I would say, as well. And you guys have been really great about publishing your results with a lot of transparency, releasing it open source.
12:34So whatever you're comfortable sharing, can you talk about what your overall technical architecture, so to speak, is? and do you think that the architecture to kind of get to this promised land is you know pretty much baked and it'll be variations on the theme of where we are and we just need to collect a ton of data or do you think that you know the architecture is still still being figured out i would say so we can maybe start with like a little bit discussing where we are at now and then we can like go into the details of like how that might change so at moment, the architecture is very analogous to how VLMs are built today that probably most of you interact with on a day-to-day basis, right?
13:15Type something in and put an image in and ask it to read what's on the image and so on. And we've kind of started from the same standpoint of, there's a model that's trained on internet scale data and it's ingested image data and text, and And we're adding all this robotics data. And our training actually predominantly now is on robotics data, on data that we have collected ourselves. We have a little bit of that internet data in the mix, but the majority of it is robotics data. The architecture is kind of this vision language model. And we add something on the side, which is what we call the action model, the action expert, the part of the model that actually then has to drive the robot.
13:55So that basically looks at the image and the instruction is getting and has to perform the task, has to send commands to the robot. And so broadly, it's like a transformer model that is a fairly large model, up to like some billion parameters at this point, that we use, that we pre-train on our robotics data and on internet data. And it is trained largely initially from human demonstration data. Carol mentioned this earlier a little bit, and we have this demonstration data, teleoperated data of humans trying to get the robot to do stuff. So that's the architecture that it looks like now. And roughly, the scaling that we're getting is from scaling our data.
14:36And we use models similar to what comes from the VLM world. How that might change, I think, is an open question. I think there's lots of opportunities in adding more capabilities to these models that we're also exploring. You can imagine that you might want more context in these models. You might want more cameras added to the robots that the model needs to be able to use. You might want to have a better understanding of the physical world in the sense of understanding exactly what's in the room, what can break, what is easily movable, and so on. So there's lots to be done, I think, in both capabilities and also changing the architecture around.
15:21And I wouldn't be surprised if in like five, six years, we look back and we say, oh, you know, maybe the backbone of the model that we used at the time, which currently comes from this VLM land, has changed. Maybe we've moved on and we use something slightly different. I think that will evolve over time. But I think the foundation of like the data and how we bring it into the model will probably stay like this. Got it. And should I think about it as it's pixels or signals in and then actions out? Is that like a single big neural net? It's one big model, yeah. It's really just basically images in, text in, text out and actions out at this point.
16:04And are you, I guess, do you have a separate kind of locomotion versus manipulation stack? And maybe this might be a good time to talk about kind of just the historical evolution in robotics and the various different waves of learning and how it pertains to your stack. Yeah, so for a long time, even before learning arrived here, people thought that robotics is one of these problems where you can, if you put enough people on it, enough engineers, they can think really hard about it and eventually write the code that will have the robot do anything in the world. And people have tried really, really hard to do it this way.
16:42And then it turned out that the world is just way too complex right like you can't just write every single every single case you'll encounter in the real world so that doesn't work and and also as we were you know trying to work on on that version of the problem what ended up happening is people did what they usually do they try to break down this problem into smaller sub problems so like rather than working on the full robotics problem you would say there was a perception aspect of the problem there's a control aspect of the problem there's the planning part of the problem and this is almost growing to different communities there's a planning community there's controls community there's they have their own conferences their own problems and all of that so then as we realized that you know it's not really possible to to handwrite all of these rules people thought that we should learn them we should learn them from data which is seems like a really good idea right this is how we learn too um but what ended up happening is that they started learning each one of those components these these broke down components separate learning separately yeah so you would have a perception layer that is fully learned maybe you'll have a control layer that is learned maybe you'll have a planner that is learned and that showed some progress it was better than what we had before yeah but then turned out that breaking down these problems this problem into these sub components it actually is the piece that doesn't work because you know when i try to pick up this glass i don't think about it in terms of perception and then planning and then control i just i just go for it i just pick up the glass and it's just all very natural So it turned out that this pipeline approach where you have these predefined interfaces that like perception gives you the position of the object and then the planner gives you the trajectory and the control executes it.
18:22Those interfaces are the pieces that broke down. So everything that we thought we knew how we work was always wrong. So then we then arrived to the next stage of this where we said, well, maybe just breaking down this problem was a bad idea to begin with. So let's just train the whole thing end to end. right so we'll take whatever uh uh the the the kind of the sensory inputs as input to the network and we'll have actions as the output that's what we refer to as the end-to-end approach where you try to go straight from pixels to actions and we'll we'll have the network figure out or the learning algorithm figure out how to split it into these different components if it's even possible um and then while we were doing that we figured that it actually requires a ton data to do this and often it breaks when it requires some kind of common sense and to gather that common sense through first first person action data sets is really really hard because you would need to experience every single thing in the world to do this and that's where we stumbled upon vision language action models where we can use models that were pre-trained on internet data that already have pretty good understanding of how the world works um and we can utilize that knowledge so that we don't need to experience everything firsthand.
19:37You can just add some action components on top of it and have a common world understanding and connect it to how to actually perform things in the world. And that's more or less where we're at today. Now at physical intelligence, we figured a few other things. So how do you scale? How do you start to scale these models? How do you get them to generalize? How do you get them to perform much better? How do you have them move much faster? How do you get them to the point where you can start deploying them? But I think largely we're still in this era of how do you bring some of the common sense knowledge from the internet pre-training?
20:08How do you make these models very general so that they can work on any robot and perform motions? And can I ask for something like reasoning, right? There's so much stuff happening in the reasoning side of the large language model space. Do you get the benefits of that as part of your VLA backbone? Do you have reasoning kind of emerge as a consequence of what you're doing as you train these end to end? or perhaps I think about, you know, some of the benefits of what's happening in the LLM world. Do they benefit you or not? I mean, I think definitely the models that we have today, they are already planning actions, not just at what is the immediate action, but kind of what are the next 50 things I need to do, right?
20:48So like the next 50 time steps in some sense. It's a very short horizon. 50 steps means like a second or two, right? And it also additionally kind of decomposes tasks into subtasks in language space already. So when we ask it, oh, clean the kitchen, the first subtask it might pick out to do is like, oh, I have to drive to the counter and then I have to like pick up the glass, move the glass into the sink. So it already has those aspects in some sense, right? So it decomposes tasks into subtasks because it gives itself its own subtasks. And it predicts a little bit of a horizon of how actions go.
21:27So some of it is already there, I think. I think in the future, there will probably be more of it. I do totally expect that all the advances on our training for reasoning, all these things will also make their way into robotics. And I think it's kind of interesting to think about because it's maybe a little different than the RL for math problems that people do, for example, right? Because I think those are very easy for us humans to think of as like textual problems, right? You think through them in your head in like text. Okay, if I change this formula this way, I will get this outcome and so on.
22:07I think for the physical intelligence part of it, it will probably be a bit more than that, right? It's going to be a little bit different when you try to learn a new sport, for example, when I recently started to try to learn how to play tennis. And, you know, I don't think through in my head of like, I need to now grab the racket. I need to move it here and I need to do this swing. But it's more like you think through the motion itself, right? You think about how does your body move? how maybe your plan, in some sense, trajectories of objects around you in your head. And so those things, I think, we'll see coming to the models more over time.
22:43Yeah, I suspect that over time, right now we're in a place where we benefit quite a bit from vision language models. I think it's very, very likely that that's going to reverse. That a lot of the shortcomings that we see in LLMs today are kind of baked in or because we are focused on the text problem, on problems like math and coding. And I think robotics will offer this new avenue where you need to kind of rethink how to think about reasoning. Reasoning should probably happen in some kind of abstract space where, you know, you can reason a little bit in text, you can reason a little bit in images, maybe you can reason in trajectories or in all kinds of different spaces to arrive at the answer.
23:29And robotics provides this really nice test bed where you're grounded in the physical world. There is not that much data yet, so you kind of need to deal with some of the difficulties that come with that. But I think it will provide for new findings that will then be reapplied to the LLM world. So speaking about data, give us a sense of, I don't know, how you measure the sort of magnitude of data you've already collected and how much you would like to collect in the next year. I'm sure more is better, but what is the magnitude we're talking about? Yeah, data is one of those things that is actually fairly nuanced.
24:11It's not just a matter of quantity. Quality obviously matters, but also things like diversity. and even when you think about the quality or diversity of robot data these are not very strictly defined terms right like if you if you go for the same tasks in like 10 different ways is this diverse data or not or how do you compare it to the diversity of the data if you go for like 10 different glasses right so this is something that i don't think we as a community fully understand like how to characterize the data how to describe diversity how to describe the quality of the data, how to make it very, very rigorous.
24:50And we're also finding out that there are some aspects of the data that really, really matter. Like for instance, if you want to get to a certain performance on a task, you're not going to get there by just increasing the quantity of the data you already have. We've been working on these three different tasks for the PyPy star 06 release. And we've noticed fairly early on that if we just keep on collecting more and more data the same way that we've been collecting so far, the performance plateaus. You're not going to just keep on getting better. So you need to find either new ways of collecting it or you need to start thinking about what kind of data will result in better performance.
25:29And this is where things like reinforcement learning and things like this can really, can really, really help. Let's talk reinforcement learning and let's talk Like pi star 0.6 is the star a nod to q star? Yeah. Okay. Effectively, yeah. Trying to get to like policy star actually optimal. Policy star. Okay, wonderful. Maybe just say a word on what you guys are doing with pi star 0.6, and then we can dive into what RL means for your world. Yeah, for sure. So, I mean, I think the main, if we want to contrast it to what we talked about earlier, the main difference is that up to that point, basically, all of the robotics foundational more learning that we've done was basically demonstration data, teleoperated, going into the model, the model is trained, kind of like just imitate that data, right?
26:17And now with this new model, PystarO6, what we're using is basically RL from experience that the robot collects itself by actually running a policy. So we start with the initial policy is this demonstration trained policy, and then you deploy it. You try to actually have the robot solve the task. and then it additionally gets kind of reward signals given by humans and it can also get corrections so where the human intervenes and says oh actually you know what this is not right let's let's do this a little differently and that data that process basically that data is collected gets comes back in and the model kind of uses that data to try and figure out which of the data can I kind of kind of like should I reinforce should I do more of and which of it should I do less of and and basically improve itself over time basically that's kind of the big distinction and having that stream of real data coming in is kind of the missing piece that Carol was talking about that allows us to now escape this plateau that otherwise we were finding we were kind of like getting to.
27:19Yeah. And I guess in my brain, I think of RRL as, you know, you're hill climbing on your reward signal. And so how do you make sure you're generalizing as you hill climb on these specific tasks? The way we're thinking about this for this specific kind of problem is like you have this sort of general model and it achieves some performance that isn't great. And now your first goal actually isn't to further generalize. You want to kind of solve this specific task first, right? Like, so we deploy it and we've picked like three, four tasks. So it has to generalize across tasks. Nonetheless, the method has to generalize.
27:55But when you're actually kind of deploying it and trying to start this RL process, you really care about let's make sure I nail down this task and I kind of nail it down in a way where I can solve it from many different positions and I can deal with all the long tail of failures that I will encounter, right? So in some sense, the generalization and the performance here may seem at odds when you look at it from like, oh, wait, but now you're like just doing this one task. But really at the end of the day, what we want to do is we have the same method, the same process that deploys to each of these tasks.
28:32and then kind of gets the performance high. And then we can have all of that data across all of these tasks and we can bring that data back basically, right? So in that sense, it's not actually at odds, if that makes sense. Yeah, that makes sense. How much of the RL are you doing? It sounds like this is a real life RL. Can you talk a little bit about the approach to how much RL you're doing in SIEM versus in real life? So we have taken a quite like real world first approach as opposed to using SIEM. We are exploring SIEM, of course, as well as a research tool. But all the URL we've done for the PyStar06 paper is actually on real systems in the real world.
Read the full transcript
29:13And the reason for that is that it's actually really, really hard to model. Again, we can get back to the long tail of failures that you see when you do deployment. I can give you a lot of examples from the tasks that we've actually looked at for this release where there were failure modes that we saw that if you had just done a simulation of it, you might not have seen it. So to give you an example, we have this one task, which is you have to build a box. So this is an actual deployment task where the goal is we build these little cardboard boxes to put chocolate into such that they can then be packaged up and sent out, basically.
29:51So that's building a chocolate box, basically. And building this box initially worked great, And then there is new shipments of boxes coming in and they come in as like a flattened sheet of cardboard. And then these cardboards that came in in this new shipment were kind of not perfectly perforated. So they were sticking together. Right. And then the robot starts like grabbing them, puts them on a table to try to build this box. And it has two boxes suddenly on the table. Right. And this is something that wouldn't happen in sim if you had written like a nice simulator where you would just get individual cardboards and like fold them.
30:25And so now you have to deal with this problem. And if you just learn everything in SIEM and then try to deploy it, you wouldn't encounter it. So we encounter it. And then our kind of method can kind of figure out that, oh, actually, what I need to do is I need to separate this. And I need to move that second piece back and build the box. And we see a lot of successes for RL being applied in SIEM and transferred to the real world, especially in locomotion. and we we haven't really seen that kind of success in manipulation for for these kind of methods and i think maybe one reason for that is that with with locomotion with trying to move around it seems that the the biggest part of the problem is modeling your own body so if you can figure out how to model you yourself as a robot you're basically like almost there so uh you can do this modeling simulation exercise once because you only can do it you only have to do it for you yourself for this one robot and then you're basically done if you do it really really well it should transfer with manipulation however the problem is not how you move your own body it's how the world reacts to it you're actually changing the world around you it's not difficult to figure out how to move your hand from a to b it's difficult to figure out how this affects the objects you're interacting with and now the problem is no longer just modeling your own robot you have to model the entire world right like every single object that you might be interacting with every single task you can think of and that's where we see scaling problems and that's i think why we haven't seen those kind of methods be as effective in in manipulation what was the headline of the results from pi star 0.6 and you know where where did you see the model get after rl on the on the test that you cared about and what do you think that means about your overall training recipe going forward yeah so i think for me the most impressive thing honestly for me personally to see was just have these models run for hours at a time, recover from lots of different failures and basically just keep going.
32:28And at the same time, do that at a rate that is actually much better than the initial model that we started with. So the headline figures were we increased throughput of the policies by over 2x on these three tasks. So there's one task was this box building task I already talked about. One was making coffee with an actual kind of industrial scale espresso machine. And the other one was kind of like folding laundry. And so for each of them, we managed to like make the base policy that was trained just from demonstrations much, much faster and also make it be able to recover from theaters much, much better.
33:05And so seeing that actually in action when you just you sit there, right, we have if you go to our website, you can look at the videos. we have the robot serve coffee for 13 hours in a row or fold laundry for four hours, things like that. Actually seeing that live changes the way you think about these models. You know, changes the way at least I think about it actually being realistic that we can deploy it and that we can do it in a way where it's not just a toy demo which is shown once but is actually kind of doing the real thing fully. and that's been really a challenge in robotics that i don't think many people are aware of yeah like you know you see so many videos of robots doing cool things and you know we post these videos too they're like there's basically like anything you want robot to do there's probably already a video of a robot doing that um but you know you can take as many takes as you will as you want you can you can keep on recording until you get the perfect shot and the problem that i think everybody encounters is the reliability of these models how performant they are, how fast they can go about the task, for how long you can actually deploy them without failure.
34:17And I think this is the biggest bottleneck in terms of deploying these models in the real world. Because if they break every other trial, they're not really deployable. Right. And this is, I think, the most important breakthrough for us with this PyStar06 release, that we can actually start getting to a place where they are deployable, where we use these robots in our office to serve us coffee or we can give them to people at Pi to fold laundry in their home or we can deploy them and have them fold boxes for real and that is really, really exciting. Should we think about what you guys are doing with reinforcement learning as primarily a customer deployment reliability point then?
35:01Like you can now make sure that you can go reliably deploy the coffee making model on a customer site and it's going to be fast enough, It's not going to fail over long time horizons. So it's more of a customer deployment innovations versus like a fundamental kind of capability innovation. Or is it both? I think it's both. I think, I mean, Carol, you said this a little bit earlier. I think to some extent, the robots that we really, really want, right? The robot that you want at home, which can do your laundry, do your dishes, cook for you, drive around. and also the robot that people want in these smaller businesses, maybe solving a real problem that they have that they don't want to automate in a classical way because it's too expensive, like building a chocolate box.
35:44Those are things where the robot has to be reliable, it has to be good, and it has to have the capability to do a new task that it hasn't seen in initial training stages. I think it's unrealistic for us to assume that, you know, we can just go with like more and more human data collection, go bigger, bigger, bigger. We will do that. But there is always going to be a limit to how good and how much data you can get and how good the initial policy is going to be. So I think it is that what you said in terms of we want deployment, we need this. But also, I think increasingly over the next years, I expect we will see that we will do this deployment and that data will actually become really valuable as a source for pre-training, for making our models better themselves.
36:29And we rely more and more on autonomous data collection, is my prediction at least, over the next coming years to kind of build that host of data, that convex hull of all the tasks that we want robots eventually to do such that the models ingest this and becomes good at doing them and interpolating between them. And I think of it as a new capability. We haven't so far figured out how to learn from your own experience. There's been many attempts, but I don't think we've seen it done at scale. to the extent that actually shows a convincing result that allows you to deploy something. And this is why this result was really, really important to us.
37:10We wanted to get to the point where they can learn from their own experience. Because similarly to how we learn, you can learn a little bit from watching videos and maybe learning from others, but at some point you need to learn on the job. You need to try the thing yourself. You need to see how your actions impact what you actually want to achieve and make your own conclusions and try to learn that way. Yeah. And I think this is the first step towards that. You're reminding me of the, do you guys read the Rich Sutton Age of Experience paper this year? I love, I thought it was very profound. Do you think that this unlocks kind of continual learning in robotics real?
37:48Will this be part of that? It kind of depends what people mean by continual learning i think it's um it's definitely more continual than what we've done in the past where you know you have like a big pre-training mixture and maybe like a post-training mixture and you like you know you you sit down you work really really hard and then you come up with an artifact and like that's it yeah right like the artifact is done and there's not much you can do to change it now this is a much more of a living thing right like we we start with a process similar to this but then you deploy it and then it keeps on learning right so it's much more continual in that sense that it tries new things it tries to learn from its own experience and it keeps on getting better yeah now i think there is still room to for it to be much more continual where it can acquire new skills that way or it can be even much faster in doing this yeah um it can probably reason throughout this process so i i think there's a spectrum of like how much you can learn on the job.
38:51And this is really promising because it shows that you can do it, but I think we can make it much, much better. Yeah, I would agree. I would say we're at the very beginning of this, right? And it's definitely not continued learning in the classical sense that people would have thought about it of like data streams and then the whole thing churns and it just ultimately leads all the way to, I don't know, AGI or something like this yet. But it's a first step, I would say. We're moving in the right direction there, and there's lots more to be done. And I think I will say, even from this release, I was personally impressed and to some extent shocked how good these models actually are at picking up little things that you put back into the data.
39:34I was surprised that even with just human corrections, there was one example for tamping when we do. So tamping is a specific part of making an espresso, right? You put the beans and you have to tamp down the best part. Yeah, the best part. You have to tamp down the coffee before you put it in. I don't get it right myself. There you go. See, I'm not a coffee expert. It's a skill issue. I'm going to get it just right. That's right. And so our robot in the beginning tamped way too hard because it just happened to be the case that the initial human demonstrations were just making sure that, you know, let's make sure the coffee grounds are flat so we can put it in.
40:11and then the robot was like tamping really hard and like almost lifting itself off the table. When we looked at it, we're like, that's a bit much. And so with just, I don't know, it was I think 34 to 50 episodes. There's a really small range of corrections that humans did and we feed that data back and the model actually starts like being much more gentle and doing the correct thing. And I was really surprised by that because you think, you know, this model has been pre-trained on these millions and millions of episodes and now you're just doing a little correction and that actually works. So seeing that happen was a thing that I think is pointing towards this continual learning part, which I find impressive.
40:46Can I ask though, and the thing I'm still hung up on is generalization. So as I learn how to tamp better, does that make me better at folding boxes or not? In this specific case, no. But the mechanism is the same that you can also employ to fix the, oh, I have two boxes in front of me that are sort of stuck together and I need to pull them apart, right? Because you can get 30 corrections for the stamping part. You get 30 corrections for the pulling boxes apart. Yeah. But you get 30 corrections for, oh, you know, this box wasn't like neatly folded together. And all of this accumulates together to then give you this more generalized improvement, I would say.
41:28Okay. So it's a repeatable recipe, but they don't necessarily cross-pollinate. Yeah. I mean, I would expect that as we scale this up, we might see also things actually kind of transfer from A to B if there is motions that are kind of similar across tasks. But at this point, yeah, I would say it's more like a repeatable recipe. And we see a lot of generalization from pre-training where you train on more and more tasks, more and more data. You see that it's much easier to onboard a new task or you see tasks that appear zero shot that you didn't expect before. And this keeps on improving. doing we uh we kick off a pre-training run at certain cadence and every single time we start seeing that the model keeps on getting better because there's more data being fed in there's more improvements that we're making to the pre-training process and so on and I also suspect that as we have more and more of these models deployed doing all kinds of different tasks they also bring data back in and I think one way where we where I'm quite certain where will see more generalization from that process.
42:29That as you deploy these models, the data comes back, the models get better, you can deploy them more, then the models get better, you can deploy them more and so on. And I think maybe it's worthwhile for this point that you brought up, we haven't really talked about one crucial detail aspect of this 5 star 06 recipe, which is that the model has kind of two parts. One is the policy that is trying to improve via corrections and RL feedback. And the other part is how do you actually get this RL feedback, right? So we've talked a little bit. I've mentioned like, you know, humans might correct and that's the human correction part.
43:04And the RL feedback part is a little different and it's kind of, and already has some of these aspects of generalization that I think you're like trying to search for, which is that the way we do this is we first basically get humans to tell us basically whether a specific attempt of making the coffee or doing the box was, successful or not. So there will be human labels provided with these episodes. And then we train something which is called a value function to try and predict basically from my given point of where I am in the task, will I likely be succeeding or failing basically. And this value function is then used as kind of a baseline to decide whether for this data point should I bump that up or should I bump that down depending on whether I expect that I will be moving towards success or I'm more likely to move towards failure.
44:00And one thing that we saw when we trained these value functions, so those are trained basically from the same kind of backbone, the same kind of model, but they're pre-trained before the actual policy is trained that actually runs the task. When we train these value functions, we see that adding more data from different tasks actually helps there. And the model starts being actually really quite good at at least for for certain tasks and knowing when it will fail beforehand and before it is obvious for me for example when i look at a video of it trying to insert uh the um uh the portafilter thank you see i'm i'm not good at making and it's trying to insert a portafilter into the the coffee machine it kind of knows that it doesn't quite have the right angle before that happens so like 30, 40 steps before that actually happens, the value function kind of, if you look at the prediction, drops and saying, oh, this is not good in this specific episode.
44:57So I shouldn't include this data. Interesting. And this gets better with more data and more tasks. So this is an interesting counterpoint to the Karpathy slurping bits from a straw thing, right? Because you're not waiting for that final bits at the end. You're actually getting a lot of signal along the way. I think a rel is just like such a vast field and there's so many different approaches to it. And people often associate RL with something like a policy gradient method or, you know, very specific on policy learning approaches. And to me, RL is more of a problem definition. And there is many, many approaches that get around the problem that you're referring to, which is that, you know, you only get the reward at the very, very end.
45:40And it's not really scalable for very long horizon tasks. There are things like value functions. There are things like temporal difference learning that try to get around this problem, where you constantly make predictions and you do it in a sequential way. And this is maybe another one of these things where I think robotics can really help the broader AI community. Because we don't have the advantage of having a perfect language simulator where you can run as many simulations as you would like. Instead, you need to do it in the real world. So you need to make more efficient methods. And therefore, you need to learn value functions.
46:15and things like this. And I think this will, these will be really valuable everywhere. Yeah. Can I push a little bit on, I'd love to understand, you know, internet video seems like it's part of the recipe, but not a huge focus right now. As I see it, like, do you think that there's gold left to be mined in internet video? And then if you look at what's happening in video models right now, world models, to what extent do you think that's going to be a, you know, discontinuous jump in model capabilities and an important part of your model pipeline. Yeah. I think maybe there are two questions there.
46:51One is about the data. How do you bootstrap yourself to the point where you can start deploying? And the other question is, what about video models and kind of the world model aspects of it? So on the data point, I think we are now in this bootstrap phase where basically anything goes. like whatever you you can figure out how to like add to the model and to to to its benefit i think it's good whether you can add sim whether you can add human videos some kind of handheld devices human teleoperations i think it kind of doesn't matter you just need to figure out some way to bootstrap yourself to the point where you can deploy these models because i think in the long term there's going to be this bootstrap phase but then there's going to going to be the deployment phase and i think deployment phase will be will provide much much more data than anything you could do in the bootstrap phase so we're in this kind of like weird spot right now where we tried many different things straight to see what what sticks to just get us to the deployment threshold i see and once you can deploy i think that will vastly be be much greater than than anything you can do uh before that um so so that's also what we are sprinting towards that's why we want to start deploying these models that's why we want to do this you know with many different tasks in in many different environments so that we can just have this very powerful data engine now on the on the world modeling side of things i think the world models and aurel approaches are kind of targeting the same problem the problem of counterfactuals of how do you or a credit assignment problem right like how do you figure out which actions were the ones that actually matter for your success and how would the world have evolved had you had had you taken a different action and one way you can do this is by predicting what would have happened right like rolling out a full video of you know if i if i put this portafilter a little bit differently you know where would i end up and would this be a failure or a success or you can do this through reinforcement learning and it does it through a slightly different mechanism a little bit more implicitly but it fundamentally targets a very similar problem uh we are exploring all of those approaches and try to see you know how to how to really solve the counterfactual problem um i don't think there was a an answer yet uh but we see we see a lot of progress with with reinforcement learning that we that we've just shown with with pi star with pi star of six but i think there is probably room for for many many other approaches too awesome can we talk about once you guys get past that bootstrap phase let's talk about customer deployments a little bit what do you bring to a customer what do you sell them and then um how do you imagine that's going to evolve over time like are you selling them a fully vertically integrated robotic solution are you selling them a model that they have to figure out how to integrate into their operations like how does this all work uh the the the real answer is we don't know yet yeah um we are we are still figuring that out yeah we are still quite early in in the technology as you can as you can tell we are just starting to to even get to the threshold where we can start deploying these things so we believe we should focus on the on the technology first to figure out how to get it to the point where it's actually easy to deploy and expand this aperture that we're talking about initially and robotics the history of robotic startups is as very often gets to this point where you develop a technology for for some period of time you started with this grand vision of what it should be able to enable how general purpose it will be and as soon as you pick an application that you want to apply it to you're kind of stuck you start cutting corners you start uh figuring out very special purpose solutions just for this application and very quickly you become you know an application company that just focuses on let's say warehouse pick and place robots and that's it and we really want to avoid that future we think we have a chance to really solve physical intelligence and the the benefits of doing this will far outweigh any single applications that we can focus on now so we want to make sure that the technology is as general as possible as easily deployable as possible.
51:07This aperture is as wide as possible. And then we'll start figuring out how to commercialize it. And as you said, there could be many different ways of doing this. There's probably ways that we can't think of just yet because they will depend on how the technology goes, whether you can be a model provider, a fully vertical solution, or you sell robots or whatever else. But I think it's a little too premature to answer this question. It will give you a lot of comfort, you know, just to like pick one of those. It will give Alfred a lot of comfort. Alfred will be happy with us, but I think it's just too early.
51:41No, you guys have a grand, grand vision. So thank you for working on physical intelligence. It's a wonderful, wonderful improvement. Just for pi star 06 is just a huge sort of breakthrough. And so congratulations on all the success you've had. Thank you. Can I follow up with a spicy question? Sure. So as you said, this vision is so grand, so broad. You're doing all these different things. I'm sure you've studied all previous robotics efforts, and they've largely, as you said, applied to an application, and they get narrower and narrower. And one of the most successful cases of a large application is self-driving.
52:24And Waymo or Tesla have done enormously well. But if I had to go back in history, I learned about self-driving when Sebastian Theron was on the stage of TED in, I think, 2009, 2010. And he talked about the thing where they won the DARPA Challenge. That was 2007. And we're in 2025, and the thing barely goes from San Francisco down here. They kind of can do it now, but they take local roads. They can't even get on the freeway. If you do such a generalized job, how long is the runway or the timeline that you're thinking about to build for generalization and performance? Yeah. So there are some aspects of the problem that make it easier than self-driving and some that make it harder.
53:15Yeah. One thing that makes it easier is that we don't need to deploy it only when it's 100 % reliable, right? There's many, many tasks out there that even if you're at 95 % reliability, you're totally fine. If you have a robot in your home folding your laundry and every 100 bite them, you know, it doesn't fold it perfectly, you'll be totally fine. You just call your child to go fold the 100 bite. That's right. It's an additional benefit. We still need chores. We still have that, yeah. Exactly. And with self-driving, that's not the case, right? Like if you fail every 100th time catastrophically, that's a big problem.
53:51Yeah. So I think in terms of deploying this technology, it might be easier. Now, we also benefit from the fact that this is a different era of technology. We are at the era of vision language models or foundation models that have some common sense. And we learn a lot of lessons between, what was it, 2009 and 2025. And we can benefit from all of those. So I think that also really, really helps. and these are much more general purpose solutions than what we had in the past. At the same time, there are some things that will be very challenging, right? Like there isn't just a single application. This is a very general purpose solution that can be applied to driving, but also to manipulation and locomotion and flying and all kinds of other things.
54:37And I think it's to be seen how much harder this is. So far, based on what we've experienced, it doesn't seem to be that much harder, to be honest. It seems that if you tackle this with a very general purpose kind of mindset from the get-go, it turns out that it can generalize fairly well. And there is something about physical intelligence that we don't fully understand that allows these models to generalize between driving and making coffee and flying a drone and operating a surgical robot. even though they seem so far apart from each other and it seems that these should be all different models and different applications, these models somehow can make sense out of all of that data.
55:18And that gives me a lot of hope that maybe the problem is not that much harder and it might be actually easier. So I think it's a fair question, but I also don't want to draw the wrong conclusions from what we've seen from self-driving. That's beautiful. Congratulations. What result has impressed you the most outside of results that you created? That's a great question. yeah that's a good question actually i can start i've been i've been really impressed by the video models what you mentioned earlier i saw them a few years ago i worked on on aspects of them a few years ago and i didn't expect this trajectory to be that the improvement to be so steep like they're basically indistinguishable right now from reality and they can do incredible things um so that's been really, really impressive and really surprising to me.
56:06Yeah. I would say I'm still in awe to some extent that we've gotten to this place where we do seem to get models that do seem generally intelligent to a level that I really didn't foresee coming out of just next token prediction. I'm still amazed with this. And like every little advance that I see, winning IMO math challenges or applying it to finding new stuff in science. To me, yeah, there are so many things this year where I thought like, wow, there's still a lot of progress to be made, even though it felt like at the beginning of the year, maybe this whole pre-training business of LLMs is kind of maybe petering out a bit.
56:48Yeah, realizing that there's like this whole almost second breath of fresh air basically coming in. Yeah, I would maybe add to this, just like the fact that this whole thing works. it's kind of mind-blowing yeah i don't think we like fully realize how ridiculous this is right like you you build this like loosely brain inspired thing that has very general purpose learning algorithm you feed it data and it somehow gets it and gets it way better than anything we've ever had before and this applies to robots and it applies to vision and language and sound and all kinds of other things and like i think if you stop for a second and just think about it how it works and and that it works it's just like absolutely mind-blowing like the fact that we can have robots you can put it in a home and it kind of knows what to do in a home that it's never been to before or it can make coffee for 13 hours straight or you know things like that and this is from this very general purpose thing that that trains fully end-to-end that we don't fully understand but it seems to start to get it that to me is just mind-blowing we're in a simulation it's what that's what sonya believes that we're living in a simulation but it is interesting right like in science they teach you to take a big problem and break it up into smaller and smaller problems and then basically somebody realizes that that's maybe not the best way to train machines or robots of any kind and to be honest the whole machine learning like ai field made that same mistake actually to some extent right we were working for a long time people were working on solving individual problems very deeply basically right and then over time there is this like notion of oh if we can put it all together like do multitask learning if we could do that really really well we'd do much better and then but then the fact that that all happened just because we switched to this you know general pre-training objective and then it just all falls out that's the part that is the surprising bit, right?
58:48Do you think it's like an accordion where we go from one framework to the other framework? We take big problems, break them up into small and small ones that work for a period of time, then it stopped working. And then we're like, all right, let's go back to the big problem and try to solve it more generally and go back and forth. I don't see us going back. Yeah, I don't see us going back. I think there is a lot of approaches or a lot of people saying that, you know, you need the best of both worlds and you need some kind of way of incorporating the rules that you already know about like you know newtonian physics you don't need to learn that we already know how it works so can you just like put it somehow into the weights but uh from what we've seen so far it doesn't work if you try to do this you kind of limit the ability to to learn new things and i don't think there's the best of both worlds i think we just go all the way learning and it's kind of interesting to you know how similarly to how we learn you would think that if there was a way to pre-bake all of the intelligence, the evolution would have figured this out.
59:44You would have just been born, you know, knowing everything there is to know. And we see this with some other species, right? Like I think deer, when they get born, they're basically like as smart as they will ever be. Like they don't really learn much throughout their lifetime. But for intelligent species like humans, but also I think crows, for instance, they have these childhood periods, adolescence period, where they're not very smart to begin with, but they have to learn from their own experience. And it doesn't come pre-baked. You kind of have to earn it on your own. And I think there is something to that.
1:00:21You need to just experience the world and learn from that. And I think that's the lesson we're learning in machine learning as well, in AI, that we think we know how we think, but we actually don't. and we just need to let the algorithm learn it from data. Same thing with raising a child. I think I know how my son is thinking, but I don't. Yeah, I have a small daughter and yeah, it's just so surprising. They learn so fast. They learn so fast and you don't know where they get it from. Hopefully from the parents. Hopefully. She definitely knows some things that I didn't teach her. Thank you guys so much.
1:01:00It's a really beautiful mission you're building after. Thank you for coming to share. Thank you. Thanks for having us. Thank you. Thanks for having us.
1:01:32Thank you.
From the publisher
Physical Intelligence’s Karol Hausman and Tobi Springenberg believe that robotics has been held back not by hardware limitations, but by an intelligence bottleneck that foundation models can solve. Their end-to-end learning approach combines vision, language, and action into models like π0 and π*0.6, enabling robots to learn generalizable behaviors rather than task-specific programs. The team prioritizes real-world deployment and uses RL from experience to push beyond what imitation learning alone can achieve. Their philosophy—that a single general-purpose model can handle diverse physical tasks across different robot embodiments—represents a fundamental shift in how we think about building intelligent machines for the physical world.
Hosted by Alfred Lin and Sonya Huang, Sequoia Capital




