In short
Training general-purpose robots to act in unfamiliar real-world environments (the “embodiment gap” and “world objects” problem), using scalable data and meta-learning-style generalization so robots don’t need task-by-task reprogramming.
Guest
Chelsea Finn, Stanford researcher who defined modern meta-learning via the MAML paper (tens of thousands of citations). Co-founded Physical Intelligence (2024). Her work focuses on reusing skills across tasks and adapting efficiently in new settings.
Key claims
Robotics is harder than it looks because robots must map high-dimensional sensor data to motor trajectories and handle huge variability. Scale is necessary but not sufficient; data must match real test distributions. Generalization is best judged by reliability and evidence of unseen environments, not just impressive demos. Pi models are already in production and aim to be generalist “foundation models” for robots.
Notable examples
Months of 0% laundry-folding success before an architectural/data insight; a robot opening the correct side of an ambiguous fridge; “drawer vs oven handle” overgeneralization; pinwheel assembly showing left/right “equivariance” after a mistake; action-head architecture changes raising language-following performance (reported from ~20% to ~80%).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Challenge of Laundry Folding
1:56 to 4:02
Discussion about the complexities of teaching robots to fold laundry.
“You've said that folding laundry is the most impressive thing you've seen a robot do.”
Defining Physical Intelligence
4:02 to 6:34
Exploration of physical intelligence and its importance in robotics.
“they're never going to be able to do it.”
Data Strategies for AI Robotics
6:34 to 8:20
Analyzing the data requirements for effective robotic training.
“we have all these devices around us that like ranging from dishwashers to laundry machines, to Roombas, to eventually like robot arms, cars, and so forth.”
Learning from Low-Quality Data
8:20 to 11:12
Examining how low-quality data can enhance robot learning.
“And if we can scale large amounts of data of real robots in the real world doing real jobs, then that's going to be the data that will fuel like large foundation models of robots doing tasks effectively.”
Meta-Learning and Task Adaptation
11:12 to 13:14
Discussing the role of meta-learning in robotics and task adaptation.
“the better your model is going to be able to generalize and handle open-world conditions.”
Real-World Adaptation in Robotics
13:14 to 14:00
Insight into how robots adapt to new environments and tasks.
“Now, going back to meta-learning, the idea behind meta-learning isn't just to be able to generalize to new tasks.”
Adapting Robots to New Environments
14:00 to 14:48
Learn how robots can adapt to unseen environments using neural networks.
“side, instead moved over to the left side and successfully opened it from the left side.”
The Mechanics of Robot Learning
14:49 to 17:44
Explore how robots interpret environments and execute tasks based on learned models.
“So a robot walks in a warehouse or kitchen.”
Overgeneralization vs. Adaptation
17:45 to 20:26
Understand the balance between overgeneralizing concepts and learning to adapt in robotic behavior.
“So we actually hadn't collected any data with ovens before because we weren't yet at the point where we're ready to start cooking things in the oven.”
The Importance of Data in Robot Training
20:27 to 22:56
Discover how training data influences a robot's ability to follow instructions accurately.
“So you went from a 20 % language fall-on rate to 80 % by changing how the action head attaches to the vision language black bone.”
Show all 27 chapters
Reinforcement Learning and Scalability
22:57 to 25:20
Learn how reinforcement learning provides scalable data sources for robot training.
“And that we don't use kind of the gradients of that signal to train the backbone of the model.”
Interpreting Robot Demonstrations
25:21 to 28:00
Find out what to consider when watching robot demos to gauge their effectiveness.
“I love the idea that it's like robots, they're just like us.”
Challenges in Robotics and Generalization Capabilities
28:00 to 29:19
Explore the difficulties in deploying robots and the surprising ease of generalizing robot models across different platforms.
“which really show like over time what it's doing, but it's really hard.”
Success in Generalizing Across Robot Platforms
29:20 to 31:18
Learn about the successful training of models on various robot embodiments and their unexpected performance.
“I think the hardest bit is that I do think that, I mean, robotics is hard and actually deploying robots into the real world is also really hard.”
Expressive Neural Networks and Their Adaptability
31:19 to 33:17
Discover how expressive neural networks can handle diverse robot designs with minimal adjustments.
“So, like, what does that actually prove?”
Humanoid Robots: Expectations and Challenges
33:18 to 36:10
Discuss the implications of humanoid robots in society, their reliability issues, and the potential for simpler systems.
“People, we don't just control our own body, but we can control cars.”
The Evolution of Robot Reliability and Costs
36:11 to 37:58
Analyze how the cost and reliability of robots have changed and the future outlook for improvements.
“So Professor Maya Maderek, who we had on Possible, and she was talking a lot about robots in the physical form.”
Software vs. Hardware: Safety and Deployment
37:59 to 40:09
Explore the differences between software and hardware development, focusing on safety and deployment challenges.
“And so if we really want to solve this intelligence problem, we think that we can move fastest by starting with simple systems because simple systems are already incredibly capable.”
Ensuring Safety in Robotic Development
40:10 to 42:01
Learn about the principles of robotic safety systems and the use of weak robots for safe operation.
“And so they were developed with the software technology in mind.”
Safety Considerations in Robotics
42:01 to 44:02
Learn about the safety measures in the design of weak and strong robots.
“There's a lot of things that dives are very useful for, but we found ways to ensure that the people around the robot are safe and so forth.”
Impact of AI on Labor
44:03 to 46:02
Discuss the effects of AI and robotics on various job sectors, especially blue-collar workers.
“So I think it's true for both technologies.”
Human-Robot Collaboration
46:03 to 47:23
Explore examples of how robots can augment human tasks in domestic and professional settings.
“I also think that you can create, you could imagine robots playing games with them, for example, that you wouldn't otherwise be playing or other forms of recreation.”
Ethics and Liability in Robotics
47:24 to 49:44
Examine the moral implications and legal responsibilities associated with autonomous robots.
“I'm not sure the legal frameworks that we have are even asking the right questions because it's obviously deployer, constructor, environment, etc.”
Future of Home Robotics
49:45 to 51:09
Speculate on the various roles and types of robots that may assist in homes and workplaces.
“Uh, and the, um, yeah, haven't yet been kind of faced with a situation where we would really, um, want to strongly consider any of them.”
Personalization in Robotics
51:10 to 53:59
Discuss the importance of customization and adaptability in personal robots across different life stages.
“And it might be something that, you know, for example, parents might even go, oh, we have a kid and here's the chatbot for the kid as a way of kind of guide and help.”
Optimism for the Future of Robotics
54:00 to 56:01
Reflect on the potential advancements robotics could bring to society in the coming years.
“So, Rapid Fire, is there a movie, song, or book that fills you with optimism for the future?”
The Future of Robotics and Human Productivity
56:01 to 57:31
Explore how robotics can alleviate human labor and enhance productivity.
“or were inspiring to me in various ways to, I don't know, like competitive mountain climbers that are rock climbers that scale like buildings and rocks that are like, I would have no ability to do myself.”
Transcript
Automatic transcript. May contain errors.0:00Chelsea Finn:I think the biggest risk is that everyone fails, that robotics as a whole fails, because robotics is so hard. There's so many pieces you have to put in place for anything to work. I think that when watching a robot demo, if there are details about how it was done, it's really important to read those and actually understand how that was developed. Was it developed in a way that is going to, in the long run, stand the test of time and be scalable? Pi models are already running in production. It gives me optimism that we are at the point where this technology is mature enough to be useful. Most of the AI revolution has happened behind glass.
0:34In search boxes, chat windows, image generators. The world we actually live in is made of objects that slip, doors that stick, and rooms no model has seen.
0:43Aria Finger:That's why robotics remains one of the deepest tests of intelligence. It's one thing to describe a warehouse. It's another to walk into an unfamiliar building, read a new label, pack a box it's never seen, and recover when something goes wrong without being hand-programmed for any of it. Chelsea Finn has worked on that problem from both sides of the frontier. At Stanford, her research on meta-learning helped define one of the central questions in modern AI. Her MAML paper has been cited more than tens of thousands of times. At Physical Intelligence, the company she co-founded in 2024, that question has gotten very literal.
1:22Aria Finger:Can a single model generalize broadly enough that robots don't need to be reprogrammed for every task, every warehouse, or every failure mode? The factual version of that story includes months of 0 % success rates on laundry folding before a single architectural insight unlocked the capability. Pi has since demonstrated a sequence of increasingly capable generalist policies. Today's conversation is about what it will take for AI to leave the screen and what that transition reveals about intelligence itself. Chelsea Finn, welcome to Possible. Welcome. You've said that folding laundry is the most impressive thing you've seen a robot do.
2:04I want to start not with the result, but with the months of 0 % success rates before it happened. What were those failures actually teaching you?
2:15Chelsea Finn:Yeah. So I think the first thing I'll say is that robotics is really, really hard. And it's really easy to underestimate how hard it is because we are so good at manipulating all sorts of things around us. With our hands, it comes second nature. We don't even think about how we go about flattening a shirt and folding it when we're folding laundry. So it's really easy to take for granted the fact that it's not too hard for us to manipulate things. But actually, for a robot, you need to translate all of the sensor readings, all of the different RGB pixel values into a vector of numbers, a large vector of numbers for all the different joints over time for the robot to do.
2:54Chelsea Finn:And the thing that I think specifically I found about laundry is that there are so many different ways for even just a single shirt or a single set of shirts to be crumpled and configured. and dealing with that variability is very challenging because the robot needs to understand how to kind of translate all of these different configurations of a shirt into actions that will actually make progress on the task. And so one of the things we had found previously is we were able to train robots to do tasks in narrow situations and once it broadened out to be, even for a single shirt, but broadened to be a much wider range of configurations, the problem gets a lot harder.
3:36Chelsea Finn:And the, whatever you, we started with something simple and we had some results where if you start with a shirt flat, it's able to fold it. And usually in research, it's good to start with something that works, then make it incrementally harder. In this case, it was just a scenario where we went from a flat shirt to a crumpled shirt, made it way, way harder. And that's where you do a little bit of being your head against the wall for a few months before you actually start to see signs of life.
4:01Aria Finger:I mean, honestly, watching your videos, I was like, no, what if the shirt's inside out? they're never going to be able to do it. And it's like, what are these things? It's so funny because you're like, wait, so a car can be self-driving and drive down the highway at 60 miles an hour, but the robot can't fold the shirt. Like it's just such an interesting disconnect. And so you're taking on the physical world. And like you said, like the physical world is so hard. What made you first think like, okay, physical intelligence is like a whole nother discipline. And it's like where you want it to be.
4:32Aria Finger:You want it to be in the real world with physical intelligence and maybe what's your definition of physical intelligence? Yeah.
4:37Chelsea Finn:So when I started working on robotics, it was maybe 11 years ago, 12 years ago at this point. And I was really fascinated by this problem of AI. How do we develop computers that are intelligent, that show the intelligence that people have? It's just like still, I think, so fascinating. And also it's like so much potential to have an impact on the world and is having an impact on the world today. And I felt like, at least in terms of how the field was organized at the time, that different areas of AI were often solving like a small subset of the problem. Like computer vision was really focused on object recognition or object classification or detecting text in the world or various things like that.
5:20Chelsea Finn:I did a project on text detection with ultimately the goal of helping visually impaired people. Like when I was an undergrad, for example, working on computer vision. And then even in natural language processing, there was like semantic parsing and so forth. And I was frustrated or like dissatisfied by the fact that these fields were organized around something that wasn't the end problem. It was only a part of the problem. It was like a means to an end, not the actual full thing. And I really wanted to work on something that was encapsulating the entirety of a problem. Because I think that when you actually look at the entirety of the problem, you solve it in a different way than if you were to try to break it up into a sub-problem.
5:55Chelsea Finn:And so that was one of the things I found really appealing about robotics. I also find that the fact that the real world, like it's so, there's so many challenges with it when you actually are forced with not just building a brain that processes images, but actually building a brain that uses that to kind of translate that into actions and actually have a physical impact on the world. So those were some of the things that really drew me to robotics in particular. And then in terms of like how to actually think about this notion of physical intelligence, I think it's really the ability to control any physically actuated device in any way that is that device is physically capable of doing.
6:33Chelsea Finn:And I think that right now, we have all these devices around us that like ranging from dishwashers to laundry machines, to Roombas, to eventually like robot arms, cars, and so forth. And we actually don't, We actually have so much ability to design all these different physical devices that have their actuated, have motors and so forth. But the bottleneck is always actually making those smart, making them know how to accomplish some goal. And so I think that if we have physical intelligence, we'll be able to kind of breathe intelligence into all of those different devices and any basically physical mechanism that we could imagine.
7:13One of the things you've been precise about is scale is necessary but not sufficient. Industrial data has diversity. YouTube has an embodiment gap. Simulation is not real. And that's a pretty thorough indictment of the obvious sources. So what does the right data strategy look like for this project?
7:33Chelsea Finn:So I think that in terms of machine learning, the first thing that you learn in machine learning classes is that you want your training data set to match the distribution of your test data set. And if the training data that you're collecting is reflective of the scenarios that you're going to see at test time, then machine learning will work. And if there's a mismatch between those distributions, then all bets are off. And so the kind of principle there is that we want the data that we're training robots on to reflect the real world situations that they're going to be evaluated on. And I think that there is no getting around that.
8:11Chelsea Finn:I think that other data sources can supplement that and help provide knowledge to robots. But this kind of gets back to like robotics is hard. You have to really you have to solve the hard problems to really make it work. And if we can scale large amounts of data of real robots in the real world doing real jobs, then that's going to be the data that will fuel like large foundation models of robots doing tasks effectively. Well, and obviously having the shirts be crumpled versus straight is a very good example of a micro example of that. So language models have a natural training signal, like the next token, and a massive corpus.
8:55What's the equivalent organizing principle for physical AI?
8:58Chelsea Finn:So I don't think that there's necessarily a direct analog per se between language model objectives and robot objectives. and people have tried to make analogs as well. You really want to be optimizing for what the robot is going to be tested for. And I think one of the convenient things about NextTokenPrediction is that you can actually frame a lot of useful virtual assistant chatbot translation tasks as NextTokenPrediction tasks. And I think that we want to be in the world where we're framing real robot tasks as an objective and a task for training these machine learning systems. And that likely means something that is going to be predicting actions and outputting actions because at the end of the day, the robot needs to figure out how to control its motors to accomplish a task.
9:48Chelsea Finn:One other thing that I think is perhaps interesting that has been a bit of a guiding principle in natural language is that even data that is not really directly doing the task can be very useful. So next token prediction on low-quality internet data can actually play a large role in pre-training large language models. And I think that we actually started to see some of the same principle hold for robotics where lower quality robot data actually can also play a large role and actually improve performance of a downstream robot model, even compared to if you exclude that lower quality data and only include the higher quality data.
10:29Chelsea Finn:And the interesting thing there, I think it might not work for the same reason that it works in language models. But one of the things that we found is that, and kind of my intuition behind it is that if you show like a low quality data of folding a shirt, for example, maybe the shirt was folded in a way that using kind of a different strategy, maybe the final quality of the fold was not very good, maybe it was slower at completing the task. But when you add more diversity to data with different strategies, you'll actually see more variations of how the task is completed. And then if the robot makes a mistake, it might actually end up in one of those variations that was only seen in the low-quality data.
11:04Chelsea Finn:And that low-quality data shows it still had to make progress on the task. And so it essentially gives you more diversity of inputs to the model. And the wider coverage you have and the more diverse data you have, the better your model is going to be able to generalize and handle open-world conditions.
11:19Aria Finger:So you literally wrote one of the definitive papers on MAML. I think we said it was cited 40 ,000 times. and, you know, MAML was about adapting quickly to new tasks from prior experience. And Pi, physical intelligence for our listeners, is doing something similar but at a totally different scale. How much of, you know, the original thinking is still in what you're building now and like what has changed?
11:46Chelsea Finn:Yeah, so I started working on meta-learning during my PhD because when I was working on robotics, I was frustrated by the fact that we would train every task independently from scratch. So we would train the robot to like hang up a shirt on a coat rack. And then we would like just throw everything away and then train the robot to insert a cap onto a bottle and then repeat that process. And there was no, nothing shared across the tasks. There was no reusability. And you would think that after you've learned some base set of motor skills, you should be able to learn the next task more quickly. And I think that the, And certainly, like, we're still using the same sort of principle of how do we reuse across tasks?
12:25Chelsea Finn:And that's really the kind of actually one of the central theses of the company of can we build a general purpose model and that building a general purpose model will actually be easier than trying to tackle an individual single narrow task. And that's also what has driven progress in language models. If you want to develop a really good machine translation system, you don't only collect data from machine translation and train only on that. You start with a really powerful general language model. And actually, oftentimes now, the general models are actually more effective at the specialized tasks than things that are special purpose built for that task.
13:04Chelsea Finn:And I think that the same principles will hold for robotics for actually the reason that we talked about before, where it gives you more diverse data and more coverage over scenarios. And so, yeah, this sort of general pre-training, I think, is really important in this reusability across tasks. Now, going back to meta-learning, the idea behind meta-learning isn't just to be able to generalize to new tasks. It's also to be able to adapt. I think that we still, I think we want to see that in robotics as well. And we've seen it to some extent with kind of pre-training and fine-tuning. But I think that we also want robots to be able to adapt really efficiently to new environments, to new circumstances.
13:43Chelsea Finn:We've seen small proofs of concept of this in robotics, where in one of our recent projects, recent projects, we found that a robot was able to adapt to open a fridge where it was actually ambiguous whether the left side or the right side was the opening one. There wasn't a handle. It kind of tried to open the right side. It was failing, realized that it needs to try the left side, instead moved over to the left side and successfully opened it from the left side. And so it's able to do this sort of like in-context adaptation or really fast adaptation on the fly. We haven't yet seen this proven out on a really large scale where the robot can like kind of arbitrarily adapt to new circumstances and to new tasks.
14:22Chelsea Finn:But I think that we are well on our way towards that. And I'm optimistic about developing that sort of capability in robots in the future. You know, part of the thing, I mean, I think we've all had that experience with a refrigerator of going, oh, wait, it's not this side, it's the other side. So it's actually a good parallel. So let's dive in a little bit for our listeners who may not be as familiar with, kind of like what the system for robots is. So a robot walks in a warehouse or kitchen. It's never seen. What is actually kind of adapting in that moment? The weights, the plan, the way it's reading language.
15:00What's the set of things in terms of the way it's operating in a totally new environment?
15:06Chelsea Finn:Yeah, so the way that we approach the problem is to, like models like ChatGPT, Gemini, and so forth, is to train a big neural network that takes as input a sequence of images and language command and potentially some other information like the joint readings and then outputs how it wants the motors to move and specifically like what you want the angle of each joint to be for the next like half second or so and actually there's multiple like it's a trajectory that it predicts and so a lot of the adaptation or a lot of the generalization to different scenarios is happening within this neural network and so it's implicit it's not kind of explicitly broken down in any way and then from there the in order to handle a new environment there's actually a lot that the model has to do it has to be able to first kind of implicitly interpret where objects are interpret the 3D position of those objects to some extent, like the height of the table can vary across environments, the lighting conditions can vary, which affects perception.
16:11Chelsea Finn:It needs to also be able to translate the kind of a language instruction of what to do into like how that relates to the perception and ultimately figure out how physically it should move its arm in a different way for this new environment. And that all happens just within the neural network. Now, one of the beautiful things about using kind of end-to-end neural networks for this is that you don't need to explicitly have a 3D model of the whole environment. And I think that surely people don't like form a 3D mental model of their environment to like complete a task. There's shortcuts you can take.
16:43Chelsea Finn:And so that allows these systems to, this is kind of the whole principle behind end-to-end training. It allows you to actually more effectively do the task because you don't have to accurately predict exactly what is the friction of the table or what is the center of massive water bottle and so forth.
16:58Aria Finger:So it's funny. I've been recently making my kids make their own peanut butter and jelly sandwiches. And you would think that would be straightforward, but like you've never watched an eight-year-old with a knife just being unable to like spread peanut butter on a piece of bread. And so I imagine it's sort of similar to when you're watching the robots and you're like, oh my God, do it. No, not that. Like, oh, pick it up. And so one of the videos you had showed a robot its task. It had to put some dishes in a sink and then it was asked to put a spatula away in the drawer, sort of a simple task. Even my kids could do that.
17:29Aria Finger:And instead, it opened the oven and put the spatula in the oven because like, oh, it's a drawer. I'm going to open the oven and put it in. That feels like a bug. Like, is this like a window into what the mental model of the robot is? Or what does that tell us about how like these robots are seeing the world?
17:45Chelsea Finn:Yeah, I guess a few notes. So we actually hadn't collected any data with ovens before because we weren't yet at the point where we're ready to start cooking things in the oven. And so whenever it saw a handle, it was overgeneralizing to different circumstances, where it assumed that if it saw like a horizontal handle, that looks a lot like a drawer. And I think there's actually, I'm not a neuroscience or psychology expert, but I think there's actually behavioral studies that show that at early ages, people actually also have a tendency to over generalize concepts in ways that aren't correct. And I think it's actually in some ways a positive sign because it means that the model has some notion of like invariances and some notion of like you'd much rather that than you'd then get the opposite of that which is overfitting where it can only open one drawer and it can't handle any other drawer.
18:31Chelsea Finn:So that's kind of a first note on that. The other thing that I'll remark is we've also seen really interesting notes of generalization in other circumstances as well. So one recent example, we actually haven't shown this publicly on anything, on any blog post or anything, but I found it really interesting, which is that we were recently trying to get the robots to assemble a pinwheel, which involves basically there's kind of a piece of paper that's kind of pre-cut in a certain way. There's a pin, there's a stick. You need to kind of put the pin in the center of the paper and then fold the four flaps and then attach the stick to it.
19:12Chelsea Finn:And we had collected data pretty, like we tried to collect pretty high quality data, but one particular strategy for doing the task where it always involved picking up the paper with the left gripper and picking up the pin with the right, inserting the pin in that fashion. And one of the things that we found is that the robot actually made a mistake and it dropped the pin and it dropped a pin actually on the other side of the table. And we found that the robot was able to actually, in that circumstance, what the robot did is it actually picked up the pin with its left gripper and inserted it into the paper with using its left gripper.
19:48Chelsea Finn:And we actually never showed it any data of how to basically translate that sort of idea, that notion of completing the task with its left arm. And so it seems like it actually learned this sort of equivariance between the right arm and the left arm. And we never explicitly told it that you can do everything with your right arm that you can with your left arm. and so forth. It kind of learned this sort of equivariance between them. And I think it's another example of generalization. In this case, maybe it's not overgeneralization because it was able to do it. And it's actually a very, like it's a less than one millimeter hole to insert.
20:18Chelsea Finn:So pretty precise as well. And I think it's a sign that these models are really learning about what are the things that are constant and invariant across different circumstances, and what are the things that vary. So you went from a 20 % language fall-on rate to 80 % by changing how the action head attaches to the vision language black bone. That's a surprising result. So the architecture affecting whether or not the robot understands what you're saying, what does it tell you about how fragile language grounding is in physical systems? So there's a lot of bits that go into getting a robot to follow instructions.
20:52Chelsea Finn:And the biggest bit is that it actually has to do with the data. So neural network systems are trying to find correlations in the data. And the, there are a lot of circumstances where you actually don't need to look at the instruction in order to know what to do. You can just kind of look at the scene and it's really obvious that you're going to be asked to do something just based off of the scene. And because of that, and especially, this is especially true, just this is true in general. Like oftentimes, like if there's a bunch of dirty dishes in front of you, then probably it's like a good thing to do is to clean the dishes.
Read the full transcript
21:25Chelsea Finn:Or if there's like a bed that is unmade in front of you, probably you should make it. or if there's something that's disassembled and needs to be assembled, probably you need to assemble it. So this is generally true. And it's also true, it can be exacerbated by the way in which you collect the data. Where if you collect the data in a way where people always do the task in a certain way or they set up the scene in a certain way and so forth, it can basically create these really strong correlations between the initial scene and what the robot does in a way that the model will just learn to ignore the language instruction and just do the task based off of the image.
21:56Chelsea Finn:This is the correct thing to do based on the training objective. But it's not exactly what we want because we really want it to do exactly what we tell it to do. So there's a couple of things that we found to be helpful for this. One is to try to actually decorrelate the data and find ways where you have to look at the instruction in order to figure out what to do. And then the second, which you mentioned, is actually a change in the architecture and the way we train the models, which is that the vision language models, which are basically like language models, but with kind of a vision encoder, they are actually really good at language following.
22:30Chelsea Finn:And actually sometimes they have the opposite problem where sometimes they ignore the image because it's really, sometimes you can just answer a question without even looking at the image. And so you could actually leverage the biases of these pre-trained vision language models, which are really good at paying attention to language when using them with robot data. And so the kind of key idea there was to try to kind of more natively plug into the vision language model using tokenized actions rather than using continuous actions to train the backbone of the model. And we separately have a separate kind of diffusion head, which is kind of this wave in which you often train image-generative models basically to predict continuous actions rather than these more coarse, discretized actions.
23:09And that we don't use kind of the gradients of that signal
23:15Chelsea Finn:to train the backbone of the model. You've argued that reinforcement learning is the physical AI equivalent of synthetic data and language models. Robots learning from their own attempts rather than human demos. How far does that analogy go? Yeah, so the most important bit and the most useful bit of this analogy is that synthetic data in language models is incredibly scalable. You don't, like, people are actually starting to complain that the internet isn't big enough. And if the language model can generate its own data, then you're basically just turning compute into data. And if you scale, if you can find ways to scale up compute, then you can find ways to scale up the data source that you're learning from and get more and more powerful models.
23:58Chelsea Finn:And that can recursively feed into the model and the strength of the model in different ways. And I think that analogously, when robots are attempting to do tasks, if they are doing the tasks autonomously and acting on their own accord and attempting the task themselves, learning from that data, that data is also going to be incredibly scalable and more scalable than, for example, trying to teleoperate robots in like a fully human supervised fashion. So it's going to be, I think it's a data source that should be, as robots, as these models start to develop a base level proficiency, it should be a really, like a massively scalable data source for training these models.
24:40Chelsea Finn:And that's especially valuable in robotics where you don't have an internet of like robot motor control data to start with. so I think that that's like the most important point now there are challenges uh even if the robot is acting autonomously you need to make sure the hardware is reliable and doesn't break down or that there are safety uh like precautions in place that the robot doesn't damage itself or damage its environment uh and the and it does also still need to be interacting with the real world to be maximally useful whereas in some language model scenarios you actually kind of do a self play thing where it talks to itself for example rather than talking to a person um and so there are some differences, but the most important bit is that it's a really scalable data source for training models.
25:21Aria Finger:I love the idea that it's like robots, they're just like us. They see a task, they don't read the instructions, they jump right in. Like, I'm putting together an IKEA table. I know what I'm doing here. It's like they make the same mistake. RTFM. Yes. And so when you're like a person who's not familiar with robotics, if they're watching one of these Pi videos, like, what should they be asking to understand what it proves? Like, what questions should they be asking if they're watching the video to understand, like, oh, this is so exciting because it generalized the right-hand instructions to the left?
25:53Aria Finger:Like, how can we know how good we are from watching these videos?
25:56Chelsea Finn:It's really hard to interpret robot demos. And I think that the, it is also, part of that is also, it's not too hard to fake a demo, too. And to kind of show a robot doing something really impressive, but not actually tell you how it was done. And if you aren't told how it's done, it may have been done in a way that actually isn't in the long term going to work or isn't impressive. And so I think that when watching a robot demo, if there are details about how it was done, it's really important to read those and actually understand how that was developed. Was it developed in a way that is going to, in the long run, stand the test of time and be scalable and be something that could be repeated from many circumstances?
26:39Chelsea Finn:Or was it something that was like heavily engineered, required a lot of manpower to like get that one task to work?
26:44Aria Finger:How novel was the task? Did the turkey have to be exactly there on the plate next to the mayo? It just, it feels like you can fake it.
26:51Chelsea Finn:Yeah. And in these videos, you don't see the answers to those questions. Like you don't know like, oh, if I like touch this a little bit, like what would happen? Or if the robot was in a different scene and so forth. So it's really hard as a starting point. and that the text that goes along with these videos can provide a lot of context for why it's interesting and why it's impressive. One of the most challenging things that we have found to convey is actually generalization, because generalization is a property of the training data. And we can tell you that this is a home that the robot has literally never been in before, but the video doesn't show that, because you don't know what was in the data and what wasn't.
27:26Chelsea Finn:And so we try to find ways to show that by actually showing us literally bringing the robot into the home and assembling it there and then running it or showing videos of the robot not just working in one home but multiple homes and doing many tasks in many rooms and in many homes to really kind of try to illustrate that like this was truly generalizing to those environments. Another example of this is like if you want to see reliability, do you just see a video, like one video, or do you see a time lapse of the robot doing it continuously in an uncut shot? We often have been trying to like aim towards uncut videos as well, which really show like over time what it's doing, but it's really hard.
28:06Chelsea Finn:And to back up the, in language models, I feel like you can really understand how well it works by actually interacting with it. With robots, we're not there yet because you need a robot in order to interact with it. But at the same time, I think that we have already started to see people play with these models. So we open source the PIO5 model and people are actually using it and building on it, which is really, really cool. And I think that it's a testament to the fact that these models are actually useful and are actually improving on the state of the art. So when so far as you can actually interact with it, that's when you get a real understanding of how it works.
28:42Well, and part of the thing that's very interesting about, not just technologically, about the generalization capabilities, building up baseline learnings, composability, et cetera, is not just technically interesting, but it's also business model interesting. So, you know, Pi's argument is that every robotics application is historically required building a company around that one application. And a general purpose model layer changes that. So what are the challenges to navigate that and what's the strongest argument against the general thesis?
29:19Chelsea Finn:The thesis, like, feels so obvious to me right now. I think the hardest bit is that I do think that, I mean, robotics is hard and actually deploying robots into the real world is also really hard. And I think the biggest risk is that like everyone fails, that robotics as a whole fails. And that these models are, like even this approach won't get us there because robotics is so hard. There's so many pieces you have to put in place for anything to work. So that's my first thought. A second thought is that for things like ChatGPT and software, you get distribution so easily just by putting things on the Internet.
30:04Chelsea Finn:And so many people have computers, so many people have phones and can interact directly with these systems. And in robotics, we need distribution. Like you need to find ways to distribute. And I think that in some ways, maybe things like self-driving cars will be the closer analogy to the kind of rollout of this technology. And it won't be like immediate. We might not have a chat to BT moment where like everyone is starting to actually like or a huge population of people are interacting with these systems. So those are maybe my biggest points. I think that the problem is really hard. You need really reliable hardware.
30:44Chelsea Finn:And even today, we don't have hardware that is as reliable as a car is, for example, and so forth. So there's so many aspects of it that are hard, let alone the completely unsolved problem of developing intelligence as well.
30:59Aria Finger:So we were just talking about how it's hard to sort of view a video of a robot and know, like, was it a novel situation? Like, what was gamed? How did we do it? But you guys fine-tuned your model on a robot you'd never seen from data you received remotely without even knowing exactly how its actions were represented. And it worked. So, like, what does that actually prove? Because you did a lot of that hard stuff. Yeah.
31:26Chelsea Finn:So one thing that we found that I think is perhaps surprising is that the ability to generalize to different robot platforms and robot embodiments is surprisingly easy. We actually so this was actually technically not the first time that we had done this. So we had previously even before starting physical intelligence, we had a project where we're specifically trying to train on train models on multiple embodiments. In the past, people had only trained models for like one robot platform and assumed that like, if you're trying to get to work across lots of platforms, that would be really, really hard.
31:55Chelsea Finn:And we actually like in that project, we I mean, there's definitely challenges. But the one of the things that we found was that it was ended up being a lot easier than we expected. and we took a model, we iterated on a bit to train it on like multiple platforms. And then we then scaled up that model to a bigger model. And when we scaled it up, we actually didn't tune the hyperparameters at all. We just changed the data mixture. We had previously just trained on one platform. We changed it to train on multiple platforms. We trained one model, sent that model to our collaborators at different universities.
32:24Chelsea Finn:They ran it on their robots. And in most of the scenarios, the model that we sent them was better than the model that they had developed for their project on their robot.
32:33Aria Finger:And so just so people understand, this is different universities who have different robots that do not look the same, have sort of different specifications, and they're using your technology and it is still performing better than what they had trained in-house. Exactly.
32:48Chelsea Finn:And the robots, they look different. They might have different numbers of joints. They might be larger or smaller. Their cameras are set up in different ways, too, like where they're mounted, where they're like even like completely different orientations of the cameras, the height of the table, the setup. There's so many different things. And one of the things that we found is, yeah, if you if you have an expressive enough neural network model that can fit lots of data and you feed in data from lots of embodiments, the model is really good at just being able to handle this. And in some ways, like, people are really good at this, too.
33:18Chelsea Finn:People, we don't just control our own body, but we can control cars. We can control video game characters. And we can very quickly, with a bit of data, learn how to control all of those different things. And it's kind of analogous to that. Although it is interesting to watch people who are not familiar with video games try to do video game characters. But yes. With some experience. Yes, with some experience. So Pi has been clear that there's no simple commercialization timeline. So what are the challenges and risks in possible failure look like? There's these long horizon bets can sometimes be very difficult to pull off.
33:55Chelsea Finn:Yeah, so certainly longhairs and vets are very difficult to pull off. All sorts of challenges as well. The one thing I will say that makes me optimistic, I think I've been like talking a lot about how robots are hard, but I will see that Pi models are already like running in production on multiple different platforms for like one of our own robots, but also for other partner companies as well. And so that, I think, gives me optimism that we are at the point where this technology is mature enough to be useful. There's still a long way to go, but I think that that gives, I think that that definitely gives some optimism that this technology is starting to work.
34:36The, in terms of the challenges and the different routes, the, like I've mentioned before, we think that like this, like solving this intelligence problem that is, hasn't been solved before is the biggest challenge.
34:50Chelsea Finn:and that is the thing that we need to focus on. And so we're orienting ourselves really around that problem and doing things that will help us advance the intelligence of the models, whether that be, of course, doing a lot of research and collecting data and so forth and studying those models, making it stronger and stronger. But also sometimes we find that hardware is a bottleneck and we need to actually improve the reliability of the hardware working on or the teleop device or something like that because that's bottlenecking our ability to make progress as quickly as possible. So we're really oriented around that.
35:25Chelsea Finn:We feel like if we take on customers now or too early, that will slow us down and make it harder to address that core scientific and technological question. At the same time, I think that one thing that's really nice about this is that the model will get stronger if it has real data and data of real use cases. And so it's actually very well aligned to make this model, to continue to develop the intelligence of these models and also figure out how to deploy robots. Because if we're deploying robots and getting real data, that's going to make the model stronger. And so I think it's actually, while we're not focused on revenue, getting customers and so forth, we are trying to deploy robots and using that to get good data.
36:06Chelsea Finn:And so we're both learning about how to actually make this stuff useful and how to make the model stronger at the same time.
36:11Aria Finger:So Professor Maya Maderek, who we had on Possible, and she was talking a lot about robots in the physical form. And what she talked about, which, you know, sort of seems obvious, but the more humanoid a robot is, the more human expectations it carries. And especially if the robot is big. Like, people respond differently if a robot is your size, if it's three quarters of your size, if it's small. Like, all of that carries sort of expectations. And also, like, different levels of scariness. Like we've seen, you know, the Jetsons housekeeper is humanoid. We've seen the Terminator who is humanoid.
36:45Aria Finger:Like no one's really upset about a Roomba. Like there's different expectations when it comes with different kinds of robots. And so you've worked across all sorts of different embodiments of robots. How do you think about the humanoid bet? Like is it definitely going to be humanoids? Is it definitely not? Somewhere in between? Like what is your thinking?
37:02Chelsea Finn:Humanoids are a lot of fun. I actually have one in my office at Stanford. We've done some projects with it at Stanford as well. I think that it's really cool. The different, yeah, they're cool. So there's that. And then I think it'll be one of the embodiments that are deployed. I think that there's a lot of things that are nice about them. At the same time, I think that there's a lot of things that are really unappealing about them as well. So I talked about how robots are not very reliable, like the hardware. And in particular, the price of robots has been dropping substantially. You could buy like robot arms that are far, far cheaper than like orders of magnitude cheaper than they were 10 years ago, which is amazing.
37:48Chelsea Finn:But these robots, while they are fairly inexpensive, they break all the time. And this is even for just a single arm with like six motors, a gripper and an RGB camera. and as you make a hardware system more complex it becomes less and less reliable and so if you're going from like even if it's six motors and like two cameras is unreliable then something like a humanoid is like off the charts unreliable uh and even for the humanoid that we have at stanford it actually takes two to three people to do a single experiment on it uh in order to to safely do an experiment ensure that it's safe ensure that um it's not going to overheat uh really easy to overheat motors as well.
38:29Chelsea Finn:And so if we really want to solve this intelligence problem, we think that we can move fastest by starting with simple systems because simple systems are already incredibly capable. You can do so many things. We've actually been surprised by some of the things that we can do with just like single motor grippers and so forth. Like we trained a robot to light a match, for example, without using any force feedback or tactile feedback. We've been able to teleoperate robots to empty dishwashers to fold laundry. And it's actually very hard to find tasks that you're not able to do with these kinds of systems.
39:05Chelsea Finn:And so if you get the reliability benefit and the simplicity of the overall system, that's a huge win. And if you're still able to do all those capabilities, then it means that you're able to move a lot faster on solving the intelligence part.
39:19Aria Finger:You said the price has dropped significantly in the last 10 years. That's what we expect from technology. How has the reliability changed over the last 10 years? Stayed the same? Gotten worse? Like, is that something else that might get better and better over time?
39:31Chelsea Finn:I think it should get a lot better. And so there are more expensive robots that are more reliable. So you can buy like a$20 ,000 robot that is developed for industrial use cases and is quite reliable, but those robots, they're not really designed for more dexterous, fine manipulation. They're really designed for just repeating the same motion blindly over and over again in a factory. And so oftentimes those design considerations, we still actually use some of those robots because they're nice for some things and they are reliable, but they're not really designed in a way that we would want to use for a really broad range of tasks because the intelligence hadn't been developed yet.
40:15Chelsea Finn:And so they were developed with the software technology in mind. And so I think that I'm hopeful that the reliability should go up significantly. We, I don't, I don't think that there's, okay, I'm not a hardware expert, but I don't think there's anything fundamentally new that has to be discovered or developed to make them more reliable. That doesn't mean it's not hard. I think it is really hard, but it just takes work and it'll take time. Well, comparing software and hardware, you know, software I gave us a culture of ship fast, measure patch, based on the assumption that a bad output is annoying, but not dangerous.
40:51There's obviously some asterisks around that, around if you're talking to a young person about suicide or other kinds of things. But generally speaking, that is a kind of a culture of development. How do you think about the line between good enough deploy and safe enough deploy in the physical world and when those aren't necessarily the same thing.
41:16Chelsea Finn:The main difference is physical safety, of course. And there are many principles in engineering that are really useful here where if you have safety precautions that are redundant with each other, then you can put together a safe robot system even if some bits of the software, like the neural network that's controlling it, are unpredictable or unreliable. Uh, and so the, like, once we have those systems in place, it allows you to, um, to, like, develop these systems. Even if the, like, the neural network part is unreliable, it allows you to develop things that are safe and so forth. Uh, and we've, we've taken safety seriously in a number of different circumstances.
41:57Chelsea Finn:Uh, and we've actually started to give our robots knives, for example, to, like, slice zucchini. There's a lot of things that dives are very useful for, but we found ways to ensure that the people around the robot are safe and so forth. The last point that I wanted to mention is that one of the things that we found somewhat convenient is that also if you use robots that are weak, that don't have super strong motors, that also is another kind of built-in safety layer. Because they're just actually physically not capable of causing a lot of damage as well. And so that's like one additional thing.
42:29Chelsea Finn:And there's a lot of tasks that you can do without super strong motors, for example. I think that once you make a really strong robot, you actually, that no longer is the case. And you need to think even harder about some of these challenges. And some of the things we put in place for weak robots would also apply to strong robots. But that's one additional redundancy aspect that we see in our robots.
42:50Aria Finger:I can't wait for the social media clip that's like, we've started giving our robots knives. It's like everyone's going to freak out. There is obviously sort of a conversation, perhaps a backlash about AI when we're talking about LLMs. Like the conversation has been about knowledge workers, lawyers and coders and doctors. And a lot of the sort of backlash and conversation has been, why are we taking our jobs away? Why are we taking human agency away? Like what are people going to do in the future? And this hits sort of a specific group of workers that wasn't hit in the past by automation, wasn't hit by, you know, machines and factories.
43:26Aria Finger:And so when we're talking about robots, this hits like a different class of people. This is people who drive cars or clean houses. And is there like a different moral question that goes on if we're going to be displacing this other group of workers who's perhaps even further from the creation of this AI?
43:44Chelsea Finn:Yeah, lots of thoughts. I think first, I do think that language models are impacting work of workers that are like more blue collar workers, for example. Maybe not blue collar is the right word, but people that are doing manual data entry, for example. I think that there's a lot of examples of things that are more digital, but also don't involve really high degrees of intelligence that I expect to be affected by language models if they aren't already. So I think it's true for both technologies. I think that ideally we should find ways for these people to be part of the conversation. Another thing that makes me quite optimistic is that a lot of technology is not replacing people, but actually making people more productive and augmenting them.
44:26Chelsea Finn:And I think that we might see the same with robotics as well. And I think we've seen this with language models too. Even with the workforce we have, there are massive labor shortages. There's people doing jobs that are really unpleasant that they don't want to be doing. There's really high turnover for a lot of these jobs as well. And so I think that there is a lot of potential to increase productivity overall, to increase quality of life. And then the, I also think that with language models, and I think we'll see this with robots, is that people are actually using these models for things that they wouldn't have done otherwise, that aren't currently being done by a person.
45:01Chelsea Finn:And so I think that we'll see that with robotics as well.
45:04Aria Finger:I think one of the things that I think about is that so much of the language model side is, we hope, some of it's taking away drudgery. You get like talking about manual data entry. It's like, oh, I'm so excited. I don't have to do the manual data entry. So for my job, I can focus on the stuff that's important. And, you know, thinking about, I just, I wish I knew the exact numbers. They were saying that the invention of the dishwasher and the vacuum freed mostly women from like 10, 15 hours a week of this manual labor that they were doing inside the home, which enabled them to work outside the home, hang out with their kids, have fun.
45:37Aria Finger:And so you could imagine with robots, like there is so much of that drudgery, even if it's not dangerous. There's just drudgery that could be taken care of. I mean, I am thankful every day for my Roomba. That means I have to vacuum my home less. And so I think there's better examples out there than just housework. But even that one, if you could save whatever, two, three hours a week from doing something, it would be a huge unlock for certain segments of people.
46:03Chelsea Finn:Yeah. I also think that you can create, you could imagine robots playing games with them, for example, that you wouldn't otherwise be playing or other forms of recreation. What are some instances? I always think about how AI can be amplification intelligence and work with human beings. What are some examples that come top of mind in physical intelligence? Yeah, so I'm really excited about human-robot collaboration, and I think that that's the most natural example of this. And I've been trying to get some of my PhD students excited about it, too. I think that any scenario where people are collaborating with each other, I think, is a very natural example.
46:43Chelsea Finn:And one thing that comes to mind is if you're cooking, maybe in your own kitchen, can a robot retrieve ingredients? Can it chop ingredients to your specification? Can it clean up the dishes while you're cooking, for example? And so that's one example. I also think that there's actually even more direct collaboration you can imagine where, for example, surgeons often are using robots. And if you could actually not have them fully control it, but actually help stabilize their motions or help kind of augment them to do it like slightly more precisely, give them like a lever to kind of more precisely, more accurately, or maybe a little bit more quickly do a surgical task, then that's something where there's actually a more, they're both controlling, both the AI and the human are physically controlling the robot.
47:31One of the questions in all this stuff, which is, I believe, completely unanswered, although that's why I'm asking the question, is who should be liable when a general purpose robot creates harm in a context that its deployer didn't anticipate? I'm not sure the legal frameworks that we have are even asking the right questions because it's obviously deployer, constructor, environment, etc. Any of this stuff that you've seen some early possible suggestions on?
47:59Chelsea Finn:I don't know. I think that we have this notion of insurance companies, for example, that I think might translate to AI systems at some point. But that wouldn't necessarily kind of tell us the blame or liability. I also think that while we're starting to see these models work in production, we still don't know exactly the shape of the technology and the shape of in the ways in which it's going to be deployed. Is there going to be like a base model and a deployer and so forth? Is it going to be more vertically integrated? I think we don't know those yet. And so the and I think we may be able to draw from from other frameworks, like, for example, like people build furniture that can like if it's like in a way that could fall over, then that there's some liability there.
48:45Chelsea Finn:And so maybe there's some things that we might draw on there or cars or autonomous vehicles and so forth. But, yeah, definitely I think it's early days to figure it out.
48:54Aria Finger:So sort of similar to that, again, you guys are sort of pre-commercialization. Your robots aren't being used in the real world yet. But people always talk about surveillance, wars, border enforcement. And some of this is great. You know, people talk about cameras being able to catch speeders. People talk about, you know, can we make war safer for people? Do you think about sort of the lines that you're like, ah, I'm drawing that line. I won't go into defense. Or sort of how do you think about that sort of sticky moral question?
49:26Chelsea Finn:So maybe first, we are actually, we have robots actually in production that are running our models. Some of our robots and some of robots from other companies as well. So it is actually there. Uh, the, we haven't, we, we aren't currently deploying in any of those applications or really pursuing any of them right now. Uh, I think that there are so many other applications that are, are really exciting and, and really impactful. Uh, and the, um, yeah, haven't yet been kind of faced with a situation where we would really, um, want to strongly consider any of them.
50:00Aria Finger:So we've been talking a lot about robots in the home. They can help you cook. They can clean. Like, how many different robots do you think a person might have in the future? And this could be both sort of in their personal life, in the home, and also in jobs. And various jobs might have many different types of robots that are helping them.
50:17Chelsea Finn:It's really hard to predict, of course. But I think it will be more than one. I think that there will be a diversity of platforms. I think that there is something to economies of scale and the ability to, like, really mass-produce devices. at the same time, we already see scenarios where we have a microwave for one thing and a dishwasher for another thing and so forth. And it seems like as a global community, we have a way to manufacture all sorts of different devices and so forth. So I could imagine it being really diverse and where you have some that are really small and some that are large and some that are like, take kind of one longer form factor versus a more compact form factor.
50:57Chelsea Finn:So I can definitely imagine a world where it's very diverse, maybe you have different robots that are in the home versus in workplace environments, grocery stores, and so forth. So one of the things that I thought a lot about in the chatbot side is how you might have a chatbot that goes with you your entire life. It's like your, you know, companion. And it might be something that, you know, for example, parents might even go, oh, we have a kid and here's the chatbot for the kid as a way of kind of guide and help. Do you see something like that for robots? I mean, you're doing this kind of general purpose and like the question of like, what could be a, you know, not invisible friend, a physical friend, you know, that could also, you know, and not just the one that might go through a person's life with them, but also kind of like how robots will then play different roles as you get to different stages in your life?
51:53Chelsea Finn:Yeah, I think personalization first is super valuable, and both in terms of efficiency, in terms of user satisfaction. You kind of expect it to know stuff about you and to use that to answer appropriately. And different people have wildly different preferences, including very different preferences around how you want things done physically as well. And I imagine that in the future, I think a robot should be capable of being, at the very least prompted in the same way that we like prompt a language model where like I where they give a very detailed prompt like okay today like I really want you to like I really like cooking so please don't like um cook any meals for me but I really like I dislike cleaning a lot and so please like do the dishes whenever you see them not done I also like I don't know um once a week please make sure the trash goes out and and all these other things uh the so I think that that sort of customizability is something that we expect, that I would expect to see in these models, at the very least through prompting, if not through longer-term memory as well.
52:54Chelsea Finn:And then I think that definitely I could imagine, I mean, within the realm of personal robots, there's also the huge other realm of commercial robots and so forth. But for more personal robots, I think that there are extremely different tasks at different ages, ranging from changing diapers to helping someone get up off of the couch or out of bed. And I think that there's also really tremendous value that could be had, especially on the tail ends of those spectrum as well. And maybe that's also a good example of something where people might do something that they wouldn't be doing otherwise. I think that if someone had a robot and that could help them stay independent for longer in their own house, they probably would ask it to do stuff that they wouldn't ordinarily ask a human caretaker to do because they have privacy, because they don't feel bad about kind of imposing on that person and so forth, especially if that robot is stronger and doesn't get fatigued or anything or is able to work at all hours of the day and so forth.
53:55Chelsea Finn:So, yeah, that's, I think, one example application that I think is, like, really cool and would look very different from a human doing that job. So, Rapid Fire, is there a movie, song, or book that fills you with optimism for the future? So perhaps surprisingly, I don't watch hardly any movies and I hardly ever read books. But I have a song, which is also probably very atypical, which is Aurora Awakes, composed by John Mackey. No lyrics, but it's a song that I have played as part of a wind ensemble in the past and I find it to be like... What do you play? I played trumpet for like maybe almost 10 years, something like that.
54:42Chelsea Finn:Awesome.
54:42Aria Finger:So cool. What is a question that you wish people would ask you more often?
54:46Chelsea Finn:I think anything technical. I love chatting about robots. And maybe not all the time. But yeah, I think that that's the most fun part of my job. I don't like email as much like probably most people. And if I get to actually be engaged with the technology and how to actually get something to work, whether it's in the weeds and the details of like where you're like positioning a camera to like how much data you're training on and so forth. I really enjoy trying to get things to work well. So where do you see progress or momentum outside your field that inspires you? I think this is a hard one. I feel like I'm not an expert in other technical fields to really judge if the progress was truly impressive or not.
55:35Chelsea Finn:I think the things like the COVID vaccine development, I think, were really fantastic to see and seeing a large group of people really come together to develop something that could have such a massive impact. I think also outside of technical fields, I think that I often find maybe not even fields, but like individuals to be quite inspiring, whether it be teachers who are really passionate about their job and who taught me a lot. or were inspiring to me in various ways to, I don't know, like competitive mountain climbers that are rock climbers that scale like buildings and rocks that are like, I would have no ability to do myself.
56:18Aria Finger:So our final question is, can you leave us with a final thought on what is possible to achieve in the next 15 years, if everything breaks humanity's way? And what's the first step in that direction?
56:29Chelsea Finn:I think that naturally, I think towards robotics. And I think that the, if robotics is successful, I think that there is like so much drudgery that could be, and people will be basically freed of having to do things that they don't want to do, but have to get done. I think that people will be more productive. And if you're doing stuff that you enjoy you're usually better at it as well and so the ability to have um yeah robots complete all sorts of tasks and like do is like i think maybe even more so than helping like people in first world countries and homes like there's all sorts of other like labor um exploitation and so forth that's also like um incredibly terrible to hear about and i think that the Yeah, I think there's so many opportunities for physical tasks to be completed by technology.
57:28Chelsea Finn:And yeah, it'd be amazing to see that. Thank you for joining us. Yeah, thank you for having me. Thanks. Possible is produced by Pallet Media. It's hosted by Ari Finger and me, Reid Hoffman. Our showrunner is Sean Young.
57:52Possible is produced by
57:53Aria Finger:Special thanks to And a big thanks to And a big thanks to
From the publisher
Useful robots won’t be programmed one task at a time; they’ll need to adapt to unfamiliar objects, environments, and robot bodies. Physical Intelligence cofounder Chelsea Finn (named last month to the TIME100 AI list for 2026) joins Reid Hoffman and Aria Finger to explain how general-purpose robot models learn to act in the messy physical world. She explores why lower-quality training data can strengthen a model, what months of failed laundry-folding attempts revealed, and how the same system can work across different robot platforms. A robot mistaking an oven for a drawer becomes a lesson in the strange line between useful generalization and obvious error. They also examine why robot demos can mislead, how hardware reliability can stall progress, and what it will take to move AI off the screen and into the world.
Congratulations to Chelsea, recently named one of TIME’s 100 most influential people in AI:
https://time.com/collection/time100-ai/2026/chelsea-finn/




