In short
The Generalist Podcast - Episode Summary
Episode Title Why Robots Still Struggle With Simple Tasks (And What Might Finally Change That)
Guest Karol Hausman - Co-Founder & CEO of Physical Intelligence
Episode Description Karol Hausman, co-founder and CEO of Physical Intelligence, discusses the challenges faced in the field of robotics, particularly why robots struggle with seemingly simple tasks, and the innovative approaches that may change that trajectory. The conversation dives into the parallels between robotics and advances in language models, emphasizing the need for general-purpose AI that can operate across various tasks and environments.
---
Key Concepts
- Early Inspiration
- Background: Growing up in a small town in Poland, Hausman was fascinated by robots after watching *Star Wars* and reading related literature.
- Turning Point: A lecture by Sergey Levine inspired him to pivot from a PhD in robotics to deep learning.
- Robotics vs. Language Models
- Historical Lag: Hausman discusses how robotics has not kept pace with breakthroughs in language models.
- General vs. Specialized Robots: The argument is made for building a general AI brain for physical tasks, rather than focusing on specialized robots.
- The Importance of Real-World Data
- Data Collection: Emphasizes the need for real-world data in training robots, highlighting limitations of simulation-based approaches.
- Deployment as Data Source: Deploying robots in real-world scenarios can create a powerful data flywheel, enhancing their learning.
- Reinforcement Learning
- Return of the Technique: Discusses how reinforcement learning is making a comeback in robotics due to improved exploration methods and the ability of models to learn from failures and successes.
- Value Function: Introduces the concept of a value function that helps robots assess their likelihood of success in tasks.
- Reliability Challenges
- High Standards: Robots must operate with a higher reliability standard than language models due to the unforgiving nature of the physical world.
- Intelligent Compensations: Robots may compensate for lack of precision through advanced learning techniques.
---
Main Takeaways
- Human-Like Learning: Hausman draws parallels between human learning (like language acquisition) and robot training, both of which often rely on immersion rather than strict rule-following.
- Philosophical Influences: Hausman references historical philosophers and their ideas about intelligence and reality, reflecting on how these ideas shape modern robotic endeavors.
- Innovation Through Experimentation: The culture at Physical Intelligence emphasizes continuous experimentation and learning rather than adhering strictly to a predetermined path.
---
Pivotal Moments
Personal Journey
- Hausman shares defining moments in his academic career, from discovering a robotics lab that aligned with his vision to pivotal conversations with mentors who shaped his understanding of robotics.
Company Formation
- The desire to create a dedicated organization focused solely on solving physical intelligence led to the founding of Physical Intelligence.
- Collaboration with co-founders who share the same ambitious vision was essential to the establishment of the company.
---
Closing Thoughts
- Hausman leaves listeners with the idea that true innovation often arises not from meticulous planning but from exploring diverse paths and being open to unexpected outcomes. He recommends *Why Greatness Cannot Be Planned* by Ken Stanley, emphasizing the serendipitous nature of innovation.
---
Episode Resources
- Physical Intelligence: [Website](https://www.pi.website/)
- Follow Karol Hausman:
- [LinkedIn](https://www.linkedin.com/in/karolhausman)
- [Twitter](https://x.com/hausman_k)
---
This episode provides a deep dive into the intersection of robotics, artificial intelligence, and the challenges of achieving general-purpose machine intelligence, offering insights from both a technological and philosophical perspective.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe State of Robotics Today
0:00 to 1:19
Explore the disappointing reality of current robotics technology and its lack of intelligence.
“The more robots you deploy, the better the models should get, because I think the models will keep on getting better.”
From Chess to Physical Intelligence
1:19 to 2:15
Understanding the paradox of AI excelling in chess before mastering simple physical tasks.
“long before we taught them to reliably fold a towel or carry a coffee cup.”
Karol's Journey into Robotics
3:56 to 6:39
Karol shares his early fascination with robots and the disappointment he faced in the field.
“So go to granola.ai slash Mario and use code Mario for three months free.”
Influence of Early Mentors
6:39 to 8:06
Discussion about the impact of mentors and teachers in Karol's formative years.
“I don't know if that's as universal as you might imagine it to be.”
The Shift to Mechatronics and Philosophy
8:06 to 11:10
Karol reflects on his transition to studying mechatronics and philosophy, shaping his career.
“There were multiple people that now looking back, I'm really grateful to that either saw something or just encouraged me to dig deeper or motivated me or give me some opportunities that I wouldn't have otherwise.”
Philosophy's Role in Understanding Reality
11:10 to 13:51
Exploring the intersections between philosophy and the nature of reality in Karol's thinking.
“of trying to find those people, move from place to place as I tried to find them and then eventually start working on it myself.”
Philosophical Foundations in Robotics
14:00 to 15:00
Explore how Spinoza's ideas on God and reality relate to robotics and AI.
“What was it about him that seemed deep to you or true?”
The Inner Game of Tennis and Robotics
15:00 to 16:40
Discuss the parallels between learning tennis and robotic motion control.
“I think more and more you realize that there's some underlying structure to things.”
The Journey to Intelligent Robotics
16:40 to 19:20
Karol Hausman shares his path from disappointment to discovery in robotics.
“So you learn all of that structure so that the thinking is that if you know all of this structure, you can follow those rules and then speak like a native speaker would.”
Finding Opportunities in Robotics
19:20 to 21:40
Karol recounts pivotal moments that led him to work on intelligent robots.
“And for a long time, it felt like no one was.”
Show all 31 chapters
Shifting Perspectives in Robotics
21:40 to 24:20
The significance of mentorship and shifting focus in robotics research.
“So yeah, I think this is like yet another person that I'm really grateful to.”
The Power of Insight in Research
24:20 to 26:00
Discussing the transformative moments that redefine research focus.
“Was that like literally the same day, the same week?”
Combining Robotics with Language Models
26:00 to 30:04
Exploring how large language models can enhance robotic learning.
“Yeah, maybe you could tell that story because that is a fascinating one.”
The Taylor Swift Experiment
30:04 to 32:32
Discover the significance of a simple experiment linking robotics and internet knowledge.
“And even though the demonstration itself was very unimpressive, just the fact that you can combine these two knowledge sources, that was really, really, really impressive for us.”
Forming a Team to Tackle Physical Intelligence
32:32 to 35:06
Understand the strategic decisions behind assembling a team to solve physical intelligence challenges.
“And then you have this moment where for the first time you see the light at the end of the tunnel.”
Finding the Right Fit: Lockie Groom's Role
35:06 to 37:34
Learn about the process of integrating Lockie Groom into the team and what he brings to the table.
“like hardware operations and things like this.”
Defining the Thesis of Physical Intelligence
37:34 to 41:33
Explore the foundational thesis of physical intelligence and its implications for robotics.
“Were there certain philosophical alignments that you needed to find in someone who you hadn't worked with before?”
Gathering Data for Physical Intelligence
41:33 to 42:00
Discuss the ongoing challenges and strategies in data collection for developing physical intelligence.
“And the two thoughts that we had from the beginning is that one is intelligence has always been the bottleneck for robotics.”
The Importance of Real-World Data in Robotics
42:00 to 45:30
Learn why physical intelligence emphasizes real-world data collection for robot learning.
“learning, large diversity of data, a lot of real world data, and do the research necessary to figure out how to build these models.”
Challenges of Simulation vs. Real-World Learning
45:30 to 49:00
Understand the difficulties of simulating real-world interactions for robots.
“physical intelligence has done maybe more than others is this real world data approach.”
Reinforcement Learning's Role in Robot Training
49:00 to 53:30
Discover how reinforcement learning enhances robot training and performance.
“And that's what we've seen in locomotion or, you know, backflips or dances or things like that.”
Navigating Commercialization Challenges in Robotics
53:30 to 56:00
Explore the commercialization pitfalls in robotics and maintaining a broad vision.
“To circle back to a topic you alluded to earlier, you mentioned that when you first started chatting with Lockheed, he understood the need for patience on commercialization.”
Optimizing Data for Robot Learning
56:00 to 57:20
Learn about the importance of data diversity and quality in robot learning.
“Is it about scale and, you know, getting as much of that data from, you know, in-person deployments as you, not in person, but real-world deployments as you can?”
Scaling and Research Challenges
57:20 to 58:40
Explore the challenges and opportunities in scaling robot capabilities with existing research.
“I think there is a few of those things that we learned.”
Simulation and Its Impact on Robotics
58:40 to 1:00:00
Understand how advancements in simulation can enhance robot evaluation and learning.
“You know, in this sort of thought experiment where we get no new research, how far do you think that that sort of takes us?”
Surprising Developments in Robotics
1:00:00 to 1:04:00
Discover the rapid advancements and surprising milestones in robot capabilities.
“NVIDIA is probably the best example of a company that is putting a ton of money into improving its sort of simulation engines and getting that to higher and higher fidelity.”
Reliability and Performance in Robotics
1:04:00 to 1:07:20
Discuss the challenges of reliability in robots compared to language models and chatbots.
“Is that sort of something that you see a fair amount in this case?”
Compensating for Sensory Limitations
1:07:20 to 1:10:02
Examine how robots compensate for sensory limitations and how this informs their intelligence.
“So I think we will probably get to similar models in robots where they will be able to predict how well they're doing or how likely the success is much better than we will.”
Compensating for Touch: The Role of Intelligence in Robotics
1:10:02 to 1:12:06
Learn how robots can compensate for the lack of touch through intelligence.
“And instead we put wrist cameras and it turns out that they can compensate for the lack of touch just fine.”
The Limitations of Current Robots
1:12:07 to 1:12:24
Understand the current shortcomings of robots and the importance of intelligence.
“We show some of the most impressive demos ever on robots.”
Book Recommendations for Innovation
1:12:25 to 1:13:36
Discover a book that encourages innovative thinking and unexpected achievements.
“Well, that was a very interesting and gracious way of answering a very baggy question from me, so I'm glad I asked it because that was fascinating.”
Transcript
Automatic transcript. May contain errors.0:00The more robots you deploy, the better the models should get, because I think the models will keep on getting better. They'll be able to absorb more and more data. As I started looking into robotics and started studying it, it was very disappointing. It seemed like there are no intelligent robots out there and no one was working with them. Every robot you would see would be this pre-programmed machine that just goes from A to B and does it repeatedly and has no intelligence whatsoever. That was maybe like a wake-up call. So many people excited about this future that we saw in the movies, that we read about in the books, and it seems like no one was working on it.
0:29We had this one particular experiment with a Coke can in front of a robot and three pictures of different celebrities in front of it. And the prompt we gave to the model we built was put a Coke can on a picture of Taylor Swift. The robot picked it up and then slowly moved it towards Taylor Swift, all from Internet data. That was the moment where it clicked for us, where you can bring in a lot of prior knowledge from LLMs, from the Internet, and connect it to robot motions. And it felt like it opened another door. Maybe, just maybe, if we do everything right, if we combine it with internet knowledge, if we scale it up, if we do all of the pieces that need to be done, it might work.
1:03And at that point, it became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence.
1:18We taught machines to beat the world's best chess players long before we taught them to reliably fold a towel or carry a coffee cup. That paradox is at the heart of the robotics industry, responsible for many of its false dons over the years. Carol Hausman believes that he's found the way forward. The CEO of Physical Intelligence has raised more than a billion dollars in venture funding to build an AI brain for the real world, one that helps robots generalize accurately across different tasks and environments. He's doing so by taking a very different approach from many of his best-funded competitors, focusing on gathering data from actual interactions rather than through simulations.
2:02In today's episode, Carl and I discuss lessons learned from the inner game of tennis, why reinforcement learning is making a comeback, and the challenges of creating true physical intelligence. I'm Mario, and this is The Generalist.
2:32Karol Hausman:With Brex, you get high-limit corporate cards, easy banking, and high-yield treasury, plus a team of AI agents that handle manual finance tasks for you. They take care of things like expenses, all according to your rules, so you can move faster while staying in full control. One in three startups in the U.S. already runs on Brex. You can, too, at brex.com slash Mario. I'm really excited about today's sponsor, Granola. Simply put, Granola is the AI notepad for people in back-to-back meetings. I've been using Granola for over a year now, and honestly, it's a tool that has transformed the way I work.
3:11Karol Hausman:Granola takes meeting notes for you without any intrusive bots joining your calls. You can jot down rough notes like you always do, and in the background, Granola transcribes and turns those notes into clear, useful notes when the meeting ends. You can also chat with your notes, which is one of my favorite features. If someone says something on the call that you didn't quite catch or want to learn more about, Granola can help you out. It's an amazing way to be better informed during a conversation without having to interrupt everyone else's flow. You can also have Granola review all your recent conversations to pull out to-dos, write a weekly recap, or surface interesting ideas you might have forgotten.
3:50Karol Hausman:Another thing I love. To get started with Granola, head to granola.ai slash Mario. And for new users, you can get three months free with the code Mario. So go to granola.ai slash Mario and use code Mario for three months free. I'm so excited to have this conversation today. In studying what you're building with physical intelligence, naturally much of the conversation is around the technology. But the more I looked into your story and the company, it feels like what you're building is really personal and connects to some of your background in philosophy and all these other things. So I'd love to start maybe with some of that first.
4:33When were your earliest memories of being excited by robots? Probably with watching movies as a kid. I think I was a big Star Wars fan. I was reading all the books, watching all the movies multiple times. And I really loved the way that robots were portrayed in that world, the diversity of them, different functionalities, the way they interacted with people, the way that there are good robots and bad robots. But I don't know if I can pinpoint a specific point. I remember just growing up being fascinated by that concept, that you can build something that is so human-like, and I always thought that if you were able to do that, it would allow us to understand ourselves a little bit better.
5:20and I was fascinated by kind of the things that I think people are just fascinated as they're as they're growing up you know asking deep questions about the nature of reality how the brain works what does it mean to to think and I thought that robots probably are going to have some answers to these questions because reading other kinds of books or studying philosophy didn't have good answers to these questions. Then, more importantly, as I started looking into robotics and started studying it, the first impression I had was that it was very disappointing in that every robot you would see would...
6:01Basically, it seemed like there were no intelligent robots out there and no one was working with them. Every robot you would see would be this pre-programmed machine that just goes from A to B and does it repeatedly and has no intelligence whatsoever. So I think that was maybe the most defining moment for me when I saw that, because that was like a wake-up call of, you know, there's so many people excited about this future that we saw in the movies, that we read about in the books, and it seems like no one is working on it. You said that, you know, in some ways you had, you know, that you're asking the classic questions of every child about the nature of the universe and the human mind.
6:39I don't know if that's as universal as you might imagine it to be. I think probably there are some kids who really become fascinated by that, but maybe not all. Were your parents scientists? Were there things that they did that sort of encouraged that natural curiosity that you must have had? Not really, no. My dad was a mechanic, and my mom was an entrepreneur. So not really. There wasn't a lot of science at home. I think it was mostly coming from watching movies and being fascinated by it. I also went to, so I come from a small town in Poland. I went to high school that had very high talent density for that area.
7:22and I remember that being an environment where a lot of these questions were being asked and it was very encouraged. So I found a group of friends that were really fascinated by similar things and I think that would really increase my interest in that area. It was just a thing that you would talk about in high school. You had a tweet that I saw where you were talking about Fei-Fei Li's biography and how it sort of spoke to you in some sense that you know the immigration story and also the fact of, you know, having teachers that believed in you along the way were sort of parallels to your own life.
7:57In high school, was that when you sort of had that first batch of those teachers who maybe saw some of that, you know, really remarkable early talent? I think it was throughout. It wasn't just high school. It was throughout. There were multiple people that now looking back, I'm really grateful to that either saw something or just encouraged me to dig deeper or motivated me or give me some opportunities that I wouldn't have otherwise. And I think at every single stage, there was someone I could point to. And reading Faye Facebook just reminded me of that. And I remember after I read it, I thanked some of them and thanked them for everything they did for me.
8:38So yeah, it was a long journey from Poland through Germany to the US, going through multiple places in the U.S. and it felt like at every single stage of the journey there was someone kind of rooting for me and helping me. You know, before you realized that you wanted to work in robotics in particular, as an undergrad you studied sort of computer science and philosophy, but I'm curious if there were interests before that. Like as a kid were you thinking I'm going to be a astrophysicist or, I don't know, a neuroscientist. Yeah, I was really into physics, actually. Right after high school, I applied to a lot of physics and math programs.
9:19And I thought I was going to be an astrophysicist or a theoretical physicist. I was really fascinated by it. I was reading a lot of books. I was really into physics in high school. But then a lot of my friends decided to go for engineering degrees and we're all gonna go to Warsaw and I thought that I could maybe try to study both or I could start with engineering and then go back to physics if it turns out to be boring it wasn't like a very intentional choice at that time it was more like these are really smart people I really want to stay with my friends and then and robots are really really cool so let me try that and then see how it goes so I went to I went to Warsaw to study it was called mechatronics at the time, which was basically like a mixture of electronics, mechanical engineering, computer science.
10:06Then I also studied philosophy. But I think that's when, once I moved to Warsaw, that's when I started realizing that robotics is a thing, but it's a thing that is very disappointing. I remember studying robotics in Warsaw. The kind of the dream job you could get was being a roboticist at an Avon factory that was right next to Warsaw. That was like the job that everybody was competing for if you could work for avon at the on the production line that's like that's how you made it in the robotics industry and i remember i went there once and saw that production line and it was the most uninspiring thing i've ever seen with robots as i said before robots just moving from a to b not being intelligent whatsoever and i think that's when when it started clicking for me that i would really like to do something beyond that and it seemed like no one was working on it, which sparked my interest even more.
11:02So that's where I started digging who is doing intelligent robots and is it even a thing? And that's then what defined, I think, the rest of the journey of trying to find those people, move from place to place as I tried to find them and then eventually start working on it myself. The sort of study of philosophy, as you've looked back on that time, maybe it's a line on the CV at this point and you're like, you know, I don't reflect on it necessarily that much. But I'm curious if there are writers that you've ended up coming back to more and more as you start to do this sort of work around, I don't know, the nature of consciousness or the, you know, the certain senses that make us more or less human.
11:44It was a really interesting period. So I think there was a few realizations I had while studying it. I think one was that a lot of philosophy is the history of philosophy. It's not philosophy itself. It's talking about what others thought about years ago. From the outside looking in, it always sounds like Plato or people like that. They were these geniuses that kind of figured everything out way before their time. But then when you actually read their works, and that was part of the deal of studying philosophy, you realize that there's a very small percentage of things that kind of sound right.
12:27And there are a lot of things that they said that are absolutely wrong. Yes. And that was quite eye-opening. That, you know, the history of philosophy shows that there are people who made many, many mistakes along the way and thought that the world worked a certain way when it's now very clear that it doesn't. They had some really interesting observations. And I think we often over-index on those. But we were wrong a lot. there were still some philosophers that I remember studying that I think were very relevant and I was a big fan of. One of them is Spinoza. That was, I think, my favorite philosopher when I looked back at what I studied.
13:06But then there were some subjects, very, very few, that were more about philosophizing today rather than looking at the history. And I think those were the ones that were the most fun for me. There was this one subject called ontology, which was a study of things, of how they are. I think it's even difficult to comprehend what this means. But it was one of those subjects where you just sit and contemplate reality. It just was very fun, even though it was very unstructured and it wasn't clear what is right and what isn't. It just seemed like it touched something real. and probably does the subject I remember the most.
13:51That's where I had the most fun, much more so the history of philosophy. So interesting. And these things I love to explore because there is, to me at least, such deep parallels between some of the things that you're trying to build to fruition. Why Spinoza? What was it about him that seemed deep to you or true? The one thing that really resonated with me was this idea that, you know, at that time, a lot of philosophers were contemplating the nature of God and the nature of reality. And there was often this idea that God is this separate entity that either was the creator of everything or had some kind of way of controlling everything that is happening in the universe.
14:34And I think with him, what was so interesting was that he said that everything around us, like all of reality, that is God. That is what it is. And there is some underlying structure to it. And that structure in itself is where the beauty lies and where the intelligence lies. And I thought that was very thought-provoking. And as you study physics and things like robotics or AI, I think more and more you realize that there's some underlying structure to things. and in some ways that structure is quite counterintuitive, like why it's there. Why is it so that you can describe all the complexity of the world in a simple equation rather than in a million different equations?
15:22And the more beautiful and short that equation is, the more real it seems or the more accurate it is. And I thought this was kind of the closest to what it feels like today given the current state of science. You had a tweet where you were talking about a book that I tried to read as much of it as I could in preparation for this called The Inner Game of Tennis. And you sort of mentioned how, I mean, you were really talking about tennis, but it seemed like some deep parallels to robotics. You sort of make the point that they talk about the fact in this book that if you really focus with your conscious mind on improving your stroke, you sort of get to this local maximum where maybe you are doing a little bit better at hitting a forehand, But really what you need to do to get to an optimal stroke is sort of allow the unconscious mind to take over in some respect.
16:16How do you think about emulating this aspect of almost the unconscious and the role it plays in fluid motion and, I don't know, proprioception and all these sort of things when it comes to robotics? Were you thinking about robotics when you were reading that book? a lot yeah definitely I was I don't think it just applies to notion I think it's there's a big difference between how we learn things and how we think we learn things or how we teach things to others you know even if you look at something like learning a language the way we we thought you're supposed to teach it or the way we would want to teach it to to others in school is by defining all the rules and you know you first learn about grammar and all the different concepts and how tensors work and what word follows what other word.
17:03So you learn all of that structure so that the thinking is that if you know all of this structure, you can follow those rules and then speak like a native speaker would. But it's not how we learn language, right? Like you're just immersed in it. You just hear everybody using the structure that they kind of intuitively know. And then all of a sudden you emerge with a full understanding of how language works, even though you can't really pinpoint any of these rules. There are many native speakers that speak perfectly, follow all the rules exactly, but have no idea what these rules are, what the underlying structure is, what the declination is, what tenses are.
17:41They have no idea of any of that, but they still follow them perfectly. with that book I think that was that was shown but in the in the in the motion aspect of things where you know you could read all you want about how to play tennis or you can try to describe you know where your elbow should be exactly where you hit the forehand or how to adjust your swing but the only way to really learn it is by just doing it and kind of immersing yourself in it and I think there is a lot of parallels to the robotics world and to the world of AI, how we thought that the machines should think or how they should learn things versus how they actually learn them.
18:22You, I think in 2023, say that you had a moment where for the first time you really felt like you could see a bit of the future of the robotics industry and things were really clicking into place. But you had really been in that field for quite a while already, it strikes me. So I wondered what the period before that moment was like. Like, was that quite a lonely time to be working in a space where you don't know how it's going to play out? So maybe I can go back and tell you a little bit more about the journey. I would love that. So you understand that kind of the feeling. After I moved to Warsaw, as I mentioned, it was very disappointing to see the state of robotics.
19:05And for a long time, it felt like no one that's working on the intelligent robots that I really wanted to work on. And I remember at the time, I was searching for really anything, anybody who is doing something that is a little bit closer to what I imagined robots to be. And for a long time, it felt like no one was. Then I remember, I finally found someone, I found a professor in Germany that was writing papers that seemed a little bit closer to what I would imagine robots to be. And there was this master's program in in Munich called Robotics, Cognition, and Intelligence. And I thought if I apply for this program and get in and they don't do intelligent robots, then I'm basically lost.
19:48If the program with that title doesn't do anything that I imagined, then something is very, very wrong. So I applied and moved to Munich. And I remember the first day of classes, I went to one of the very first classes, and I got stuck in the metro somewhere and was super light, so missed the entire class. And I showed up. No one was there. The class was over at that point. But I saw a janitor, and I was looking for a bathroom, so I asked him where the bathroom is. He pointed me to the third floor of the building. I remember it very vividly. And as I went upstairs, I remember seeing these big, big doors and there was some mechanical sound coming from behind them.
20:32I don't know why, but I just decided to like open them and see what's going on there. The sound was very intriguing. And as I opened these doors, I see these two basically humanoid robots doing, one of them was making popcorn and the other one, I believe, was spreading butter on a toast. And that was the moment where, you know, this is something I've been searching for at this point for years. So that was the moment that was probably one of the most powerful moments I had during that journey, where I finally found it. There's robots that look intelligent, that do something that is very human-like.
21:07There are people who work on this. So I immediately left the room, because there's basically no one there. And I knocked at the room next door and asked if I could work there. and basically I was completely unqualified. I had no idea how to program these robots. I was not familiar with any of the tools, but the person who was there was another one of those people that gave me a chance and asked me to just come in on Monday and see what I could do. Wow. I got involved, got a job at that lab. That's how I got into finally working on intelligent robots. So yeah, I think this is like yet another person that I'm really grateful to.
21:46and then afterwards I moved to the US got into a PhD program studying intelligent robots and at that point I was so stoked that I finally get to do this and I finally found a small group of people that work on those things it didn't matter that it doesn't work that well I was so excited that I finally get to meet the people that are working on this I can learn from them, I can push it forward and I can finally work on the thing that I always wanted to work on So I did this throughout my PhD, and I had another moment like this during my PhD, this kind of defining moment where I was working on a certain sort of problems where it was referred to as active perception.
22:29And I can go a little bit deeper into this if you're interested. But at that point, it kind of felt like there are ways to write papers, to progress in your PhD. but there isn't anything that felt like it could actually solve this problem. It all felt like the world is too complex. You can push it to some extent, this technology, but you won't really get to the finish line. They're not kind of adding up. You don't really see a path out of it. And there was another moment where a postdoc candidate stopped by our lab. And at that time, when a postdoc candidate stops by, you are supposed to as a PhD student show them your work and kind of talk with them as part of the interview.
23:16So I showed to that post kind of that stuff that I was working on and he listened to all of it and then at the end of it he said that I should drop all of this and switch to deep learning. That's the way to solve the problem that I was actually working on. At that point I just thought to myself that he had no idea what he's talking about and he probably didn't listen to me because at that point I think I spent already two years or something like this working on the topic I was working on. But then later that day, I went to his lecture where he showed what he was working on. And that was another one of those moments where it was extremely eye-opening because for the first time I saw something that could actually work.
23:56Where it was less of like, here's a little paper over there and a little paper over there, but they don't add up. But it's something that kind of brought all the pieces together. and yeah that was another one of those moments like the similar one to the one I've experienced in Munich where at that point I decided to drop my PhD topic change it completely start collaborating with that person and do everything I can to push deep learning in robotics and that person was Sergey Levin who is now my co-founder at Physical Intelligence and one of the pioneers of deep learning in robotics and yeah I never looked back after that but to kind of this is a long way of answering your question but it didn't feel lonely in that you know i was just really excited to be working on that problem that's such good stories there the first one is almost like you know um when a wizard is accepted to hogwarts or something it's like you finally found your you know collection of of magical people and and the magic that they were doing the sergey story is also so interesting to me because you seem to have gone from skeptical to convinced very, very fast.
25:08Like what was the cycle on, on that? Was that like literally the same day, the same week? I think it was within that lecture. I think like in the middle of that talk, I was like, yep, this makes sense. This is a hundred percent that whatever I was doing is wrong. This is the right way. Uh, I want to do everything I can to work on, on this topic. And I had feelings like this. I think every researcher has a feeling like this every now and then, where they just see something and it clicks. And at that point, it's a very powerful feeling. You kind of want to drop everything you've been doing, and you just feel like you found something that is much closer to truth than what you thought before.
25:50I think we all had a similar moment afterwards when we were working on combining large language models with robot learning. And that was, I think, one more of these, another of these powerful moments where like things start to click. That was the Taylor Swift demo? That's right. Yeah, maybe you could tell that story because that is a fascinating one. Yeah. So what happened afterwards as I started working with Sergey, there was a few of us who were really getting into that field. It kind of felt like we were a small group of renegades within the robotics community because there were a lot of problems with deep learning at the time.
26:28And some of these problems remained. Like it was not interpretable. There was no way of really making it modular. It seemed like nobody fully understood how it worked. It wasn't very sample efficient. And there were all of these problems. So like at that time, it was extremely unpopular to write any deep learning in robotics paper. And it was kind of like between these two worlds of machine learning world, where any robotics paper seemed kind of like the wrong paper for that venue. the robotics world that never fully embraced or didn't embrace deep learning at that time. But it was still really cool to work with a small group of people and push these methods forward.
27:06So then afterwards, the only place that was really embracing that view was Google Brain. So I decided to join Google Brain right out of my PhD, continue working with Sergey and others on those set of methods, with the basic premise being that the way to really get it to work is to scale it up. So we're scaling it up, we're trying to figure out how to really get it to work at scale, but again it started feeling like there is so much complexity in the world that if robots need to, the only way for robots to learn all of this complexity is to experience all of it firsthand, it'll be very difficult to scale.
Read the full transcript
27:45It will be very difficult to have them learn about logic and how to break down a task if they had to do it all by themselves. So then I think around 2022, I would need to look back to see what year it was exactly. We started combining it together with Brian Liechter, my other co-founder. We started combining these robotic methods with large language models. This was before the chat GPT moment. This is where people just started experimenting with large language models and kind of understanding them better and better. There was this one particular demo. So we were really excited about this because there was one, we thought that this would be a path to bringing in a lot of prior knowledge that robots didn't experience firsthand that we learned from the internet into the robotics world.
28:32And that would kind of solve this problem of you having to collect so much data to understand how the world works, because that understanding is already embedded in large language models. So if you figure out a way to combine these two, that should really, really help. And we had this one particular experiment where we were testing how much transfer do you get from this internet-scale knowledge to robotic behavior. And usually when you work on robots, you work on this very specific task and then you test that specific task. So you're very rarely surprised. It's mostly disappointing because you really want this task to work.
29:10you've been working on this task for a very long time and then it usually doesn't work but if you work hard enough eventually you get it to work but very rarely you are in a situation where you didn't expect the task to work or like it's not the task that you're working on and it works so very early you're positively surprised so we set up a few of these experiments trying to test how much knowledge transfers and one of them was this experiment with a coke can in front of a robot and three pictures of different celebrities in front of it and the prompt we gave to the model we built that combined LLM and robot models was put a cocaine on picture of Taylor Swift.
29:46And one of the pictures was Taylor Swift. And you can see this video on the internet. It's a pretty pathetic demonstration of what robots could do at the time. But the robot picked it up and then slowly moved it towards Taylor Swift. And that was another one of these moments of huge, huge excitement even though if you watch the video it's totally unimpressive. Because this was the robot models had never had the chance or had never had any of taylor swift their data it had to understand the concept of taylor swift connected to the image of taylor swift and then connected to the right motion that would move coke and to the picture of taylor swift all from internet data so that was the moment where it clicked for us that would actually work where you can bring in a lot of prior knowledge from llms from from the internet and connect it to robot motions.
30:36And even though the demonstration itself was very unimpressive, just the fact that you can combine these two knowledge sources, that was really, really, really impressive for us. And it felt like it opened another door. Yeah, it feels like, I don't know if you're familiar. I'm sure you are. I can't believe I'm asking if you're familiar. You know much more about this than I do. But there was some experiment in like the 70s, I think, called like Shardloo or something like that, where the entire approach was to try and teach AI all the rules of the world programmatically and to say, you know, you know, here is the category of what a bird is.
31:12It has these characteristics. And then here's the category of a bat. It has some overlapping character, you know, all of these sorts of things. But but essentially, you know, people are familiar with this concept now because of the way we use LLMs. But it was almost as if the robot suddenly inherited all of these rules and, you know, pieces of knowledge from tying them up with the LLMs that allows the Taylor Swift demo to happen. Yeah, that's exactly right. And I think there's maybe two points there. One is that we made the same mistake in robotics. We wanted to write all of these rules. We thought that if only we had enough of these rules, the robots would be able to follow them and do the right thing.
31:51But kind of like we said about the inner game of tennis, you can't just write all of the rules. You kind of have to do it. And there is some underlying structure, but you can't just fully put your finger on it of what it is. You need to learn it from data. And I think people thought about it similarly in language as well. They thought that, you know, if only we could write all of the rules, that would be enough. But it turns out that there is, you know, trillions or billions of these rules. And sometimes we can't fully even express them language you just need to learn them and if you do then you would be able to follow them even though you still don't understand what fully what they are so i think we're learning this lesson over and over again and so you know you have this almost third eureka moment for yourself where the technology is you know taking yet another jump by that point it sounds like you have most of the the physical intelligence co-founders around the table but how did you sort of pull the the last few members aboard and decide to make that leap yeah i think at that point it started becoming clear that it could be possible and you know if you worked on something for so long at that point that was basically my entire adult life 15 years or so and for a long time you thought that there was no solution to this problem there were some times where it felt a little bit more tangible but it never felt like it could actually be solved.
33:18And then you have this moment where for the first time you see the light at the end of the tunnel. Like maybe, just maybe, if we do everything right, if we combine it with internet knowledge, if we scale it up, if we do all of the pieces that need to be done, it might work. If you see that light at the end of the tunnel, you can't unsee it. You want to do something about it. And at that point it became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence, is to solve this problem. It can be solved as priority number 20 in another organization.
33:53The only reason to exist for this organization would be to solve physical intelligence. So then the next question is how do we actually do it and what are the conditions to make it happen? And what became clear immediately is you need to have truly the best people, truly the best researchers to make it happen. Because the second best team in those areas doesn't really work with a company like that. You need to have the top top people in the field. So that one actually wasn't that hard because we were already working with each other for a very long time. So I happen to be friends with the best people in the field already.
34:33So it was just a matter of making sure that we all want to do it and we're ready for it. the other condition was that you need to get access to a lot of funding it would require very long very long term bet and investors that are fully aligned with this taking some time and this starting as a research company and not being oriented around revenue or around the short term revenue so it was the second requirement and the third requirement was just like build a very very build an incredible company with the right focus with the right people, with the expertise in all the other areas that we didn't have expertise in, like hardware operations and things like this.
35:14I spend most of my time figuring out these two points. How can we find the right investors that can really help us in this adventure? And they're fully aligned with how we want to do this. And then how can we find the right people to to fill all the missing pieces. How did Lockie Groom join the crew? Because he's obviously had a very impressive career as an operator, as an investor, but isn't someone who one would automatically think, you know, this is someone who's going to devote the next chapter of their life to robotics. So I didn't know Lockie before we started thinking about starting a company, but then, and he should tell the story rather than me, but I believe what was, what was happening is that he's been investing in that point for a few years.
36:02And as long as he's been investing, he always thought that he doesn't want to be investing forever. He wants to find something else that he can fully devote himself to. And he wants to build things. And the one area he was always fascinated by was robotics. And he was seeing over the past few years that there was a lot of innovations happening there. There was some kind of moment. Robotics was having its moment. and particularly he was impressed he was seeing over and over papers coming from chelsea finn's and sergey levin's lab um as well as our papers from google so he he talked to a lot of um a lot of his friends that if ever chelsea sergey or carl or anybody from those teams are thinking of starting a company please connect me to them i would at least want to invest or at least talk to them so that's when we got introduced by um by a friend and we pitched to locky as a way to um to to have him invest and at the end of that pitch it became clear that he would like to do much more than just invest this is the opportunity he's been looking for for all the years he's been investing and he really wanted to to dive in um so then the next step was us trying to get to know each other as quickly as possible, spend as much time together as we could.
37:20And then as we were doing that, it also became clear that this is the person that I've been looking for that could help us on the commercial side, on fundraising side, on operation side, kind of one of the missing links that we really needed. Were there certain philosophical alignments that you needed to find in someone who you hadn't worked with before? What were the sort of, you know, even beyond maybe the tactical pieces, the traits that you saw in him that made you feel like, yes, this is someone I can bring into this very trusted group where we've already built all this history together? There were a few.
37:59I think maybe the biggest one was that he immediately got it. At that point, there were a lot of investors or other people that I talked to, other business people that I talked to. and it was very hard to get that idea across that idea that you you have to do research right you need to get you need to build the technology first and you can't be distracted by short-term revenue and if we do do this right this is going to completely change the world and it's going to be most valuable business of all time but you need to you need to have the patience to let it let let us do it the right way rather than short-circuit that path and and kind of cap the ceiling and he got it, I think, within the first minute.
38:42So that was very reassuring. And then on top of that, what I really liked about our conversations was that in many of those previous conversations with other people, I always felt that I'm the one pushing the ambition of this project, of how far it could go, if I set it up right. And I think this was the first time where I had somebody else set up even higher ambitions for us. That was really cool to see because it felt like we're all pushing in the same direction. He will make us even better and more ambitious. So I was really glad to find someone who will, you know, make sure that we are not going to, that we're not going to short circuit this.
39:23He's fully aligned in doing this in the biggest way possible. He makes us more ambitious and he's the person in the areas that we really need more expertise. And he'll make sure that those areas are fully aligned with how we want to develop this company. So it was basically a no-brainer at that point. You had this insight that this was the right moment to go and solve this problem. But at least from the outside, one could imagine many different form factors that might take. Like another version of this company might be saying, hey, we're going to try and build the best humanoid robot ourselves.
39:59Or, you know, we're going to take this approach to getting the data we need to run these models very effectively versus another approach. How did you sort of land on the version of physical intelligence as it is today? I think from the get-go, we had an idea of, we had the thesis for the company. And that thesis was around all the research that we've done up until this point. That similarly to what happened with language, it's not going to be specialist models that's really going to solve this problem. It's going to be generalist models that work across all kinds of different tasks, all kinds of different environments and all kinds of different robots.
40:39And it was similar to what we've seen in language, where you would think that the best translator would be just specialized for translation, or the best coder would be just specialized or just trained on code. But it turned out that the way to be the best at all of those fields is to train one generalist and generalist that takes in poetry data and coding data and translation data, and then it turns out to be much better than all of those specialists at those specialist tasks. and we started seeing something similar in robot learning in robotics that if only we could collect enough data if that data was very very diverse and if we do it the right way we should be able to build the best generalists and if we solve that that is really going to allow us to have this world of many diverse form factors and and robots that can be very intelligent so i think the kind And the two thoughts that we had from the beginning is that one is intelligence has always been the bottleneck for robotics.
41:39Rather than trying to start a robotics company that focuses on a specific robot, how can we tackle this problem head on and just focus on the intelligence? And then the second thought being that the way to solve intelligence is to take the lessons that we learn from vision and language and other fields. and really take the foundation model approach, which includes things like cross embodiment learning, large diversity of data, a lot of real world data, and do the research necessary to figure out how to build these models. So you sort of land on this idea of, you know, the AI brain for these different robot types, to put it sort of simply.
42:17How did you start to think about the right way to gather the data? Because that's clearly such a big piece of it. And obviously there are different pieces, different players that take different approaches there. So I imagine you must have had to reason through that in a million different ways to land on the one you have. I think we're still reasoning through it. I don't think we have all the answers yet. I think the important piece about how to think about physical intelligence is we are not very dogmatic. It's not that we sit down and think very, very hard and then come up with the solution and this is our bet.
42:54I think the better way to think about it is that we are really truth-seeking and we run experiments and we don't know the solution and we know that we don't know the solution. So we want to follow the scientific method and really try to find the truth and really try to find what works and what doesn't. Because I think there is a true answer out there. We just need to be very humble in finding it. So that's, I think, how we arrive at the current set of answers. We run a lot of experiments. We try many different ideas and we see which ones stick and then we double down on them. Now in terms of, based on what we've seen, based on the evidence we've seen, how we think it's going to work or how I think it's going to work.
43:36These models will need to be able to absorb very diverse data sets. Where it's less about picking the right data or picking the right way of collecting data. More about building the engine that allows you to absorb all kinds of data. whether it's videos of people, whether it's teleportation data from robots, data from handheld devices video data really, or simulation data really anything. And the more data they can absorb the better the model will be. And I think we're kind of at this stage of robot learning where we kind of try to throw anything we can at these models build them in a way that they can absorb as much of it as possible and get them to the threshold of being able being deployable, where you can actually deploy them in the world and have robots out there collecting data for real, doing economically valuable tasks.
44:28And I think once you're at that stage, once you're at that threshold of they can actually work and do valuable tasks, that's when you enter this next stage of now deploying robots at scale and deploying it in multiple verticals and diverse environments and actually delivering value. And I think that second stage is actually going to be the stage where we get most data from. Where the more robots you deploy, the better the models should get, the more you can deploy them and there's a natural flywheel to it. And what's exciting about the moment today, this moment right now, is that I believe we're very close to this threshold.
45:04We're already above that threshold. And that's something that is really, really exciting. Because I think the models will keep on getting better, they will be able to absorb more and more data, but they will also have this sustainable source of data that is very, very valuable because that's the data that is the most real, that is the closest to how you actually want to deploy the robots. For folks that maybe haven't spent as much time digging into this or sort of coming to it fresh, I would say that one of the experiments as you think about this data collection that physical intelligence has done maybe more than others is this real world data approach.
45:38Why is that so important and so valuable to get and so valuable to enter that phase too that you talked about that can sort of start the flywheel? We have a few theories why this is really important, but I think maybe the meta point here is that if there was a different path that worked much better, like maybe for simulation or from videos or something else, we would happily pick that path. So again, it's not that we sat down and thought that this is the best path and this is the only thing we're going to do. We run a lot of experiments, we tried out a lot of ideas, and this is the one that seems to be working very, very well.
46:15Why do we think that this is important? I think it's kind of difficult to describe it in the absolute. It's, I think, a little bit easier to compare it to other alternatives. So one popular alternative is simulation. And this is, in particular, an area that had a lot of success, especially in locomotion use cases, where you have robots walking around or doing stunts, like backflips and things like this. Most of these methods are training simulation first. And I think for those kind of tasks, the main complexity is about how you move your own body. It's less about interacting with the world and more about how do I move my legs correctly so that I can walk or I can run.
46:58And in that sense, as long as you model your own body accurately, you should be good. If you model it very, very well, your own particular robot, and that transfers from simulation to the real world, you should be able to learn that behavior in simulation and then that's good enough to them to work in the real world. Now, when it comes to the problem of manipulation, where you're manipulating the world around you, the difficulty is less about modeling your own body, like how you move your arm from A to B, but it's more about modeling how the world will react to it. So the problem of manipulation is more about this interaction with the world that you have, or the interaction of the world that you're interacting with.
47:40And I think in this case, it's just much harder to simulate the world around you than it is to simulate your own body. It's no longer about just a single robot that you need to simulate. You need to simulate everything. And we don't really know how to simulate everything at scale, how to do this accurately enough and scalably enough so that it would work for anything. For every single one of tasks or objects, it takes really long to get it exactly right, to get all the friction parameters right, to get the simulation behaviors right. And it's just not as scalable. That's been our finding so far.
48:14And so it's sort of the case to try and distill that, that, you know, with simulation, you can do some of these, I don't know, maybe this isn't the perfect word, but sort of coarser, larger actions that are more self-contained. But once you start trying to, yeah, implement picking up the coffee cup or the towel or whatever it might be, you're starting to rely on a simulation that would have to be so good, you'd effectively have to, you know, create a simulation that is perfectly similar to reality. And that's where the real world data starts to become so important. Yeah, you would need to simulate all of the external world.
48:53Yeah. Any one of those tasks. And that's just too costly, too difficult to do. But as long as you can get away with just simulating your own body, I think that works perfectly fine. And that's what we've seen in locomotion or, you know, backflips or dances or things like that. There was a moment, I think, about a year and a little bit ago where you talked about reinforcement learning making a comeback. And that has since become a really interesting part of the way that physical intelligence seems to work with recap. What were you seeing at the time that made you think, you know, this technique might have sort of a second or additional wind here?
49:33And why has that been so useful for you? Yeah, there's been a long history of applying reinforcement learning to robots. And the problem of reinforcement learning is such that you need to be learning from your own experience. And while gathering that experience, you need to encounter some successes. And then you want to increase the probability of good actions that lead to those successes and decrease probability of the actions that lead to failures. What this means is that you need to have, as you explore the world, as you collect your own experiences, you need to have some of these successes, because if you don't see any of them, you basically don't really know where to go.
50:13You're kind of lost. And if you start from scratch, if you just try to command random commands to a robot, it's very, very unlikely that you'll encounter a success based on that. So let's say you're trying to teach a robot how to grasp an object, and you command basically random commands to all of the motors of the arms, it's very, very unlikely that some of these commands are going to lead to a single success of you grasping an object successfully. And that's what we often refer to as the exploration problem. But we don't really have a way of guiding the robot on how to explore so that you can encounter a success.
50:49And as soon as you encounter it, then you're on the flywheel, you can now increase the probabilities of those actions. is just to get to the very first one very, very hard. And that's been the problem of reinforcement learning for a very long time. But now with the models that we've been putting together, this exploration problem became much easier because robots now don't start from scratch. They start from a foundation model like pi0, pi05, or pi06, where they already have some intuitive understanding of how motions work and how you can explore the world around you. So even if you put a new object in front of the robot, the chances are if you ask it to grasp it, it will either grasp it or it will fail in an interesting way.
51:32So the probability of you actually doing something useful starts to become very, very high, or at least much higher than it was before. So now you have a way to explore the world, and that allows you to apply reinforcement learning methods much more easily because they can encounter successes much quicker. So this is kind of how reinforcement learning has evolved over time. And that's what allowed us to actually kind of start applying it again and seeing if we can see the successes of it. Now, from the other perspective, the reason why we thought reinforcement learning would be important is that the way most of these models are trained today is mostly based on imitation.
52:10So you collect a lot of data, whether it's in simulation or from human videos or from real teleoperation data. And then the objective for these models is try to replicate the actions that you've seen in the data set. And the important caveat there is that it's not actually the objective that we care about. We don't care about executing every single action exactly the same as it was presented to you. What you really care about is the success of the task, whether you accomplish the task or not. And we don't really have a way of codifying this objective with these models. They're all optimized for imitation, and therefore it's very difficult to drive the success rate of those models because it's not the objective that they care about.
52:53They only care about imitating the actions as closely as possible. And reinforcement learning provides you a way to basically codify this objective, to have them now be optimizing for actually accomplishing the task successfully. Both of these things combined kind of led us to believe that we should start applying reinforcement learning so that we can improve reliability of these methods and have models that work, you know, not 70 % of the time, but 99.99 % of the time because they actually optimized for the right objective. And at the same time, it became possible because now these models can explore in ways that are more intelligent than it was possible before.
53:32To circle back to a topic you alluded to earlier, you mentioned that when you first started chatting with Lockheed, he understood the need for patience on commercialization. and you sort of referenced the fact that, you know, there's a version of this company that if you're too impatient, maybe you can short circuit the sort of long-term vision. What are the ways that that happens and how do you have to be alert to them? I think there's a long history of robotics companies doing this or just happening to them. We're not the first company that starts with a big vision and broad vision for robotics.
54:09But if I analyze what happened in the past, what often happens is you start with that vision, you start developing the technology, and then because there's usually some external pressure, you try to commercialize this technology at that point when it's not fully ready yet. And you dive into a particular application, maybe the application that has the biggest TAM or something like that. As soon as you do that, you start cutting corners on the technology itself. And you short-circuit that vision of it being a very generalist, general-purpose technology that would actually deliver the most value, but that you're just trying to deliver the most value for that customer or for that particular vertical that you chose.
54:49And that's perfectly reasonable, right? Like all of your incentives are tied to that. Your revenue is tied to that. The customer satisfaction is tied to that. Your valuation, you know, how employees think about this. So you have all incentives in the world to start cutting corners and trying to make a more special-purpose solution. that's usually what happens so you start with this you know broad vision of how general purpose robots are going to be solved but you very quickly end up becoming you know an application company like a warehouse pick and place company and this would be a heartbreaking outcome for me i think we really have a chance to to solve the the big problem the problem of general physical intelligence and if we do that you won't just enable that one application that you could have focused on initially, but I believe you'd be able to solve all of it.
55:39You'd be able to solve it for any robot to do any task. And counterintuitively, that would be the much, much better commercial outcome as well. It's all just about trying to have the right time span in your head. How does that influence how you think about the right partners for the business? Like, are you optimizing for, I don't know, a range of use cases in that case so that you're getting the diversity of data? Is it about scale and, you know, getting as much of that data from, you know, in-person deployments as you, not in person, but real-world deployments as you can? Right now we're optimizing for the rate of learning.
56:19So it's still quite early, you know, unlike maybe what you can see on Twitter today, it's not that, you know, robots are about to knock at your door and show up at your home and do everything. Yes. I don't think everything is solved yet and it's just a matter of scaling up the existing recipes. I think we can push them much, much further than where there are today. But there's still a lot of research to be done and a lot of questions that are unanswered. So right now we're mostly optimizing for speed and how we can learn as much about the problem as possible. A lot of it does come down to what kind of data to collect, how to integrate the data into the model, and how to have that loop be as tightly closed as possible.
56:59It's a long way of saying that it's a fairly nuanced question. But what we're optimizing for is learn as much as possible so that we can figure out the scalable recipe that then we can just scale as much as possible. Do you have theses at this point on what gives you the sort of steepest rate of learning, whether that's, you know, I don't know, very fine-tuned tasks or something totally different? I think there is a few of those things that we learned. We know that the diversity of data is really, really important. We know that the quality of data is really important. We know that closing the loop with the models are very important.
57:38I just feel like these are fairly broad statements. That, you know, if you had asked me this a year ago, I would probably say something similar. But we did learn a lot about what these terms actually mean. I think, you know, very often people think about the diversity of data or quality of data or something like that. but we don't really have very good definitions of what quality of data actually is what does it mean or what diversity of data actually means or how do you measure it so i think these are like fairly deep questions and we are um we are getting more and more of a of a feel of what it actually is and how we can measure it and how we can optimize for it but it's a constant push and pull on how much do you scale the existing thing?
58:24How much can you better understand that? How much can you increase the slope of improvement there? So I think we'll just continue on the trajectory as we scale it. You mentioned that, you know, there's still a need for real research in this space to achieve the ultimate end goal, but that there's also plenty to optimize. You know, in this sort of thought experiment where we get no new research, how far do you think that that sort of takes us? Do you have some rough heuristics of what kind of a robot that is able to produce? It's a question that we think about a lot. Because I think one powerful insight from the language model work is that for a long time, people thought that, you know, we still have a few ideas that are missing, or we need a different architecture, or we need different this or that.
59:12But it turned out that it was good enough. It was just a matter of scaling. and we just didn't have enough foresight to predict how scaling is going to resolve some of the issues that seem like big issues of today at small scale. So I think there is a non-truelled chance that the recipe is already there, that the recipe that we have today would work and would just solve everything and we just need to scale it in the right way and execute on it really, really well. Well, I'm not certain about this at all, but I think that's one question that is an important question that we're thinking about quite a bit.
59:47So what we're trying to do is verify it, so start scaling the existing recipe and see how it performs, how it scales, and at the same time continue to do research that at the very least could improve the slope of that scale. NVIDIA is probably the best example of a company that is putting a ton of money into improving its sort of simulation engines and getting that to higher and higher fidelity. How optimistic are you about that improving at a sufficiently high rate that it really, I don't know, closes the gap in some meaningful way or allows you to get the kind of data at scale that could improve things meaningfully?
1:00:28So as I mentioned before, I'm quite open minded about these things. I'm not very dogmatic about the path that's going to take to get there. I think simulation over the past decade plus has been improving like crazy. The simulation that we see today is much more realistic than what we've seen in the past, and it's becoming more and more scalable. I believe that the first place where we're going to see the impact of it is not going to be as much data collection as it's going to be evaluation. And one thing we see here at Physical Intelligence is that as these models become more powerful, it takes more and more time to evaluate them.
1:01:04so now not only you need to evaluate them for longer you need to take more samples to distinguish between the model being you know at 99 success rate versus 95 rate the success rate than it was between you know 50 versus 70 you just need more samples to be statistically significant but the repertoire becomes broader and broader so now you need to evaluate them on more tasks across more robots in more environments and that trend i think will continue as the models get get stronger, you will need to evaluate them for much longer in much more sophisticated ways. So I think if simulation starts to work in the way that we all hope it will, I think that's where we'll see the first impact of it, where we can start evaluating these models in simulation across diverse set of scenarios.
1:01:53And if they're realistic enough, that should correlate with real world evaluations. You know, we talked about some of the real inflection points for you over the course of your career. I'm sure that physical intelligence, biggest inflection points are to come. But over the past couple of years so far, what have been maybe the biggest surprises or the moments that have impressed you the most? I think the main thing is just how fast all of this has been moving and how much it's picking up speed as we go. So before we started the company, we thought that it was going to take us something like five years maybe to start deploying robots and we deployed our first robots i think 18 months and into the company so it either means that i'm very bad at making those predictions or um it's been just moving much much faster there's been multiple of these moments i think um our first model our first release by zero was definitely one of those moments where i thought it would take a very long time before we have robots folding laundry especially diverse pieces of laundry fully autonomously.
1:02:58This was one of these holy grail of robotics that I thought it's going to take many, many years before we get to. And we got there within the first seven months of the company. I think the next model was another one of these moments where I thought it was going to take very, very long to get robots to perform in an environment that they've never seen before. And with our Pi-05 release, we got robots to new homes that they've never seen before. and home is the most diverse environment you can imagine. So it's the hardest version of that challenge. And that started working as well. Then with PyStar06, we started seeing that these robots can perform tasks at a very high clip, at very high success rates.
1:03:38And now we have videos of robots doing extremely dexterous tasks, like making coffee on the proper espresso machines, a machine for like 13 hours without any cuts, just like 13 hours straight the robot is doing coffee over and over again and cleaning after itself so i think all of this is just moving very very fast and every single release so far has been quite surprising to me i thought it was going to take longer or that the problem is much more difficult than i than i expected but i think that's maybe like one thing about working with these models i think probably people working on llms would have similar answers that they thought it was going to take much longer to get to a certain level of capability you know when you're running a robot for 13 hours straight making espresso what does a good failure rate look like in that in that case or what are you excited by when you when you see that happening if anything and then you know something i was talking about with with a friend who's in this industry is that a lot of the times the failure for uh these kinds of robots is maybe less obvious and more like it almost doesn't realize it's stuck in some way and doesn't find a way out of being stuck in that loop.
1:04:49Is that sort of something that you see a fair amount in this case? I think in general, people don't pay enough attention to this problem of reliability. And that's maybe a big difference between something like a language model or a chatbot and a robot. When you talk to a chatbot and it makes a small mistake or it just like puts not quite the right word, it's totally fine because you're right there and you can absorb all of this, you can correct it you can ask it again or you can interpret it the right way you're very forgiving and the physical world is the opposite of that that's not very forgiving if you make any small mistake if you grab the portafilter just a little bit off or you don't insert it just quite right then the whole task is going to fail in a catastrophic way I think the bar on performance is just much much higher because there isn't someone to help you, someone to catch you.
1:05:43And I think this is a very big problem. Now, the other thing you point out is more about the model being self-aware of where it's not going well. And the power of reinforcement learning is that it very often needs to have an additional function that tries to predict how close you are to success. We call this a value function. And that value function should have an understanding of how well you're doing so that then you can improve upon it. And what we found so far is that you can train these value functions to be pretty good. You can train them to the extent where we don't fully know that the robot is about to fail yet, that everything is looking okay, but the value function already knows.
1:06:25And they start seeing that the expectation of success is starting to go down. This is actually something that we've also seen with words like AlphaGo, where we had a computer program learn how to play Go, where if you watch some of the games that this AI system had against the Lisa Dahl, the world champion at Go, what was quite intriguing is that in many of these games, you have world's experts at Go talking about and commentating on the game. They're all thinking that the game is basically head-to-head. It's very unclear who is winning, whether it's the AI or Lisa Dahl. But then you look at the value function predictions, and it's basically over.
1:07:07Like the AI system knows 100 % that the game is over, it's already won by AI and Lee Sedol has no chance. But the entire world, you know, all the experts, including probably Lee Sedol himself, think that this is a very, very close game. So I think we will probably get to similar models in robots where they will be able to predict how well they're doing or how likely the success is much better than we will. One of my favorite activities to do with my son is he loves to look at my bookcase and just pull out whatever covers he finds interesting. He's, you know, 15 months old, so he's just doing it based on, you know, what interesting faces on the front or whatever.
1:07:46And he's gotten really interested in Oliver Sacks's biography called On the Move, because I don't know, it's a picture of Oliver Sacks. It's, you know, engaging. And I was just paging through it with him like last week, and I was already starting to think quite a lot about our conversation. And it lands on this page where Oliver Sacks talks about being obsessed with proprioception and how this is, you know, the sixth sense. And it's actually the most important sense because, you know, you can strip humans of all the other five and they can more or less get along. But if you, you know, get rid of this sense of where your body is, life becomes extremely hard.
1:08:26And he talks about working with these patients who, you know, maybe have a virus and they lose this sense of their own body. And he has this one case, which made me think a lot about what you're doing and maybe what the final state of it might look like, which is this guy gets a virus and he loses the sense of proprioception, but he compensates with visuals where, you know, if the lights go off in a room, he can no longer move. But as long as he can see where he is, he can sort of walk, but not walk and talk, et cetera, et cetera. And so it's basically the way that the brain compensates for the loss of this sense.
1:09:08Is that sort of the stage we're at at the moment with physical intelligence where we don't have true proprioception? And so we're sort of compensating with these other senses. and is there a point in the future, do you think, where however a computer is able to emulate that, we have something that really feels equivalent to it? Yeah, there are many neuroscientific experiments showing something like this, how you can disable one sensing modality and then compensate for it with another one. I think in general it's quite unclear what the right sensing modalities should be. I think our intuitions are often wrong about us because also these machine learning algorithms can extract the right level of information from signals that to us seem like insufficient.
1:09:59Like a common question I get is, you know, about touch sensing on our robots. And instead we put wrist cameras and it turns out that they can compensate for the lack of touch just fine. And they probably find some kind of regularities in the data where maybe you see the deformation on the gripper or something like this that indicates to you how much pressure you're applying. and that's enough for it to fully compensate for the sense of touch. It might be that we'll find some paths where that's not the case and you really need some other sensing modality to give you that signal. But so far, I think it's been quite surprising to the entire field how far you can push it with just simple cameras.
1:10:40I think in general, though, a lot about intelligence is about compensating. And what I mean by this is kind of like predicting what's going to happen and then trying to figure out if what actually happened, how different it is to what you predicted. And it doesn't just apply to sensing modalities. I think it also applies to precision. So for instance, as a human, you're very imprecise when it comes to movements, right? Like if I ask you to put your finger on a specific point a hundred times in a row, and I measure this exactly, there will be quite a bit of variance, especially compared to modern industrial robots that can go to sub-millimeter accuracy and very good repeatability.
1:11:24But if you compare our dexterity to dexterity of those robots you know it's night and day. We're way more dexterous even though we're way less precise of a machine. And I think the reason for this is that we have physical intelligence that can compensate for it. We have enough feedback that if we just look where our fingers are or if we just feel it we can compensate for it immediately. And I think we'll see that being translated to robot design as well, where maybe the robots of the future don't need to be as precise as we had thought, or, you know, they can accommodate for a lot of things, like a little bit of backlash in the motors, or just noise, because with the right intelligence, you can compensate for it.
1:12:06And that's what we've seen here at Physical Intelligence. We show some of the most impressive demos ever on robots. But these robots are really bad robots. If you compare them to any state-of-the-art industrial machine, they're really, really bad. They're very imprecise, not very reliable, a lot of backlash and other problems, but it doesn't really matter because the right intelligence can compensate for it. Well, that was a very interesting and gracious way of answering a very baggy question from me, so I'm glad I asked it because that was fascinating. As a final wrap-up question, I always like to ask folks that if they had the chance to assign a book to everyone on earth to read and know that they would understand it, what is a book you would love to give everyone?
1:12:53I reread recently a book that I'm a big fan of. I'm not sure if this is the book I would recommend to everyone, but this is the book I think that is really, really good. It's Why Greatness Cannot Be Planned by Ken Stanley. Yes. Yeah, I think it's just a wonderful counterintuitive book that shows multiple examples of how we arrive at something really spectacular without planning for it. I find it quite inspiring and quite motivational in a way that is maybe a little less straightforward than usual. So I think it's just like a very insightful book that would help a lot of people to show, shed a new light on how to think about innovation and achieving something spectacular.
1:13:41Well, that's a perfect place to end. Thank you so much, Carl. Thank you. That's it. Thank you for listening to this episode of The Generalist Podcast. Please subscribe on Apple Podcasts, Spotify, or your preferred podcast app. Ratings and reviews help others discover these discussions, so if you enjoyed the conversation, I'd be grateful if you could take a moment to leave one. For all past episodes and more, visit us at thegeneralist.substack.com. See you next time as we continue to explore the future.
1:14:16Thank you.
From the publisher
Karol Hausman is the co-founder and CEO of Physical Intelligence, a robotics company building a general-purpose “AI brain for the physical world.” The company has raised more than $1 billion in funding to develop foundation models that allow robots to operate across many machines, environments, and tasks rather than being programmed for a single purpose. The core thesis: the same scaling dynamics that transformed language models may also unlock robotic intelligence. But only if you resist every commercial pressure pushing you toward specialization. The central challenge isn’t mechanical design. It’s intelligence: how robots learn, generalize, and interact with a physical world that is far harder to simulate than it is to describe. Before launching Physical Intelligence, Karol worked at Google Brain and Stanford University, studying robot learning alongside researchers Sergey Levine and Chelsea Finn, who later became his co-founders.
In our conversation, we explore:
- How growing up in a small town in Poland and watching Star Wars sparked Karol’s fascination with robots
- The moment a lecture from Sergey Levine convinced him to abandon his PhD research direction and pivot fully to deep learning
- Why robotics has historically lagged behind breakthroughs in language models
- The case for building a general “AI brain” for the physical world rather than a single specialized robot
- The role of real-world data in training robots, the limits of simulation, and how deployment could create a powerful data flywheel
- The return of reinforcement learning and the parallels between human learning and robot training
- The unique challenges of physical intelligence and why robots must operate with far higher reliability than language models
—
Thank you to the partners who make this possible
Brex: The intelligent finance platform.
Granola: The app that might actually make you love meetings.
—
Transcript: https://www.generalist.com/p/karol-hausman-physical-intelligence
—
Timestamps
(00:00) Intro
(04:05) Karol’s early fascination with robots
(07:38) How Karol relates to Fei-Fei Li’s biography
(08:52) What inspired Karol to build better robots
(11:19) Philosophical influences
(15:33) Parallels between The Inner Game of Tennis and robotics
(18:21) Karol’s entry point to robotics and PhD program
(25:49) Combining robotics with LLMs: The Taylor Swift demo
(30:48) The 1970s SHRDLU AI experiment
(32:33) Founding Physical Intelligence
(35:13) How Lachy Groom got involved
(39:40) How research shapes what Physical Intelligence builds
(45:22) The importance of real-world data
(49:07) The return of reinforcement learning in robotics
(53:31) The risk of commercializing too early
(55:47) Finding the right partners for the business
(57:13) Open research questions
(1:00:00) NVIDIA’s simulation engines
(1:01:57) The surprising speed of progress
(1:04:16) Reliability in robotics
(1:07:31) Compensating for missing senses
(1:12:28) Book recommendation
—
Follow Karol Hausman
LinkedIn: https://www.linkedin.com/in/karolhausman
—
Resources and episode mentions: https://www.generalist.com/p/karol-hausman-physical-intelligence
—
Production and marketing by penname.co. For inquiries about sponsoring the podcast, email jordan@penname.co.




