Sergey Levine: Current State of Humanoid Robotics, China & Future Predictions

24 Aug 2026 · 58 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Sergey Levine discusses the current state of humanoid robotics, why progress is hard to predict, what “emergent capabilities” look like, and likely deployment timelines (single-digit years). He argues robotics is in a “fundamental technologies” phase (not predictable GPT-4-style scaling), and that reliability/robustness and generalization are the main bottlenecks.

Guest backgrounds

Sergey Levine is a leading robotics researcher and co-founder/leader at Physical Intelligence (implied by discussion of their demos and projects).

Key claims

Scaling works when done at the right thing and with the right data; robotics needs a positive deployment/data flywheel but data is heterogeneous. Generalization is the core challenge, and robots must improve via autonomous experience and common-sense recovery. Robotics safety is real but less of a “hard stop” than in cars because tasks/domains can be constrained.

Notable examples

folding laundry with two shirts disentangled; kitchen cleaning where the robot opens the oven instead of the drawer; washing plates with memory errors (dropping one plate). Long-horizon espresso-making for 13 hours; box assembly at Dandelion Chocolate Factory. Waymo autonomous driving as an analogy for real-world landing. Skill transfer across robots via intermediate “thinking” in images (e.g., UR5 folding a t-shirt without t-shirt data).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Current State of Humanoid Robotics

0:40 to 2:06

Discussion on the current advancements and capabilities in humanoid robotics.

“LOMs have caused unprecedented impact and investment in that area.”

Developing Scalable Technologies

2:06 to 4:00

Exploration of how scalable technologies are essential for robotics development.

“So figure out the design, figure out, like, roughly the mixtures.”

Evaluating Robotics Performance

4:00 to 6:16

Insights into specific evaluations and emergent capabilities in robotics.

“I think a lot of the puzzle pieces are falling in place.”

Inspiration from Autonomous Driving

6:16 to 8:32

Comparisons between humanoid robotics and advancements in autonomous driving.

“Another experiment we had is washing all the plates.”

Challenges and Ecosystem Health

8:32 to 12:18

Discussion on the challenges faced by the robotics ecosystem and the importance of a healthy industry.

“in the real physical world, I think that's really inspiring.”

Future of Robotics and Scaling

12:18 to 14:00

Exploring the potential for rapid advancements in robotics through effective scaling.

“But I think that it's important to sort of fully embrace that this is like a holistic thing and not something where we can do just like one piece of it and like outsource everything else, basically.”

Data Flywheels in Robotics

14:00 to 19:00

Learn about the significance of diverse data sources for training robots effectively.

“The trick is that there are lots of ways that it can be done wrong.”

The Roadmap to Generalized Robotics

19:00 to 26:00

Explore the critical milestones needed for developing generalizable robotic systems.

“So I do whatever I do on the back end, in a lab, in whatever.”

The Role of Data in Training Robots

26:10 to 28:00

Understand the different types of data necessary for effective robot training and their implications.

“And I was reading there's different types of data.”

Robotic Foundation Models and Data Absorption

28:00 to 31:00

Learn how robotic foundation models leverage embodied data for improved learning and task performance.

“Maybe we should like start with like YouTube videos and then put robot data on top of that.”
Show all 20 chapters

Transferring Skills Across Robots

31:00 to 33:56

Discover how models can facilitate skill transfer between different robot types through innovative thinking techniques.

“When we started doing all this, I had a big long list of all the cool research I wanted to do to better accommodate different morphologies.”

Challenges and Future of Humanoid Robotics

33:56 to 37:30

Explore the reliability and robustness challenges facing humanoid robotics and what could impede progress.

“but with a twist that you have to think in the right modality.”

Foundation Models vs. Specialized Systems

37:30 to 40:12

Understand the debate between using generalized foundation models versus specialized robotic systems in various applications.

“Whereas with a robot, like the full value of it is unlocked when it's actually doing the thing autonomously.”

AI Safety and Societal Implications

40:12 to 42:00

Discuss the societal implications and safety concerns surrounding advanced AI and robotics technologies.

“And I think this is like kind of hard to tease out sometimes because obviously, like, you know, because that's the cool thing.”

AI Safety Concerns in Robotics

42:00 to 44:38

Explore the implications of advanced AI systems on safety and societal impact.

“kind of a traditional vertically integrated robotic system.”

Evolving Views on Robotics and AI

44:38 to 47:34

Discuss the evolving role of robots and AI in society and labor.

“Like, Like, you know, computers at some level are kind of like mechanical brains.”

Breakthroughs in Humanoid Robotics

47:34 to 49:56

Learn about key papers and breakthroughs shaping humanoid robotics.

“And the idea was that, hey, if you set up the right kind of low-cost robot setup, in his case it was based on these robot arms from Trosson Robotics.”

The Role of Control Systems in Robotics

49:56 to 53:48

Understand the importance of decision-making and control systems in robotic functions.

“and then in the last, I don't know how many years, I don't really hear about them much anymore.”

The Importance of Prior Knowledge in Robotics

53:48 to 56:01

Discover the significance of integrating prior knowledge into robotic learning.

“industry and give yourself some advice, knowing everything you know now, what would you say?”

The Importance of Prior Knowledge in Learning

56:01 to 56:38

Learn why prior knowledge is crucial for solving complex problems like assembling furniture.

“They do it from observation, from observing other people, other creatures, and so on.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00To me, that's kind of mind-blowing because like the base model wasn't trained on any human data at all.

0:04Sergey Levine:This is Sergey Levine, one of the world's leading robotics researchers, and I asked him all about the current state of humanoid robotics. What are the most astonishing emergent capabilities you've seen so far? I don't think anybody watching that eval thought that the robot was going to do this. I have some questions here on China. There's no avoiding it sometimes. If humanoid robotics did not succeed in single-digit years, What do you think would be the most likely reason why humanoid robotics failed? Here's the full episode.

0:40Sergey Levine:LOMs have caused unprecedented impact and investment in that area. Physical AI and humanoid robots could potentially be even bigger. And so I wanted to ask you today about where are we today with humanoid robotics? and how do you foresee this technology actually being deployed into the world? I guess with machine learning, what we've learned over the last few years, over the last decade rather, is that it works when you do it at scale. And this is like very obvious now, but it wasn't always obvious. But there's a caveat, which is you have to like scale the right thing. And, you know, initially when people started working on models for language, for example, the dominant design was LSTMs.

1:28Some people remember what those are. They were like kind of okay. They were a lot better than what came before that, but they didn't really scale as well. And then the big thing with transformers was not that transformers were somehow like particularly, you know, mathematically elegant or anything like that. It's just that they scaled better. So they were easier to train on very large amounts of data with lots of parameters. So the technology proceeds in phases. First, you figure out what you can scale. Basically, what is the scalable technology? And then you pour on a lot more of kind of an industrial scale effort, adding lots of data, adding to model size.

1:58And that's when kind of the magic happens. So when we're doing kind of more fundamental technology development, the key is to understand, like, what are those scalable levers? So figure out the design, figure out, like, roughly the mixtures. And that by itself does something pretty cool, but not like, but that's not the thing that actually changes the world. It's like when you start pulling that lever that things actually change. So with LLMs, when the first GPT models came out, with GPT-2, it did some stuff, but it was sort of like a parlor shook. So you could get it to synthesize a story about unicorns in Peru or something, and it was coherent English, but it wasn't a thing that would solve lots of real-world problems.

2:39But the folks that worked on this recognized that, hey, there's something magical that happens because as you add more data and you make the model bigger, the stuff gets more coherent and more effective. So they could see that if we do a lot more of that, then it'll become a lot more powerful. So to come back to your question, what I would say about robotics is that it's not in the like GPT-4, GPT-5 stage where it's like an industrial scale effort to kind of make the model bigger and get more capability out of it. It's in that stage where we're establishing the fundamental technologies. And because of that, what one should expect to see right now is not necessarily that each month the model gets bigger and more powerful by some predictable kind of scaling curve.

3:26It's that the scaling properties themselves are evolving as we develop the right technologies. So to bring this back to something closer to reality, I'm very happy with the demos that we're doing here at Physical Intelligence. And I think that a lot of the results that other people are coming out with are really cool. But these are, to put them in context, we should not expect these to be the things that are actually illustrating the power of scale. We should expect them to be developing the fundamental technologies that will be scaled up after that. So where I think we're at now is that we're actually kind of getting all those puzzle pieces in place.

4:00And I think it's actually very close. I think a lot of the puzzle pieces are falling in place. But it's not like what makes it so hard to prognosticate about where the technology is going to go is that it's not yet at that predictable scaling stage. It's at the stage where we're figuring out the puzzle pieces, which I think is really exciting. But it means that it's also very, very hard to foresee, like, you know, sort of what the coefficients on that will be.

4:20Sergey Levine:What are the most astonishing emergent capabilities you've seen so far? That is the thing that is the most fun. And certainly we've seen a lot more of that happening as we progress. In the very beginning, it was little things, but they were magical because in robotics, basically, prior to 2024, the stuff never happened. The little things that happened, and this was maybe at this point, about two years back, we would see things like, okay, we train our policy for folding laundry, and it takes out individual shirts out of the hamper and tries to fold them. And then one very vivid memory I have in late 2024 is we were watching one of the evals, and it takes out like two shirts at the same time.

5:05And I'm watching this, and I'm like, OK, it's done for. Like, there's no way it can possibly do this. And then it puts the two shirts on the table, disentangles, and then puts one of them back, and then starts folding the other one. And it's like, wow, like that. OK, like in retrospect, you can do some like detective work and figure out where it got that from some piece of training data. But that was like one of those moments where I don't think anybody watching that eval thought that like the robot was gonna do this It's like a little thing.

5:30Sergey Levine:It's like exhibiting the common sense you expect people to have but actually the thing that I find more interesting recently is some of the mistakes because one of the things that's That was pretty remarkable about LLMs is that once they got good enough Even the mistakes kind of made sense in the sense that they weren't like crazy mistakes were just outputs like ZZZ all the time, but there are mistakes that sort of semantically are sensible. We had an evaluation last year for Pi05 where the robot was cleaning up a kitchen, and it's told, like, put away all the utensils. Like, there's, like, some spoons, spatulas, et cetera.

6:06And it tries to open the drawer where it thinks the solar goes, and it can't really get the drawer open. So it slides over and opens the oven, which is right next to it, and then starts putting this stuff in the oven. It's like, you know, you can sort of imagine that if you ask, like, a child to clean stuff up, and put it away, they might decide to do that because, OK, it's like a container and you can put stuff there and nobody sees it. Another experiment we had is washing all the plates. So it would pick up the plates, wash them with a sponge, and put them on the drying rack. And this was an experiment on memory because it has to keep track of everything that it's doing.

6:36It has a scratch pad kind of memory where it's writing down, like, hey, I had three plates. I cleaned the gray one. I cleaned the green one. And then it drops one of them on the floor and drives the base over so you can't see. And it's like, OK, I've cleaned the gray plate. it's done. So I mean, obviously, like these are not the things that we want to see. But it's kind of interesting that some of the mistakes, they're almost like what you would associate with like a child trying to do the task.

7:00Sergey Levine:So now it just needs to grow up. You mentioned the advancements, they're kind of happening all over the industry. And I know there's a lot of people building humanoid robots. There's Figure, there's Tesla, and there's many other competitors. And you have the expertise of what is hard and what isn't. When you look at the competitors, has there been any advancement or achievement where you think, oh, that's really admirable and that's impressive? I actually think that one of the most inspiring things to me in the industry is to see the kind of takeoff that autonomous driving systems have had. Because one of the criticisms that is sometimes leveled against robotics researchers is it's like, same as nuclear fusion, Like it's the technology of the future, but it's always in the future.

7:50But that's what people said about autonomous driving too. And now like, you know, we're in San Francisco, you can go outside, you can take a Waymo and it will actually like take you to your destination and there was no driver sitting there. So like, you know, without getting too much into the technical details, I think what's really inspiring about that is just like this case in point that yes, you can actually have one of these technologies of the future and it actually does land. And I think it's not an accident that it's landing now in the mid 2020s because a lot of the puzzle pieces for large-scale ML are getting to the level where we can put them together with actual physical systems.

8:24And there's a lot of differences between driving and robotic manipulation, of course, but I think that the illustration that, yeah, we can actually land learning-based technologies in the real physical world, I think that's really inspiring.

8:36Sergey Levine:If OpenAI or Anthropic started investing more heavily into robotics, how do you think that would impact the industry? Do you think competitors would be worried about that? Robotics is an area where maybe to be fair, like the ecosystem hasn't been as healthy as it has in other areas of machine learning. And what I mean by that is that like computer vision and NLP are things that sort of lend themselves naturally to a machine learning based ecosystem because there's freely available data. People kind of have a general acceptance that they're going to be using learning. You know, there aren't concerns, as severe concerns about like safety, at least physical safety, right?

9:20You know, people rightfully are concerned about AI safety, but it's not the same as like a physical device causing some physical harm. And because of that, I think it's a bit easier to spin up a very serious large-scale ML effort in those areas. Robotics is not like that. Robotics traditionally is not a discipline that, you know, really embraces sharing of data, for example. So I think that the more activity there is around learning in robotics, the more I think it'll shift people's thinking towards this kind of future where we accept that robots will be controlled by learned models, not by hand-designed controllers, that there will be data, that data will need to be shared because there's no way that somebody can build out a true foundation model in a single vertical.

10:04And basically kind of like shift the entire thinking around robotics to look more like how we think about vision and NLP as opposed to traditional factory automation. So I think in that sense, much as I'm proud of the work that we're doing at physical intelligence, I think it'll take more than one company to kind of shift everyone's thinking in that direction.

10:22Sergey Levine:As a bystander, I see on Twitter these really impressive demonstrations of Chinese robotics. It almost feels like they're ahead in some sense, but I don't have that deep domain expertise. And so I was curious your thoughts on what do you think of China's robotics and are they further along? I think that one thing that is very useful and constructive for us to do, those of us that work on these things in the United States and in Europe, is to ask what is the lesson to learn? And to me, one lesson is that it's important to have a healthy ecosystem. And ecosystem means that there should be, you know, obviously good researchers, good engineers working on these things.

11:16There should be like healthy open source. But it also means that the different industries that contribute to robotics need to individually be very healthy. And those industries are not just the computer science, ML, and model building stuff. It's also supply chains, manufacturing, hardware, R &D. These are all very important things. And aspects of those things are things that the United States does quite well. Other aspects of them are things where the United States has sort of let things go a little bit. And I think that what we should do is we should look at what's going on in the world, look at some of the excellent results that Chinese labs are doing, that labs in other countries are doing.

12:03And we should take away that lesson that we should strive to build a healthier ecosystem. And that means investing in all the different facets that contribute to this. So I'm not much of a business person. I'm not much of an investment person. So I can't claim to know how to do this. But I think that it's important to sort of fully embrace that this is like a holistic thing and not something where we can do just like one piece of it and like outsource everything else, basically.

12:27Sergey Levine:Of all that, the pieces of the ecosystem, are there parts of it as a robotics researcher in the U.S.? If it was better, that would have the biggest impact on the advancement of robotics? Yeah, I think certainly availability of reliable, low-cost hardware is a big deal. And right now, I mean, certainly for hardware used for robotics research, A lot of that does come from China, and it's good. It's relatively inexpensive. It's of high quality and meets the standards that people generally need. But it would be awfully nice to be able to source all that domestically as well. And I don't think there's anything impossible about that.

13:07I think it's just a matter of sort of embracing the fact that that entire ecosystem needs to be supported rather than just one piece of it.

13:14Sergey Levine:You can imagine that the lab that is first to get to scale will be the first to break out. There may be this exponential growth effect where once you deploy, deployment helps you grow faster and growing faster helps you deploy and this flywheel. And so do you believe that that will happen in this industry where one lab, whether in US or China, will hit some breakout point and they'll kind of jump everyone else? There is a lot of truth to that. I think that there's an important detail to keep in mind. The detail is that you kind of have to scale the right thing. So I think that basically the statement that, like the way that I would phrase this is having an effective positive feedback loop where more deployed robots translates to more model capability, that's kind of the key.

14:14And that makes total sense. The trick is that there are lots of ways that it can be done wrong. So I'll tell you like a few obvious ones. One obvious one is, let's say that I'm a car company and I have a robotic arm that is welding cars. And it's like there on the assembly line and it welds cars every day and it gets like a million welds every month and so on. If I just use that as my data flywheel, I'm unlikely to get something more capable than a robot that welds cars. So since the robot is already there and it's already welding cars, presumably the marginal improvement for that is not all that valuable.

14:55So that's an example of how you kind of have to scale the right thing. Because data is not quite as fungible. It's not like electricity or oil. You can't just buy more of it. It has to be heterogeneous. So I think that basically it's right, but you have to scale the right thing with the right technology and the right kind of source of diverse learning. Data is more like an education program for your robot than it is a fungible commodity.

15:19Sergey Levine:If you think about that data flywheel, and I guess putting out the proper platform that would generate data that matters that would then improve that platform, when do you foresee this kind of happening in the world? I think on the technology side, like things are advancing very rapidly towards that. And I think that maybe a particular balancing act to strike that would significantly determine that timeline is kind of how much structure somebody is willing to admit. So, you know, on the one extreme, you could imagine like jumping straight to fully unstructured deployment domains like home robots.

15:55And that could be really exciting because then you have a lot of diversity like right off the bat, but the bar is a lot higher to be effective enough in that domain, to be safe enough. Safety is a much bigger issue there because it's around people in their home. On the other extreme, you could imagine much more structured tasks, maybe not quite the welding robot, but sort of like, you know, maybe the robot down the hall that does something more unstructured, right? And that could be a much easier domain in the sense that there's much less safety concerns because it might be around trained humans.

16:25The task might be more predictable and so on. But there is the marginal value of each bit of data you get in that domain is lower because there's less variety. So it's like you could take off earlier, but with a smaller slope or later, but with a larger slope. And you kind of have to like calibrate that. But my sense is that whichever end of that extreme we're talking about, it's in the single digit years rather than the double digit years at this point. So it could be that the more structured ones might be happening like now or next year, the less structured ones might be a few more years out, but probably not like a decade.

16:59Sergey Levine:A lot of people talk about timelines, and it's kind of more nebulous. And I'm wondering if you were to put down a, I guess, somewhat of a roadmap, but not so concrete, but just milestones towards that North Star of robot in my home doing? Yeah, yeah, that's a really good question. So maybe to preface this answer, I think that there's one thing that is very important to say about robotics that is very easy to miss from kind of like, I guess, like the hype cycle and the demos that people put out, which is that the hard thing in robotics was always generalization. But when somebody shows a demonstration of their system, you know, the demonstration alone usually doesn't make it clear what level of generalization is being shown.

17:48The highly acrobatic robot demos, for example, are really exciting to look at. But typically, if it's something that is like a little bit more staged, like sometimes it's literally on stage if it's like a show. Like, you know, that is obviously like rehearsed, and that's okay because it's like literally just a show, but it's not the same as doing a task every time reliably in any home. And the generalization piece often doesn't look that impressive when viewed in isolation because generalization is sort of a property of many trials, not of one trial. So you might see the robot doing something fairly mundane and unimpressive, But what's exciting about it is that it's doing it with an object that it's never seen before in an environment that has never been tested in before.

18:31And that's actually harder than like doing an acrobatic backflip that it's practiced like millions of times. So with that said, my answer to your question is that the roadmap is all about both achieving better generalization and kind of that second order effect of having a mechanism to get more generalization as you generalize. So one of the steps on that roadmap is to have a very concrete demonstration of a robotic system that gets better with autonomous experience that is collected in a setting that it wasn't originally trained for. So I do whatever I do on the back end, in a lab, in whatever.

19:12I get my model. I get my adaptation algorithm. I put it in a new setting. Maybe it's a home. Maybe it's a factory. Whatever it is, something where it's doing something real. And it does okay, but then over time it gets better and better and It gets better and better to the point where it reaches Sort of practically relevant levels of robustness without sort of capping out at like 50 % like if that I think would be a major milestone Because now that that says okay if this is truly an automated process It's improving it's getting better even as it's collecting useful experience now I can take it and I can put it in lots of different domains Collect useful experience do something that people actually want and it'll improve the model So that I think is a really major step on that road.

19:54I think another really major step is to demonstrate a very concrete and practically useful way to transfer knowledge, to transfer common sense, to achieve robustness. So that's like the kind of the other scenario where you don't get to practice. If you're driving your car on the road and you see a fire truck and you see like a bunch of traffic cones, even if you've never been in that situation before, like your common sense tells you like, hey, I should like slow down. Maybe I don't know how to react optimally, but I shouldn't just barrel on through the cones and upset the firefighters and so on.

20:26So that's common sense. And if you can apply that common sense to effectively recover from unexpected situations, like when the robot put the spatula in the oven, it should probably open it up, take it out. It knows that this is not the thing you do semantically. That, I think, is another important step because that tells us that we can use common sense to fix mistakes.

20:45Sergey Levine:I could imagine a more narrowly scoped robot. Let's say it's a humanoid robot that's one step in assembly line in manufacturing. In my understanding, you're less interested in that because that's not really a step towards that general intelligence. What I would say, it's not that I'm less interested in it. It's that I think that the real world has these leaky abstractions that make that kind of stuff a lot more complex than it seems. Let me try to explain this with an analogy. So in the 90s, when people really started working kind of full steam on autonomous driving, there was this idea that we could avoid a lot of the hard problems by kind of instrumenting the environment a little bit.

21:31Like we'll have like magnetic sensors along the highway and so on. And cars will have like a little sensor on them and a little transmitter so they can tell where each of the cars are, kind of like the way you do it with aircraft, basically. And then people thought, well, we don't really need fancy AI. We'll just have these sensors and it'll just work. And that just basically didn't go anywhere because the real world has so many messy exceptions and special cases that even if 99 % of the time, the magnetic sensors and all that stuff just allow the car to drive, the 1 % when someone steps in the middle of the road or there's a piece of trash or whatever, it just messes everything up.

22:06So the thing that actually worked was when Waymo said, hey, we're going to not try to avoid the hard problem. We're going to actually try to deploy our cars not in like middle of nowhere, but in San Francisco, like very messy. And let's just deal with it head on. And that allowed that community to make progress. And I think robotic manipulation is going to be the same way that past the fully structured world of the factory, if you want to go even a little bit outside of that, even if 99 % of the time, it's all like pretty straightforward, that 1 % when something weird happens means that you really need the full scope of the problem basically to be addressed.

22:41Sergey Levine:This kind of reminds me, I remember Figure had this demo where they live streamed a robot kind of sorting packages. When you watch that demo, does it demonstrate generalization in your opinion? Yeah, yeah, I think it does. And by the way, this is something that I find very encouraging is that, like I mentioned that it's hard to show generalization in a video. And it's clear that lots of people are thinking about that and that are thinking about how do you like, Like, how do you present something that somebody can watch and take in kind of at a glance what generalization is? I mean, you can see a lot of creative steps towards that, live demos, these kind of like really long time lapses.

Read the full transcript

23:19I think it's a great idea. And I think that's like a really nice way to move towards elevating the importance of generalization in people's kind of consciousness. consciousness. You know, when we were working on the Pi Star 06 project, the RL project that we did late last year, we wanted to do some longer horizon experiments. In some cases, it's obvious, like we had our robot assembling boxes at Dandelion Chocolate Factory. So they're like, you know, it's an actual chocolate factory, so they need the boxes. So we run it for several days. But we had this coffee task, which was the robot was using an espresso machine to make espresso.

23:56So what we did is we ran it for 13 hours making espresso drinks. And we are, you know, we tried to be like very, I guess, environmentally conscious about this. So we didn't want to like throw out the coffee. So after 13 hours, everyone in the office was like a little wiry because someone has to like drink the coffee. But it like it ran for 13 hours and it was pretty cool. It screwed up a few times, like it'll spill all the coffee grounds and then it needs to go and get like a cloth and wipe it down. But like it does it and it's, you know, nothing exploded. 13 hours went by. Probably the most negative consequence was like loss of sleep from too much caffeination.

24:29Sergey Levine:But when it spills the coffee grounds and cleans it up, it did that by itself. Well, so the way that that experiment was done is that there is a high level prompting. So like roughly the prompt is updated maybe every like five minutes or so, like in between semantically coherent tasks. So you tell it like, make espresso, clean up the machine, et cetera. So those steps, the actual like clean up the machine was commanded by a person. In principle, we could automate that. In fact, one of the things we're spending a lot of effort now is improving our high-level policy that does those commands. But for that experiment, every five minutes, somebody basically updates what it's being asked to do.

25:03The way we intended it was the commands would be like, if you go to an actual coffee shop, you say, oh, I want a latte, I want an espresso. That was supposed to be the prompt, except then you also have to tell it, I want you to clean it up before you do the next one.

25:15Sergey Levine:That makes sense. OpenAI, Anthropic, Cursor, and Vercel all use this product to make their lives better. And the problem it solves is when you're building SaaS or an AI product and you wanna sell to other companies, there's all these requirements you need to meet. There's SSO, there's SCIM, there's RBAC, there's audit logs. These are all things that take time to integrate, but aren't the main focus of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. So let's say you have a new SaaS product and you wanna sell to other companies, Work OS will solve all of these critical feature gaps for you.

25:54Sergey Levine:You can check them out at workos.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. It sounds like on the way to generalizing, data is a very important part of that. And I was reading there's different types of data. There's simulated data. You could collect physical interactive data. and I wanted to hear your take on what's the best data to get, what's the worst, and what are the pros and cons. This is, by the way, like a question that is, I guess, quite... There's a lot of discussion in the robotics community about this question and some people have, like, very opposite opinions on it.

26:36My own take on this is that a lot of different data sources are easier for the model to internalize if it can ground them in a thorough physical understanding of the world. So let me try to explain what I mean with a few examples. If you want to learn to fly an airplane, you will probably use a simulator, at least during part of your training. But the simulator makes a lot of sense to you because when you start using the simulator, you have a lot of world knowledge that you can use to ground what's going on. Like, you know that when you are using the flight simulator to learn how to fly the airplane, You're not just like playing a video game.

27:15You're trying to acquire knowledge that you will then use with a real airplane. And you understand that there's sort of an abstraction there. Same thing if you're playing like a really cartoony, like, you know, Atari game or something, right? Like, you know that all the symbols on the screen, you can sort of connect them to things that you've experienced in your life, and you can make an analogy there. So a lot of that, like, even though it kind of seems like these simulated environments reflect aspects of the real world, To us, they make a lot of sense because we kind of bring to bear a lot of our own prior experience and we fill in the blanks that the simulation has.

27:52And also, if you want to use data yourself as a person of somebody else doing something, if you watch someone, let's say, cooking a meal, right, even though you don't experience every movement they're experiencing, you have a lot of that knowledge that you bring to bear. And you're like, okay, I see they're picking up the salt shaker like I've put salt on things before So I kind of roughly know what's going on there And I can file it away at this level abstraction of like add salt without having to like figure out all their muscle movements So my point with this is that once you have that Understanding of how you do things physically with your own body and how you experience the physical world now All these other sources of knowledge can be connected up to it because that Foundation you get from your experience helps you ground everything So where I'm going with this is that if we have a robotic foundation model that is trained on lots of real embodied data that provides that grounding, it might actually be much better able to absorb other sources of knowledge.

28:46And this is actually like a little bit upside down relative to how some people think about it because it's very tempting looking at the success of like internet data for LLMs to say, well, maybe we should do the opposite. Maybe we should like start with like YouTube videos and then put robot data on top of that. But I think it's actually the other way around. And I even have a little bit of evidence for this. So my colleague, Suraj Nair, together with Simar from Georgia Tech, they had a project together a while back where they took our robot foundation model and they added human video data. But they didn't start with human video data.

29:22They actually started with a model trained on robot data and then added video data on top of it. And what they did is they looked at the representations inside the model. Basically, how does the model represent human experience versus robot experience? And they found that if you use a small model with a small amount of robot data, predictably the human experience and the robot experience are fully separated, meaning that the feature representations are different. But if you train on lots of robot data from lots of different robots, then the features are grouped much more by what task is being done rather than by whether it's a human or a robot.

29:52And when we looked at the feature plots, it was just mind-boggling because literally when you crank up the amount of robot data to 100%, they just line up perfectly. Like you see the, you know, you do this T-SNE embedding, you see the shapes of the embeddings, and it's just all task identity and like minimal sensitivity to embodiment. And to me, that's kind of mind-blowing because like the base model wasn't trained on any human data at all. But once you start adding human data, it represents it exactly the same way. And I think that's really exciting. I think that to me is like one of the strongest indicators is that if you have that good foundation of robot experience, you can put everything else on top of it.

30:27It's actually better at absorbing that.

30:29Sergey Levine:Is it important that that base model has data that was collected using that specific set of motors, specific set of joints? So far, we've obviously put a lot of effort into cross-embodiment models that can handle many different robot types. But generally, you do need data of the robot you're going to be deploying on to get good performance. So kind of the metric of Generalization there is not can you zero shot a new robot? But it's mostly can you get away with less experience from the new robot and transfer skills from other robots? Okay, so that's the current state of things now There is a little bit of a kind of surprisingly positive read on that which is even though you need data from these robots the amount of special stuff that the model is doing is is kind of minimal.

31:16When we started doing all this, I had a big long list of all the cool research I wanted to do to better accommodate different morphologies. Can you factorize the model's representation in some way so that there's a six-degree freedom arm head and a seven-degree head and a gripper head, etc.? We didn't do any of that. The model just outputs a big vector of numbers. If the robot has less degrees of freedom than the number it outputs, it just zero-pads it. There's nothing fancy, and that's it. and then just train on all the robots and outputs the correct actions based on what is seen through the camera essentially.

31:49But now, to your point about whether you can handle new robots. So far, the thing that we focused on, and I think this is showing some promise, is to be able to transfer skills between robots. And this is actually where the particular choices in how the model works seem to matter. For example, you can have a model that does some intermediate thinking. And that thinking can be done in different modalities. So you can think in text. And thinking in text is really good for transferring high-level behavioral structure. So that's basically how you understand that, hey, if I want to clean the kitchen and put away the silverware, first open the drawer.

32:34That's kind of a semantic inference. You can transfer that very well because obviously that's largely agnostic to any embodiment or anything like that. But even lower level things can be transferred if you use the right representation. So one experiment we did is we had a thinking stage that is expressed in images, where you basically dream up an image of the next milestone in the task. And with that, we could actually get a robot, the UR5 robot, to fold a t-shirt, even though we didn't have any t-shirt folding data on the UR5. Because while getting the R motions correct is very hard, because the robot basically requires very different joint angles to do the task, cooking up an image of what it looks like for it to fold a shirt is not that hard because like you've seen the robot arm in all sorts of different poses, you've seen the shirt in all different stages of being folded and unfolded, you know roughly where it should hold it.

33:21So getting a good generative model to cook up that image is pretty straightforward. And once you have the image, then from that backing out the correct actions is easy too because you can just look at the synthesized arm angle and just like back out what the angle should be. So it's not changing the problem, but it's just introducing this intermediate step that makes it easier to solve. Just like if you're solving a math problem, if you figure out like the right intermediate step, kind of the answer is obvious from that intermediate step. And I think that's really exciting because now that shows that this level of generalization across robots, and I'm sure other generalization too, can be facilitated with thinking, just like in LLMs, but with a twist that you have to think in the right modality.

34:00Sergey Levine:Interesting. So it outputs a, I guess that image is what its video sensor is seeing, and it's like the next step. Yeah, you can almost think of it like image editing. You can do the same thing with video. You can do it with video prediction. But the key is to like imagine what it would look like to progress on this task. That seems like a very human thing to do, right? Yeah, yeah, absolutely. Some things you plan semantically and some things you plan spatially. Like if you're doing rock climbing, you're probably not thinking like, hey, left arm to rock 37 centimeters to the left. You're probably more like imagining your arm reaching for the rock.

34:34Sergey Levine:Earlier in the conversation, you mentioned that Waymo was very inspiring and their kind of path to productionization is proof that you can do real world generalized robotics. And if I recall correctly, when I was a lot younger, it was kind of this early promise of this is going to happen. And then in reality, it took a lot longer. And so I guess my question is, in the case of humanoid robotics, what would make you say single-digit years it's coming versus a long tail and policy challenges as well, I could imagine? I think one big difference between how robotic foundation models address the problem and how more traditional engineered systems address the problem is that the stack is really thin.

35:27So it's not easy to train a foundation model. You need to obviously get the right data. There's a lot of work that goes into curating, labeling, all that other kind of stuff. But the actual software that runs on the robot is very, very simple. So, you know, you might have some kind of thinking or reasoning stage. You might have the model produce actions. It needs to be fast enough. But the like just if you think about in terms of raw lines of code, it's much, much lower than a more traditional AV stack. And, you know, partly that's because modern autonomous vehicles, the work on that started a lot earlier with very different technologies and evolved over time.

36:10Partly it's also because the problem is more safety critical. Like, yes, you don't want a robotic manipulator to, like, drop a fragile object. But at the end of the day, that's a lot less bad than having a car hit somebody. So that is not to say that the safety challenges with robots are not real. They're very real and it's very important to tackle in the fact It's probably one of the harder ends of the problem, but they are not as much of a hard stop to practical deployments because you can Come up with tasks and environments and domains and also physical hardware where those problems are a lot less severe so I think that that combination radically simpler software stack plus less drastic software challenges, I think, actually make it a lot easier.

36:57And because, you know, to your earlier point, that there's this kind of flywheel effect, that there's a positive feedback loop that, you know, starting to get things out in the world, even under some constraints, will actually facilitate getting them out more and more.

37:09Sergey Levine:I think a lot of people are familiar with this idea of post-mortem, you know, looking back on why something failed. But in this case, I'm curious, what would you say to a pre-mortem, And in the sense of if humanoid robotics did not succeed in single-digit years, what do you think would be the most likely reason why humanoid robotics failed? Ultimately, for these things to be truly useful, they do need to reach a level of reliability and robustness and generalization that is higher than what we typically expect from LLMs, for example, or generative AI for images and video. Because typically like these tools they are very much human interactive tools Like you get an LLM to do something and it doesn't do quite what you want So you sort of revise your prompt and you basically like you iterate with it and and that's why like even the earlier LLM tools like the first version of chat GPT Even though they were much more primitive than what we have now They were still already useful because somebody could just like keep hammering out until it basically solves their problem You know just like if you're using a search engine like you type something in the search and you don't get quite what you want You revise your query and then you get what you want.

38:20Whereas with a robot, like the full value of it is unlocked when it's actually doing the thing autonomously. So it's having to have somebody like constantly iterate for every single task is almost like antithetical to the benefit that you're getting. So I think a lot of the risk has to do with how easy is it to get that level of reliability and robustness. And that's, again, where some of the demos might be like a little bit misleading because if someone shows a demo of the robot doing something cool, I mean, obviously, if everything is presented in a forthright way, that could still be a very good indicator of progress.

38:52But it doesn't make it obvious how far or how close it is to reaching that practically relevant level of robustness. So I'm personally a big believer in using techniques like reinforcement learning that can actually benefit from autonomous experience to kind of fine tune that last few percentage points to make it go from like 95 to actually 100%. But that's really important and it's not yet a solved problem.

39:18Sergey Levine:If it did take longer than expected, it's because the bar is higher. Because the bar is higher and in particular, like those last few, the kind of the last inch so to speak, is something that requires not just really good models, but also new innovations in technology. I mean, I don't think that, you know, I'm not the kind of person that would say like, oh, we should like throw out everything that we know about foundation models to start over. I don't think it's that at all. I think that roughly the puzzle pieces that we have are actually very good puzzle pieces. But still, we should acknowledge that right now, the methods and the models need more work to cross that level of robustness.

39:55Sergey Levine:In LLMs, it feels like everyone is doing kind of the same thing, but different flavors. In the robotics industry, is everyone doing kind of the same thing? Are there any hot take architectures that are different direction? I actually think that there's a lot more heterogeneity than it might seem. One big dividing line that I think is maybe not as obvious from just kind of looking at the results is the distinction between kind of fully embracing the foundation model ethos, so to speak, versus focusing on specific like kind of vertical areas. And I think this is like kind of hard to tease out sometimes because obviously, like, you know, because that's the cool thing.

40:40But the foundational model ethos fundamentally is something like this, that if you have a particular problem you want to solve, it is better to train a more general model that can use data from a breadth of problems. And if you do it right, it'll actually be better at the specialized problem you want to solve than a narrow specialist. So, again, to come back to the LM analogy, if you want to do machine translation, don't build a machine translation system. Build a language model that understands all language tasks and throw it at machine translation. And in robotics, I think that is actually very deeply uncomfortable to people.

41:14Because if someone is actually working on an application, like they're doing warehouse automation, it is very awkward to think, oh, if I want to do warehouse automation, let me collect data of putting away silverware in kitchens. It just sounds bizarre. But that is the foundation model lesson, that if you have enough breadth, if you collect data from a wide range of different tasks, Then you will acquire those generalizable skills, and if your model is built correctly, it will repurpose those skills for whatever situation it encounters. So I think it is actually true that even if you want to build a warehousing robot, you're better off collecting a breadth of data, and it will be better at handling all the weird edge cases you might encounter even in that warehouse domain.

41:55But this is not something that's easy for people to accept, because it's just so antithetical to the principle of building kind of a traditional vertically integrated robotic system.

42:05Sergey Levine:I noticed this new interesting phenomenon with these AI companies, which is if they're wildly successful, it creates this, I guess, worry or new set of things. So, for instance, when Anthropic had a very powerful model, then the government comes in and there's these worries about safety and risk and all that. And I'm curious how you think about that. Like if physical intelligence this year had a phenomenal, incredibly capable generalized model, how do you think about those kinds of topics that might come up? Working on AI safety is not a new thing. My colleague at UC Berkeley, Stuart Russell, was talking about this stuff like over a decade ago, and lots of people spent a lot of time working on it.

42:57It's just that the trouble is when the technology moves so fast, the important problems are not just a function of like, you know, the core principles. It's also a function of how society reacts to it, what kind of tools are adopted and so on. And I think that's very, very hard to anticipate. So I don't have like a very satisfying answer here. In terms of how we are approaching it our philosophy around all this stuff is Basically one of empirical experimentation like let's get stuff out there. Let's see what happens in the real world Let's see what goes right and what goes wrong so that we have as much of a preview for You know what the technology can do what are its weaknesses?

43:40What are its strengths? and and so on but You know at the end of the day you kind of have to just like keep your eyes open see what happens and adjust as you go. It's very hard to anticipate. And I think your question, though, is very spot on, even though I don't have a great answer for you. Because yeah, if we're having this much concern and issues with AI systems that are basically limited to using computers, we're presumably going to have strictly more concerns and issues with AI systems that can do everything in the physical world that we can do. So the issues are real. It's just, you know, you kind of have to like see what happens and then adjust.

44:18Sergey Levine:And that's kind of in the scary path. In the happy path, if everything goes well and we have incredibly capable models and robots, in 10 years, is the North Star that that's the end of human labor? I believe it's a mistake to think of robots as mechanical people, right? Like, Like, you know, computers at some level are kind of like mechanical brains. But when personal computers like really took off in the 90s, early 2000s, etc., it's not like the first thing that happened is that, you know, people replaced their brains with computers. Rather, what we saw is actually a proliferation of very different kinds of computers.

45:01We saw kind of like ubiquitous computing. So you would have a computer on your desk, but you might also have one in your pocket. You might have one in your refrigerator and in your car. Like because computing became so accessible, you could have a little bit of computing in everything. And, you know, I don't think that that's what like the people that first started thinking about this stuff in the 40s and 50s would have imagined. They would have imagined like, you know, room-sized computers whose job it is to like control the, you know, the policy of an entire country or something rather than like a little bit of computer in everybody's refrigerator.

45:32So I think it's, you know, by analogy that we might imagine there might be like a little bit of physical actuation in everything and it might just be like lots of everyday things that you have to do yourself. Now you get like a little bit of help with it. I think the other example that's worth thinking about is modern coding agents, right? So I think that, you know, this is something where, of course, the jury is still out as to what the end game of coding agents is. But certainly, from the experience of software engineers today, it kind of seems like probably fair to say that most would consider coding agents to be more empowering them rather than somehow causing them to have a panic.

46:14I mean, some people might have a panic. But in general, at least from the software engineers that I've talked to and from my own experience, it's more empowering to kind of be able to amplify how much work you can do with AI tools. So I think from that, and that maybe is like a pretty direct analogy because that is straight up an example of an actual real job where AI has entered into it and has actually provided more leverage to the people doing that job. So I think that's another example that we can look to. But the truth is that I think it remains to be seen.

46:44Sergey Levine:So in LLMs, there's a few seminal papers that if you read those papers, you kind of get a sense of the lineage of the breakthroughs that mattered and understanding where we are today. In the robotics industry, are there a set of top papers that you really think kind of show the breakthroughs that people should know about if they're curious about the state of the art in terms of humanoid robotics? One thing I would point out, and this is partly a shameless plug because I am a co-author on that paper, though candidly, 99.9 % of the work on this was done by Tony, who was the lead author, is the original ACT paper, the Aloha paper.

47:27It's an interesting example because in some ways, the ideas weren't really that new, but they were illustrated in a really nice way. And the idea was that, hey, if you set up the right kind of low-cost robot setup, in his case it was based on these robot arms from Trosson Robotics. They're like$7 ,000 hobbyist arms. He set them up in a bimanual setup with a leader follower teleoperation device. And he showed that actually if you do it right without really any particularly fancy tricks, you could easily collect teleoperation data of extremely dexterous tasks that people had previously thought would require like very sophisticated hardware and all sorts of like really expensive stuff, and then set up like a fairly straightforward transformer-based model, and it could actually do a lot of those tasks.

48:13And it's kind of like an interesting thing because usually in academic research, we put a big premium on like, you know, do you have some like sophisticated new mathematical thing or some sophisticated like technical insight? And in that paper, which I think at this point has been hugely influential, the insight is really just like, yeah, just put together the right pieces and have a little bit more faith in what a simple robot could do, so to speak, equipped with a good intent learning system. And he showed things like replacing batteries in a remote control. He even got like a little mannequin foot, and he showed that you could put a shoe on it for an assistive task, sort of.

48:51Some people need help getting their shoes on, so that's good. But what people found, I think, so interesting about that paper is just how far you could get with relatively simple building blocks. And at this point, he open-sourced the code for it, and the ACT code has been used by lots of people. If someone wants a very basic starter kit for robotic learning, that's usually what they grab. And I think that it's worth for somebody who wants to get into the field to go through that paper and really understand what's going on there, because even though in some ways it's not that sophisticated, I think it provides a bit of calibration on what matters.

49:30Like, you know, the details matter, but the details don't have to be complicated.

49:35Sergey Levine:Before all this large model robotics kind of wave, prior to that, Boston Dynamics had these really impressive demonstrations and tons of mindshare. I guess I wasn't even in the field, but I think, wow, they're really doing incredible robotics. and then in the last, I don't know how many years, I don't really hear about them much anymore. Is there some shift in the industry that made that so or, you know, is that something you could explain? So the way I would explain it is this, that there are, you know, robotics at some level is about building complex systems. So even though it's very tempting to say like, oh, there's different areas of AI, there's like LLMs and vision and robotics, one of those is not like the others.

50:30Because for robots, you actually need like all the parts, everything from like how you wire up the robot, what the power source is, what does the actuator look like, all the way to how does it do like high-level planning to determine like what tasks to do next. And even though we could look at these things and say, like, oh, all of these different videos and different companies and different demos, they're all robotics, they're really kind of different parts of the stack. A lot of what the classic Boston Dynamics results show is very sophisticated hardware, very carefully designed hardware, with a traditional control approach, with very smart controls engineers setting everything up, but with comparatively less emphasis on the kind of decision-making aspect.

51:20And I think that at a particular point in time, that actually made a lot of sense because we can't build the physical body. It doesn't matter what kind of decision-making system is running on it. But to our earlier discussion about generalization, at this point, we're at a stage in the development of these things that even though we can do more on hardware, in many ways it's good enough. And the big challenge is how to have the decision-making loop that actually works and that reacts intelligently to everything in the environment. And the place where I would draw the dividing line between those is, like, decision-making loop doesn't mean symbolic decisions.

51:59It could mean low-level decisions. The question is, do you need to take the rest of the environment into account, or are you just dealing with a robot? So if you want to do a backflip on flat ground, you must have to deal with a robot. But if you want to pick up a coffee cup off of a table, even though that's maybe in some ways simpler than doing a backflip, you really have to understand what's going on in the rest of the world rather than just your own body. And that dividing line, I think the way that technology has panned out, I think it's fair to say that that is the dividing line between AI and controls.

52:29Like controls is when you have to control the robot body. AI is when you have to take into account what goes on outside of the robot. And I think that's why you see this divide because I think a lot of the demos where you mostly needed to deal with the robot itself and not the rest of the world, really good controls could allow you to, could admit a very good solution there. And another thing I would say here is, okay, if there's a lot of controls work that goes into doing some particular skill, well, there is actually something to learn from that because if you can hand design a controller that performs a sophisticated behavior, very likely you can also learn that controller.

53:08So just that proof of existence that the thing is possible, and not only possible but also simple enough that a person could build it, Because remember, people are, you know, at the end of the day, even with code these days, the kind of complexity that people can handle is not as high as the kind of complexity that AI can handle. So if a person can hand design something to do a backflip or do some acrobatics, that's a really great proof of existence that there exists some relatively parsimonious control law for doing that skill. And parsimonious does mean generalizable. So if it's simple enough for a person to design, probably there's something fairly general in there.

53:40And if you can learn it and automate it without having to have the human controls engineers in the loop, So that's good news.

53:47Sergey Levine:And then last question for you is, if you could go back to when you just entered the industry and give yourself some advice, knowing everything you know now, what would you say? One thing that I've learned over the last few years, which I think is a little different than kind of my original mindset, is I think that addressing robotics effectively requires using very broad prior knowledge. And I think that there's this idea that a lot of people in robotic learning have, which I think I shared initially, that since people learn things kind of from scratch, maybe robots should learn things from scratch too.

54:25So, for example, in some of our early work on large-scale robotic learning at Google, we had this, what we call the arm farm project. We set up a bunch of robot arms in a conference room, actually, because we didn't have a proper lab, but it was a conference room. And we had them all grasping objects. And the idea was, well, if they grasp millions of objects, they'll learn very general grasping strategies. And it basically worked. Like they could learn to grasp objects, but it was very hard to like take it further from that to the next level. So, okay, now I can pick up anything, but like, so what?

55:00It didn't serve as a very good stepping stone for more complex skills. And I think part of that was that we were approaching this like very blank slate. Like let's start from zero and see if knowing nothing in advance, the robot could start picking up behaviors. But I think that it's much, much more practical to get all this to work if you can combine robot experience with knowledge that you can pull in from other sources. Like, for example, I was very skeptical initially about the utility of language. And I think scientifically this is defensible, which is that like, hey, you know, like animals can do some pretty impressive things, like monkeys can do really cool stuff.

55:34But monkeys, as far as I know, can't speak, at least not very eloquently. So maybe our robots should also be able to do stuff and they don't necessarily need to understand language. But I think the subtlety there is what's important is not language, it's prior knowledge. that you can put in as a scaffold in your learning process. And you can pull in that knowledge in all sorts of ways. Humans don't necessarily pull it in entirely through language. And monkeys certainly don't. They do it from observation, from observing other people, other creatures, and so on. So there's lots of sources of power knowledge.

56:06But the point is that you've got to get that power knowledge in there. Otherwise you're actually faced with a harder problem than what humans and animals have to solve. Because if a person had to figure out how to assemble IKEA furniture, but they've never actually encountered any article of furniture in their entire life, like, okay, that would be like pretty difficult because they don't even know what like, what the point of this is or what the end game looks like. So yeah, prior knowledge is important. And while I'm still a big fan of learning things through experience, I think that, you know, my advice to myself would have been take prior knowledge more seriously.

56:38Sergey Levine:Awesome. Well, thank you so much for your time, Sergey. I really appreciate it. Yeah, thank you for your questions. Hey, thank you for watching this podcast. If you liked it and you want to see the show grow, please support with a comment or a like also if you have any recommendations for people you want me to bring on please drop a comment guests like barbara liskov mike stonebreaker mark brooker these were all people that i brought on because someone left a comment on another note aside from the podcast i'm working on building the ergonomic keyboard that i wish existed here's a glance at the prototype it's a split keyboard so there's two sides this is in the case but yeah we launched on kickstarter and we hit our goal within eight hours of launching i really appreciate it if you were one of the people who grabbed one of the early units we're now working on the long journey of building the tooling now and so if you still want to pick one up i've left the late pledges open on kickstarter so you can grab one there i'll put a link in the description thank you again for watching the podcast and i'll see you in the next episode

From the publisher

Sergey Levine is one of the world's top robotics researchers and co-founder of Physical Intelligence. We talked about where humanoid robotics is today, thoughts on the Chinese robotics ecosystem, and his predictions for future timelines.


• My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/

• The Kickstarter page for it: https://www.kickstarter.com/projects/ryanlpeterman/compose-simple-ergonomics-beautifully-done


Podcast links:


• YouTube: https://youtu.be/9OSbaPjv0Rc

• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835

• Transcript: https://www.developing.dev/p/sergey-levine-current-state-of-humanoid?r=n49ky


Thank you to this episode's sponsor for supporting my work:


• WorkOS: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code, check them out at https://workos.com/


Timestamps:


(00:00) Intro

(00:37) Where are we today

(04:20) Most surprising capabilities so far

(07:03) The most inspiring real world robotics

(08:36) If OpenAI or Anthropic got into robotics

(10:22) Chinese robotics

(13:15) Will one lab breakout from the rest

(16:59) Thoughts on a concrete roadmap

(21:03) Generalization and demonstrating it

(26:04) Types of data and which is best for robotics

(34:34) Why humanoid robotics differs from Waymo

(37:10) If humanoid robotics failed here is why

(39:55) Are there hot take modeling architectures in robotics

(42:05) Thoughts on AI safety in robotics

(46:44) Top robotics research paper recommendation

(49:35) Why is Boston Dynamics less top of mind

(53:47) Advice for his younger self

(56:42) Outro


Where to find Sergey:


• Google Scholar: https://scholar.google.com/citations?user=8R35rCwAAAAJ&hl=en

• Website: https://people.eecs.berkeley.edu/~svlevine/

• Wikipedia: https://en.wikipedia.org/wiki/Sergey_Levine

• X/Twitter: https://x.com/svlevine?lang=en

• LinkedIn: https://www.linkedin.com/in/sergey-levine-5a31a24/


Where to find Ryan:


• Newsletter: https://www.developing.dev/

• X/Twitter: https://x.com/ryanlpeterman

• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/

• Threads: https://www.threads.com/@ryanlpeterman

• Instagram: https://www.instagram.com/ryanlpeterman

• TikTok: https://www.tiktok.com/@ryanlpeterman


Referenced in this episode:


• Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA / ACT paper): https://arxiv.org/abs/2304.13705

• Emergence of Human to Robot Transfer in Vision-Language-Action Models: https://arxiv.org/abs/2512.22414

• Summary of human to robot paper: https://www.pi.website/research/human_to_robot

More from The Peterman Pod

All 60 episodes
Sergey Levine: Current State of Humanoid Robotics, China & Future PredictionsThe Peterman Pod · 58 min
Listen in VO