π0: A Foundation Model for Robotics with Sergey Levine - #719

18 Feb 2025 · 53 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Summary of Podcast Episode: π0: A Foundation Model for Robotics with Sergey Levine - #719

Podcast Details

  • Podcast Title: The TWIML AI Podcast
  • Episode Title: π0: A Foundation Model for Robotics with Sergey Levine - #719
  • Host: Sam Charrington
  • Guest: Sergey Levine, Associate Professor at UC Berkeley and Co-founder of Physical Intelligence
  • Date: [Date not provided in the transcript]

Episode Overview In this episode, Sam Charrington interviews Sergey Levine about π0 (pi-zero), a general-purpose robotic foundation model developed by Physical Intelligence. The discussion covers the model architecture, training strategies, data collection methods, and future developments in the field of robotic learning.

Key Concepts and Discussions

  1. Introduction to π0
  2. Objective: To create general-purpose robotic models that can serve as a foundation for various applications.
  3. Vision: Achieve the capability of robots akin to those seen in science fiction, allowing them to adapt to a wide range of tasks without needing to develop a new system for each specific function.
  1. Model Architecture
  2. Components:
  3. A Vision Language Model (VLM) paired with a diffusion-based action expert.
  4. The model integrates spatial representations necessary for controlling robotic movements.
  1. The Training "Recipe"
  2. Focus on Pre-training and Post-training:
  3. Pre-training: Involves collecting a diverse set of real-world data from human operators using teleoperation rigs.
  4. Post-training: Fine-tunes the model with high-quality, consistent data to improve task performance.
  1. Data Collection and Challenges
  2. Human Input: Data is primarily sourced from various human operators controlling robots in realistic environments.
  3. Synthetic Data and Reinforcement Learning: Exploring the integration of synthetic data and reinforcement learning to enhance model capabilities.
  1. The Role of Reinforcement Learning
  2. While not heavily featured in π0, Sergey acknowledges its importance for future iterations of robotic foundation models and aims to integrate RL into the development process.
  1. The FAST Tokenizer
  2. A new tokenizer that improve action representation by optimizing the ways actions are tokenized, utilizing concepts from data compression to enhance model performance.
  1. Open-Sourcing of π0
  2. The team released π0 as an open-source model to encourage experimentation within the robotics community, hoping to spark innovation and variations in its application.
  1. Use Cases and Future Directions
  2. The episode discusses potential applications of π0, including household tasks like laundry folding and assembly.
  3. Future advancements may include improving robots' capacity to follow complex instructions and adapt to novel environments.

Key Takeaways

  • General Purpose Models: The podcast emphasizes the importance of developing general-purpose robotic models to streamline the creation of robotic solutions across various domains.
  • Innovative Training Methods: The conversation highlights the need for innovative data collection strategies and training methods to ensure robots can learn effectively from diverse experiences.
  • Continuous Learning and Adaptation: Future iterations of robotic models will need to address the challenges of generalization and the ability to adapt to new tasks and environments seamlessly.

Conclusion The discussion between Sam Charrington and Sergey Levine presents exciting developments in the field of robotics and artificial intelligence. The introduction of π0 represents a promising step towards creating adaptable, intelligent robots capable of performing a wide array of tasks, positioning the robotic community to leverage these advancements for innovative applications in real-world scenarios.

For more detailed show notes, visit [TWIML AI Podcast Episode 719](https://twimlai.com/go/719).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00In robotics, every time you want to tackle a new robotics application, you have to start an entire company around, or you have to start an entire research. cloud. So each robotic application is just like an enormous amount of work. And if we can have these general purpose models that can serve as the foundation for a huge range of applications, that would actually allow us to get robots to the next level. It would get us the kind of generalist robots that we like see in science fiction, basically.

0:35All right, everyone, welcome to another episode of the TwiML AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Sergey Levine. Sergey is an associate professor at UC Berkeley and co-founder of Physical Intelligence. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Sergey, welcome back to the podcast. It's been a little bit. Yeah, it's been a while. Thank you for having me. I'm looking forward to digging into our conversation. We will be talking about all things physical intelligence, and in particular, we'll be digging into the Pi Zero model that you launched last fall and recently open sourced, as well as some other work that you've been doing at the company.

1:15I think the last time you were on the show was January of 23. So a couple of years ago, talking about reinforcement learning and robotics as part of our AI trend series. What have you been up to since then? I guess a lot of that is going to be founding a company. Yeah, it's been a bit of an adventure. So, you know, I'm probably, I wouldn't have imagined myself founding a company, honestly, a few years back. But as I and a number of my colleagues saw how things were advancing in 2023, we decided that the pieces were really falling into place to make a very serious and larger scale push towards true robotic foundation models.

2:01And we felt like to really get robotic learning to advance to the next level, it needed to be something bigger than what we could do individually in a more academic context. So that's why we really pulled the trigger on this thing. And yeah, I'm excited to tell you more about it. Maybe we can start by having you share a little bit about kind of the hallmark of physical intelligence's approach. Like, what are you really going for here? And what are the ways in which you're trying to differentiate? So at physical intelligence, we are very committed to the mission of actually building general purpose robotic foundation models.

2:33So, you know, there's a lot that you could do with robots. Like, you know, robots can automate warehouses. They can, you know, you can have self-driving cars, all that stuff. But we're much more interested in what it would take to get things to the next level, where you could actually have general purpose models in the same way that like, you know, like chat GPT is general purpose. So it used to be that we would have fairly specialized systems for natural language processing or fairly specialized systems for computer vision. But with general purpose foundation models, you could have a single system that could be adapted to a wide range of different behaviors.

3:03And it's tremendously powerful, especially in robotics, because in robotics, every time you want to tackle a new robotics application, you like have to like start an entire company around or you have to start an entire research lab. So each robotic application is just like an enormous amount of work. And if we can have these general purpose models that can serve as the foundation for a huge range of applications, that would actually allow us to get robots to the next level. It would get us the kind of generalist robots that we like see in science fiction, basically. So that's not a short term goal.

3:30I don't mean to like make this seem simple. It's not like we're like on the verge of having that tomorrow. It's a very long-term goal, but a lot of the pieces that we think are necessary to make this happen have been falling into place over the last few years. And now is the time to like really ramp it up and make a very serious push for it. And what are some of those key pieces? In robotic learning, if we kind of step back and look at the challenges that have made robotic learning hard to use in the real world. Well, robotic learning has enormous promise, like the idea that a robot can just go into some environment, like figure out on its own how to solve a problem.

4:00That's enormous. But in practice, that dream has never quite panned out before because a few things have been missing. One big one is machine learning works best when there's a lot of data. And in robotics, that has always been a very deep tension because while you can go on the internet and get tons of images and tons of text, there's no like internet of robot data where you can go and just like get lots of data for your robot. You have to create that data yourself. So that means that everybody who wants to get their robot to do something now is put in this position where they have to create big data sets or create small data sets and use models that don't work very well.

4:34So in practice, that's always been kind of a showstopper in robotic learning. Now, the big advance there is the community has figured out much better how to create transferable general purpose models so that now we actually have some hope that we can create models that can control a wide variety of different robots that can then be very efficiently fine-tuned to a new robot or a new application, or in the future, are maybe even prompted in zero shot to perform some new times. Now that drastically lowers the barrier to entry. Other challenges that have been really big, so that was challenge number one, other challenges that have been really big in robotic learning, generalization and common sense.

5:10So for a robot, unlike chatbot, for instance, when the robot goes into some physical environment, has to deal with whatever's going to happen there. It can't just say like, oh, you know, I don't know what's going on here. I'm just going to bug out and hope someone else deals with it. if someone like, you know, if it's driving along the floor and there's a sign that says like slippery floor, do not enter, it has to, even if it's not designed to read signs like that, has to be able to parse that situation and react intelligently. And this is where things like vision language models come into play, where previously we wouldn't have had any idea how to handle this.

5:41Now we can actually draw on general purpose semantic priors, learn from internet scale data from language models and vision language models and address those kinds of situations. use. The third really big one has to do with robustness, reliability, performance. And this is where advances in reinforcement learning are really making it much more feasible to get highly precise, highly performant, and highly reliable systems. So I think with those three things, we're in a much better position now to actually make the push towards making these things a reality. I was going to ask specifically about reinforcement learning.

6:16You've been on the podcast three times previously, I think. And each of those times, RL has been a big focus. And I think the last time we spent a lot of time talking about the Simna real gap, which was kind of a promising idea for allowing us to collect a lot of data and simulation and then use that in the context of real robotics. And looking through the PI Zero paper, RL isn't mentioned a lot. you know, in what ways does, well, let me back up with the question. A, to what degree is RL a part of the Pi Zero work? And then more broadly, how do you see RL coming into play with regard to, you know, this idea of robotic foundation models?

7:02Yeah, that's a really good question. So Pi Zero, you can think of it as really kind of a first step towards robotic foundation models. So, you know, we're pretty proud of the work. We think we have like a really great demo. We think that we did a good job of it. But the truth is that it's a beginning, like it's not an end. It's the beginning of our journey. And to use an analogy, we've seen in the world of language models that there were a number of steps that we have to advance through. So first people figured out transformers. They figured out good architectures. Then they figured out things like, you know, the early language models, things like GPT-2, models that could not really solve any practical problem, but at least have that like plausibly human-like text generation.

7:47And then there were a number of algorithmic pieces. People had to figure out instruction fine-tuning. People had to figure out RLHF. And, you know, RLHF was used famously for chat GPT, but the ideas were initially prototype in much simpler settings. So a lot of those pieces have to fall into place. And you can think of pi zero as one of those early steps. It's not the complete solution. It doesn't have all the ingredients. And what we're doing at Physical Intelligence is we're developing multiple ingredients, many of them in parallel, and they will connect more and more in the future. So in the same way that reinforcement learning came into the world of language models later on, once the basic foundation was already pretty solid, in the same way that I expect reinforcement learning will make a really big difference for our foundation models at physical intelligence once all the pieces are ready to be integrated.

8:31And in the meantime, we'll develop them in a way that allows us to most advance the science. So let's dig into R0. Talk a little bit about R0 from a model perspective. You mentioned that it's built on and based on VLM's vision language models. Talk about what surrounds that VLM and how you think about the model as a whole. The first thing I would say there is I think it's very important when we talk about foundation models to remember that the foundation model is not, it's not just about the model itself. It's not just about the architecture, but it's also about the recipe. And in fact, for folks that actually work on this stuff, a lot of the most impactful research on foundation models does nothing to change the architecture whatsoever and actually instead refines the recipe, the kind of data mixture that is used, the way the data is curated, all that sort of thing.

9:21So I think there are a lot of clever things that we did for the architecture. But in truth, I think that the most significant innovations were more with the recipe than with the architecture itself, meaning that I wouldn't be at all surprised if a very different architecture, but with the kind of refined recipe we use, would also do very well. So there isn't necessarily something that special about the architecture, but I'll talk about the architecture as well. So there's a very particular challenge that you have to address if you want to adapt vision language models to robotic control, which is that vision language models, they're very good at understanding images.

9:56They're good at understanding text. They're also good at generating text. The problem is that robotic behaviors, especially dexterous, sophisticated behaviors, require precisely representing spatial movements. And that is not something that is very easy to express in text. So I could tell you, you know, verbally, like, here's how you play tennis. Like, yeah, you hold the tennis racket in your hands and you move your muscles, but that's not really going to transmit to you the actual precise details of the move. Move your muscles carries a lot of weight there. Yeah, exactly. There's like a lot wrapped up in that.

10:27It's doing a lot of work. Fortunately, we actually have a very good methodology for representing continuous spatial information. It's just not necessarily with transformers, but with diffusion models. Diffusion models are the kinds of models that have been used for things like image generation. They've also been used to generate spatial structures for things like protein design. Like actually the best protein design systems in machine learning today are based on diffusion models that will actually generate positions of atoms in a molecule. So we use the same kind of techniques. And the challenge was to figure out how to connect these to vision language models.

11:01So you have a vision language model that processes image and processes text. And then there is a second component of this model that will actually attend to the internal activations inside the VLM run diffusion and produce continuous trajectory snippets, what we call action chunks, to actually control the robotic arms. And that gets you a lot more ability to handle complex, multimodal behaviors, very dexterous behaviors, very precise behaviors like you need for folding laundry. So that was pretty important to make this work well. And there's a lot of details that go into making that efficient, like it has to run in real time.

11:32So there's a lot of kind of things that you have to get right. Yeah, I was going to ask about that real-time point. but specifically I picked up on you saying continuous. Did you mean that in a sense of continuous versus discrete or continuous in a sense of time? Yeah, so this is actually interesting because it's something that we have actually been experimenting with. Both senses matter here. So for the original Pi Zero model, it was discrete in time, meaning that it still discretizes time into specific steps, but it handles continuous spatial information. So the way this model works is because it has to run asynchronously, it actually outputs a little chunk of about 50 time steps at 50 hertz, so roughly a second of movement.

12:16It doesn't actually run it for the full second. The next inference run will land maybe like a third of a second to a half of a second later and override it. So there's a window of 50 forward-looking action steps. Okay, three times a second roughly. But in regard to the time discretization, one of the things we've been working on, and this connects to our more recent work on better tokenization for vision language action models, is that once you start thinking of it in this way, you can also represent the behaviors in continuous time, for example, in the frequency domain. And that turns out to have some very appealing properties because it allows you to represent things in a way that is more compressed and therefore easier for the VLM to learn.

12:56Interesting, interesting. We'll definitely come back to that. So you've got a core vision language model. In your case, this is based on GEMMA, the GEMMA model. And then you kind of attach the action expert model to that. And I thought it was interesting in the paper you talked about one way of thinking about this is like a MOE. Can you elaborate on that a little bit? Yeah. Yeah, so this was actually very inspired by some really interesting work that came out of meta research, where they were concerned with a pretty different problem. They were actually concerned with building models that could produce both text and images.

13:38And they developed a few different models, including chameleon and transfusion, that would produce text via standard, like kind of softmax, the way that everyone else does it. And then for the tokens that corresponded to images, they would actually have a different set of weights that would be part of the same transformer, like in a mixture of experts, except that in a mixture of experts, like the weights kind of all, you know, all the different experts, they produce text. Here, there's a separate expert for producing images and a separate one for producing text. But they can all attend to each other.

14:09So they will look into each other's activations, which makes the whole model very expressive. so we uh we were very inspired by this design but instead of using it for generating text and images we actually uh built on this design to produce these uh vision language action models where we'd have an expert that was from the from the pre-trained vlm and then a separate expert which was actually much smaller so that it could run in real time that was responsible for producing the actions so you can kind of think of it as almost like to use a this is the kind of thing that my academic colleagues might judge me for.

14:41But to use an overly human-like analogy, you can think of the little action expert as like the little visual cortex of this model. And the VLM part is like the visual cortex. So the little action expert is the motor cortex and little VLM part is the visual cortex. Prior to the fast work, which you mentioned and which dug into thinking about tokenization or rethinking tokenization, how did you represent action tokens in the original pi zero. Yeah, so in the original pi zero paper, the actions are represented as continuous vectors. So that's where the diffusion model that's running through the action acts but it just directly produces continuous numbers.

15:21Like they're literally joint angles like in negative pi to pi. So tokenization is something that came later as part of the model. Got it, got it. So with that in mind as kind of an overview of the model, let's dig into the recipe which you mentioned, you know, rightly so, is an important piece here. So, well, for background, like what's the recipe for language models in very broad terms? Language models are trained on lots of data scraped from the web and that's like next token prediction. That's what everybody knows and loves. And that's kind of actually the easy part. I mean, it's maybe like hard from a system standpoint, but in terms of all of the curation recipes, like, yeah, you just take everything you can get your hands on.

16:01The challenge is like scraping enough data, But yeah, but you're not being like really smart about creating just the right, you know, instruction examples. And then comes the hard part where, which is typically referred to as post-training, where people do a lot of really smart things to generate synthetic data, to curate, to get human experts to produce data. And a lot of the improvements that you see with like AI assistance comes from people being really smart about post-training. but both parts are important because post-training by itself wouldn't work because the model gets all of its prior knowledge from that pre-training phase and pre-training by itself wouldn't work because the pre-trained model is just like a text completion engine like you prompt it with some somebody saying something not terribly intelligent and it will say things that are very unintelligent right so it's just like kind of emulating the internet basically so with robotic foundation models people haven't thought about it that way even though it might seem kind of obvious but typically people haven't gotten that far because of the there weren't large robotic data sets.

16:57The fundamental data problem was in the way. Yeah, exactly. So people would take whatever data they have for their robot. And because it was data that they were typically producing in small amounts, it was like typically high quality data. And they would just like try to really overfit to it and really nail those particular behaviors. We collected a much larger data set and it was much more heterogeneous. So it was much more natural for us to start thinking about a pre-training and post-training separation. We didn't actually design this from the ground up, like kind of in the paper, we have this narrative that like, oh, like, you know, this is the right way to do it.

17:26Like, clearly we thought this through. What we actually did is we added more and more data. And then we found that like, yeah, the model training is taking a really long time. So it's better not to like retrain it each time, but to have like a little pre-trained thing that we'll fine tune. So it kind of grew organically. It's like a little bit of convergent evolution that we ended up on a similar recipe. And we found a lot of really interesting things that maybe in hindsight are kind of obvious, but to me, we're like, we're pretty cool. We also found that using very high quality data, is actually bad.

17:56And the reason that using very high quality... For pre-training in particular, as opposed to task post-training? If you use only high quality data. Only high quality data. If you use only, this might seem kind of bizarre, like if you use only good data, it's a bad idea. And it's kind of interesting why. The reason it's a bad idea is because the robot is never perfect. The robot will make mistakes. And when it makes a mistake, it finds itself in some situation that doesn't happen in the good data. So if you have low-quality data, mediocre data, yes, you see some mistakes, but you also see the recoveries from those mistakes.

18:28And that ends up making the robot much more intelligent. So you really need both. If you have just the mediocre data, the robot doesn't know how to do the task well. If you have only the high-quality data, then it won't know what to do when it messes up. And is the pre-training data, that is all unsupervised, just hours and hours and hours of robots performing tasks in video? or is there some degree of task supervision? So in this case, all of our data came from people controlling the robots. We are actually studying autonomous data collection as well. But in this prototype, all the data was collected by our robot operators across a very heterogeneous range of different tasks.

19:06Now, people naturally have a lot of variability, both in how they perform the task and how good they are at it. And, you know, some people are good at some tasks, some people are good at others. So what we actually found to work pretty well is to take all of the data we've collected ever across all the tasks and across all the different robot types, use that for pre-training, and then put a non-trivial amount of effort into curating the data for post-training. And for post-training, it's very important to get data that is obviously of high quality because it shows the robot how it should do the task, but also data that has a degree of consistency.

19:37So you want the robot to perform the task with a consistent strategy that reliably works well. So that's where it takes a bit more care to pick out the right kinds of behaviors. The amount of post-training data is not typically actually all that large. So it depends a lot on the task. Harder tasks need more data, but it's between 2 and 20 hours. The pre-training data is about 10 ,000 hours, so it's much, much larger. And so you mentioned, so this pre-training data is collected from humans operating the robots. What does a typical frame or sample of data contain? Yeah. So the way that the data is collected, and this is something that we do put quite a lot of thought into, is we use a teleoperation rig.

20:20You can think of it as a leader follower setup. So there are these, they look a lot like robotic arms. They're these leader arms that a person holds. They're kind of lightweight arms that are meant to track their movements, which they can use to demonstrate behaviors. and the behaviors that they demonstrate, we typically try to make them kind of as realistic as possible. Like it's supposed to be kind of like reflective of a job. Like maybe the job is to like, I don't know, replace the paper towel roll on the paper towel holder or fold all the laundry or like assemble cardboard boxes. So these episodes, you know, they'll range from a few minutes to tens of minutes in length.

20:54The episodes can be segmented and tagged with language. So that allows us to get the robot a degree of language and instruction understanding. And then they're processed into training data and used to train them all. You also used as part of your pre-training data set, the OXE data set, OpenX Embodiment data set. My understanding from the paper is that it was less than 10 % of the data. Can you talk, was this just your source of bad data? Talk about like where that comes in and why. Yeah, this is an interesting question. So, you know, one thing that I'll preface this by saying is that, you know, Pi Zero is partly it's like a technology demonstration.

21:36Like, you know, we wanted to show the kind of cool tasks the robot could do, but partly it is actually genuinely intended as a step towards general purpose foundation models, which means that it needs to be able to understand all sorts of different robot types and all sorts of different skills, not just the ones that we demonstrated in our videos. And we've been using the model, actually, in all sorts of ways, like, for example, fine-tuning to new robot types, where we released a demo with another company called Astrobot, where we fine-tuned our robotic foundation model to their robot with a small amount of data.

22:05And it's a humanoid robot. It's very different from the robots we have. So we really wanted to make sure that as much as we could, we could get PyZero to understand a variety of robot embodiments. which means that we basically took all the data we could get our hands on from as many different robots as we could find and try to fold it into the pre-training data set got it so the big contribution there is just the number of types of robots that it had whereas you trained you built your pre-training data set on uh eight or nine i believe that one had more on the order of 20 yeah exactly and and the number is growing regularly right so uh we are adding additional robot embodiments all the time, both from our own experiments and from data that we get from researchers, from partners, from all sorts of sorts.

22:47And so you've got this pre-training data set, you pre-train the model and you do the fine tuning based on specific tasks. And is that post-training, rather? Post-training, is that, you know, simply kind of additional iterations with higher quality data, or is there more that goes into that part of the recipe? Yeah, so in the current prototype, it is, the actual training part of it is pretty straightforward. It's just supervised learning, take the pre-trained checkpoint and fine-tune it. Most of the care goes into selecting, curating, or collecting the data in the right way. And this is the kind of stuff where we, you know, for example, for some of the more difficult tasks, it might even be hard for humans to do them.

23:34So we might have, there might be a particular person who's really good at doing that task. And maybe all of the post-trained data just comes from that one expert person. Whereas the pre-training might contain data of the same task even, but just from less expert humans. So I think this provides kind of an interesting sketch for how we could see robotic learning in the future, where it's natural that as the tasks get more complex, there might be a lot of value in getting the best possible data from experts, which might actually be a lot better than what that expert could do if they were like actually doing the job routinely.

24:06Like someone can probably fold the t-shirt much more effectively if they just have to like do it really well for three minutes than if they were doing it for like 100 hours, right? So you could actually ask somebody like really just go all out and do this really, really well. And then we'll teach that behavior to the robot in the post-training phase. Yeah, let's talk a little bit about the laundry folding demo because I think for a lot of folks that, you know, saw the initial Pi Zero release, that was super compelling. I can't even remember what year CES this was, but Samsung or one of the big companies did a laundry folding robot.

Read the full transcript

24:39This was at least five years ago, probably more. And, you know, first of all, they had to set up the clothes, you know, just so. And then even still, the folding wasn't all that great. And it was a purpose-built laundry folding robot. You know, so a lot of care and, you know, curation of the scenario and poor results. Whereas, you know, in your example, the robot takes the laundry out of the dryer into a bin, pulls it out and starts folding it. You know, talk a little bit more about, you know, that specific example and, you know, to what degree you, you know, is it, you know, is it zero shot? in the sense that you didn't do anything specific to get the robots to perform well on that task rather than relative to other tasks beyond the post-training?

25:35Or was there care somehow that went into that scenario and other scenarios to make the robots perform well on them? Yeah, that's a good question. So there was a lot of care, but no care went into the code in the sense that the code that is running on the robot for laundry folding is exactly completely 100 % identical. to what's running when it's like building the cardboard box or cleaning the table. The care went into carefully working with the robot operators to come up with a good strategy that the robot could execute well. So again, this is coming back to that point about post-training, like, yeah, it really matters what kind of data you get for the post-training phase.

26:17Now, fortunately, you don't need an enormous amount of that. You need an enormous amount of pre-training data. And there you can be, you don't have to be nearly as careful. But for the post-training, like, yeah, the strategy of folding the shirt, there's like five different ways you could do it. And that particular way of doing it is better. It's more reliable. The robot can more easily get the shirt into that position. But at the same time, you know, one of the things that makes laundry folding a really great illustration of this pre-training, post-training principle is that while you could have a really smart strategy that you've designed and you say, okay, like this strategy, if you guys do this for like 20 hours, the robot will be pretty good at this.

26:49you can still get into lots of unpredictable situations. So you're still relying very heavily on that pre-training data to get the robot out of the weird situations and into the ones from which those nice strategies will work. And in the videos that we released, there's a few things that even kind of surprised us. Like we tried to, you know, obviously when you build a robotic system, the first thing you want to do is you want to mess with it and see how you can get it to fail. So Michael, who's one of the researchers working on this, he, in the videos, you can see him, he puts a shirt in the basket the robot takes it out starts falling then he takes another shirt and just drops it on the table and the robot just picks this thing up and just drops it back in right so it's like yeah get this out of here i'm working right it's even more interesting because at first it's it tries to just continue working yeah and then it realizes that hey i got to get rid of this thing and it's just whoop yeah exactly and uh like the post-training data doesn't have that obviously because it's a small amount of very carefully controlled data where everything is just right.

27:44But yeah, probably somewhere in the pre-training data, somebody messed up a little bit, took out two shirts by accident and put it back. Maybe they didn't put it back in the same way. So the robot has to generalize, but that diversity of lower quality data illustrates a lot more of these issues. And as long as you have a good generalizable model on top of it, that can extrapolate from those patterns, then you'll actually get those kinds of emergent behaviors. And that even surprised us. Like we kind of thought like, yeah, that'll generalize some cool way, but we didn't expect quite that kind of nice extrapolation.

28:11And along the lines of curating that training data, if you think about like folding a shirt in terms of stages, one stage is getting the shirt into the position, the next stage, then you execute your folding steps and then there is putting it onto the stack. Does the training data isolate those individual steps or is it all end-to-end run through the process? It's end-to-end. And so we basically, we instruct the operators to perform the entire task. We do, afterwards, for the language condition stuff, we do have human labelers segment the data and annotate it with text, but the collection is just fully intent.

28:57And I think this is important because if you, you can sort of imagine that for now we're doing this in the laboratory, but in the future you could have a robot that is actually like out there in the world really doing real work, maybe initially under teleoperation, but that teleoperation is creating the data that would later help it become more autonomous. Yeah. Yeah. I think the question came from the idea that, you know, once you've got the table laid out, your scenarios may be more well-defined, but when it's in the shirts in the basket, you've got a lot more, you know, think of it combinatorially, like a lot more positions that the shirts could be in that kind of thing.

29:32And so you might need more data of, you know, getting the shirt out of the basket and flat onto the table. But it sounds like you didn't do that. And maybe a follow-on question is, you know, are there areas that you worked with or introduced synthetic data into the process? Yeah, I think this question is getting at something really important that is important to think about when we're collecting robotic data and something that people sometimes don't put as much thought into, which is that very broadly speaking in machine learning, the one thing we know works is when training matches test so getting real authentic data is really really important um and we try actually pretty hard to make the data collection process for our systems as realistic as possible and as representative of what would actually happen uh for real tasks like even to the point where right now we're doing some experiments collecting data for like kitchen type tasks and the robot is like literally in our kitchen like our building has like a little office kitchen and the robot is like cleaning up like the actual office kitchen When we're doing the table cleaning task, we would like eat our lunch, put our lunch there, and the robot would go and like clean up the actual lunch.

30:42So the more realistic we can make it, the more training will match tests and the more the model will actually generalize to real world situations. Now, at the same time, to your question about synthetic data, something that I'm pretty excited to study in the future is the degree to which having this initial foundational understanding might actually make it easier to incorporate synthetic data. It's like if you're playing a video game, right? We understand how the real world works. So the somewhat abstract cartoony environment in a video game or in an animated film makes sense to us. But we're coming to that image with a lot of the physical priors we learn from the real world.

31:19And it may very well be that by having this really nice foundation of real world experience, it might actually be easier to incorporate less realistic data in the future because the robot will represent its experience in this way that abstracts away those difference. The Pi Zero model, I think the full model is 3.3 billion parameters. Can you talk a little bit about kind of the application of scaling laws and how you see that evolving? Yeah, that's a great question. So we wanted to start with a pretty lightweight model because, well, we didn't want to wait for a long time for things to train. We also have to run this thing in real time for inference.

31:54so we weren't particularly like deliberate in saying like this is the optimal size we just picked kind of the smallest model that seemed like it had you know enough of that internet scale knowledge baked into it but one of the technical challenges that I do think is really exciting to study in the future is how we can both have the benefit of large models for the more elaborate kind of semantic reasoning the you know like all the fancy chain of thought stuff solving complex problems and still connect it up to like a little motor cortex model that can run fast enough to control the robot. And there's been a little bit of work academically, including in my group and other research groups, that shows that if you have these vision language action models, you can actually do the same kind of reasoning stuff, the same kind of test time compute tricks that people have used for LLMs.

32:42So we had a paper, for example, called the Embodied Chain of thought from my lab at UC Berkeley, where we trained a vision language action model, a different one, predating Pi Zero. This is based on OpenVLA, which would actually reason through things in the task. So it would say like, okay, you're asking me to put the banana in the plate. Okay, well, to do that, I should first find the banana in the image, find the plate. I know the bananas are yellow, so let's find like the yellow thing. I need to find where my hand is. Okay, my hand is over the banana. That means that the right thing is to move down.

33:10Move down means coordinate, coordinate, coordinate, and then execute the task. And this actually helps. This helps a lot to get good results, especially in unfamiliar settings, because in those unfamiliar settings, while the robot might not kind of instinctually know what to do, if it reasons through the task, it can succeed more reliably. What was the performance like? It sounds very slow. It was slow, yeah. I mean, this was also on an academic compute budget, let's say. So yeah, it kind of like moves, pauses, moves, pauses, but it's 50 % better. So you do get a big improvement. Obviously, now the systems challenge needs to be overcome.

33:45I'm curious if one of the things that is clear, you know, both in our conversation and the paper is that the team has been really good at like taking tidbits of, you know, pieces, innovations from lots of different places and pulling them together into this work. Was there anything that came out of the recent DeepSeek, you know, R1 stuff that, you know, was inspiring or that you are looking forward to playing with in terms of improving this model? Yeah, that's a really good question. I mean, it's something that we have been thinking about a lot. I don't think there's anything concrete that I can really talk about now because, you know, when you have kind of vague ideas, you want to like at least test them out first before you risk saying something stupid.

34:25But a lot of us found that to be really inspiring. And I think that especially the way that they have this fairly simple recipe where it's just like pre-train and then use RL. Obviously, like the fancier recipe that's a little more conventional is like a little bit better, But just the fact that I pre-train and then use RL works so well, that's pretty cool. And I really wonder how some of those could be adapted to things like these VLA reasoning models where the robot could actually use RL to train itself how to think more carefully. But this is just speculation at this point. So you mentioned FAST a couple of times.

34:59Talk us through the motivation there and how it builds on what you're doing with PyZero. Yeah, yeah. Yeah. So we used a diffusion-based model for Pi Zero, which we had to kind of adapt with VLMs. The action expert in particular. Yeah, the action expert, exactly. But prior work on VLAs used discretization. And the reason that it used discretization is because VLMs naturally output discrete tokens. So if you want to adapt the VLM most directly to control robots, you take your actions and you basically turn them into text. In fact, in the first vision language action model ever developed in the RT2 model, they were literally like numbers.

35:37Like you represent the action as the number 132 and that is converted to floating point and then run on the robot. The trouble with this though is that we know from language models that it really matters how you represent those tokens. Like much as we'd like to imagine that language models are so brilliant that they can just handle like whatever you throw at them, the way that you represent your tokens makes a huge difference for performance, not necessarily because there's any particular knowledge baked into it, but because it makes the model train a lot more effectively. Basically, different letters in English occur with different frequency, and therefore they carry different amounts of information.

36:15And it's much easier for a neural network to learn when all the outputs are kind of equalized in terms of how much information they carry. So it really helps to represent that text in a way where every single token has about the same amount of information. And then the model learns very quickly and generalizes effectively. And of course, there's no reason to think that that wouldn't be true for actions either, except that that's not how actions were being tokenized before. If you're just representing them as numbers, the occurrence of those numbers on the internet has absolutely no bearing with how frequently those actions occur for robotic data.

36:49So the foundation of modern tokenizers is basically compression. Like it turns out that the way that you get every token to have about the same amount of information is to compress your text because the optimal compression will provide as much about as much information, about the equal amount of information for every single bit. So we can take inspiration from compression methods for continuous data to develop a tokenization for actions. And we thought about this and one place where you see compression of continuous data is images. So you use like JPEG compression. JPEG compression compresses images.

37:24JPEG compression basically represents an image with different frequencies. So different frequencies occur in different extents in different images. You put your image in a frequency domain, roughly speaking, and then you compress the weights on those frequencies. So we decided to try that. Essentially, it's like run JPEG compression on your action chunks. And that actually gets you a much better tokenization. It gets you the ability to represent the same actions with the same fidelity, but with a much smaller number of tokens. That by itself is not actually as important. What's important is that now those tokens contain about the same amount of information.

38:00and when you train on this your model trains about four times faster than it would if you used uh the the naive tokenization that's a that's a really big difference and it's not just that you spend less time waiting like if your model trains four times faster you can train for the same amount of time it'll train four times better right so that's a really big deal uh that allowed us to train much better models especially for uh tasks that required a lot of generalization to language following uh this part we're still trying to understand like what's the connection between this in language following.

38:29Like, you know, maybe it's something like if you fit the data better, you'll get to understanding language better. But one of the things this allows you to do is take, for example, the open source Droid data set, which is a big data set of Franca robotic arm manipulation that was collected across a number of different universities with language annotations, and actually get a policy that will generalize to a new Franca arm in a new location. We took the Pi Zero Fast model and we sent it to our friends at all sorts of different universities. There was a video that I posted on Twitter from some folks at UPenn that ran this.

39:02And it's like literally the first time they load this up on their Franco robot, they tell it's someone like pick up the pineapple and put it in the basket. And it actually goes and does it. And like, you know, the folks running this are not used to models working out of the box. Like usually robotic models don't work out of the box. So it's pretty cool that we could get that by just figuring out a better way to compress actions. And when you say that you train the model using this idea, what is the model in this case, the entire model, just the VLM, just the action expert? Yeah, so the model at this point, for the Pi Zero Fast experiment, is just the VLM component.

39:37And I don't know if that's necessarily a good choice, like we honestly didn't even really experiment with it very carefully. You could apply the same Fast Tokenizer to other VLMs, so we used it with the OpenVLA model, for instance, and the tokenizer is like available on Hugging Face, so anyone could grab it and run with their own VLA model. And so does the action expert change at all? Or is that the same with and without the fast tokenizer? Basically nothing else changed. Like the only thing that changes is the tokenizer. Okay. And so the mapping from the, you know, the fast tokens to the actions, that is still learned, but it's just learning a different thing, you know.

40:18Yeah, that's part of the tokenizer. Part of the loop, right. Okay. Interesting. Interesting. And so talk a little bit about the open sourcing of all this. That's kind of the most recent thing that you guys did. Yeah, yeah. So we figured that now that we actually have a model that can perform some pretty interesting tasks and also a model that other people could actually run. So we had like a little closed beta where we sent this to a few other folks, a few universities, a few companies. that this would be a really great thing to share with the community. Now, obviously, we have our own reason that we want to do that.

40:57Like, we want to see how people use robotic foundation models because that'll help us learn how to make them better in the future. But we also think that this is a great way to galvanize a lot more interest in this stuff because we saw how with language models, just getting the pre-trained checkpoints out there and allowing people to fine-tune their own models created this, like, wave of creativity where people come with all sorts of new things to do with them. So we're really looking forward to see what people will do with pre-trained robotic foundation models. So we really want to get that out there.

41:24We have a few demos where, you know, it'll run on some robots that will run on Droid. The truth is that it is like a very early prototype, I think, in the grand trajectory of robotic foundation models. So probably most things that people will try it for won't work. But just from seeing the kinds of experiments people do, I think we'll all learn a lot about it and we'll figure out a lot of new ideas for how to make robotic foundation models more applicable in the future. And so what specifically did you provide in open source? I know you provided the weights for the models. Are you providing any of the details around the training recipes and or data sets, those kinds of things?

41:58Does it allow you to fully replicate PyZero or does it allow you to use what's already been done? So the open source repo includes the code for fine tuning. It includes the base checkpoint. It includes another base checkpoint for the fast version of the model with the tokenizer. and then it has a few example fine-tuned models like it has a model fine-tuned for droid it has a model fine-tuned for aloha and these are really kind of intended as like demos if you just want to like try it out and see uh the main use case that we anticipate is for folks to take the base model and then fine-tune into their own robots because the robotics community is still very fragmented like everyone's setup is different so if you are lucky enough to somehow have a setup very similar to droid or very similar to a setup that we had then you might be able to try to run it in zero shot.

42:42But really, the intended use case is to collect some of your own data, fine tune it, and then try to use it to solve your task. And in order to do that, would you need a setup that offered the same kind of operator guided data collection methodology? That's a really good question. And this is actually one of the things that we hope to understand better. So, you know, we explored a few different ways to collect data. We kind of have our own intuition for what works and what doesn't work. But if somebody tries to fine-tune the Pi Zero model, they'll collect data. Maybe they'll do it the same way that we do.

43:16Maybe they'll do it differently. And I think it'll be really interesting to see what kind of recipes work and what kind don't. So we know what recipe worked for us. We described it in the paper. So hopefully if someone tries it, it'll work for them. But maybe we'll try other things. And then we'll find out that maybe you can get away with a lot less data if it's more consistent. Or maybe there's some other kind of thing that we didn't know was true that's actually true. So we kind of just really want to see how people experiment. And you, most of the robots that are depicted in the paper and in the demo videos are arm-based robots.

43:45You mentioned some work on a humanoid based robots. Are there other form factors that you see folks experimenting with? Yeah. So we ourselves have successfully been able to run the model. Well, either we ourselves or with our collaborators and partners on robots that include humanoids, single arm, dual arm, and mobile platforms. In principle, the model supports very flexible action representations, as long as your action is less than 32 dimensions. So, you know, we have to pick a maximum, so we pick 32. I've gotten videos of people that have successfully run this model with five-finger hands. I'm not entirely sure how they did it because no one has actually told me, They just send me videos like, look, it's working.

44:30In principle, it should be possible to run it for like navigation and things like that. We haven't done that ourselves, but that's very much within the kind of the constraints of the model. So as long as it fits to the dimensionality, somebody could try it. I'm really curious to see what works and what doesn't. I'm sure there'll be some limitations. Like if you train on an octopus arm, well, that's probably a little too different, but I'd be curious to see what happens. When you think about kind of the robot, the platform landscape, like are there accessible hobbyist types of or enthusiast types of arms that you could try it out on?

45:04Yeah, that's a really good question. So we actually were very lucky to get some help from Hugging Face, who helped us with doing a PyTorch port of the model. And they also have a very low cost arm. I can't vouch for its quality. I've never actually used it, but they seem to have been able to do some pretty cool things with it. So if anyone is interested in like a really low-cost robot, checking out LaRobot and their PyTorch port of PyZero, as well as some of the, you know, I think they've actually tried it out on their arms. That could be a good way to go. But honestly, like even the nicer arms that we've been using, the arms that we used for things like the Aloha experiments, these are not all that expensive.

45:46I think Aloha was like on the order of 20K. So the whole system is on the order of 20K. The arms are, I think, somewhere in the$6 ,000 range per arm. So if you just want the follower arms, like the minimum thing, I think you could probably get away with like 15K or so. But the cost is, you know, seems to be going down every year. So I wouldn't be surprised if like next time this year that these things are even cheaper. Okay. Awesome. Awesome. By the way, one of the things I'm really excited about is like, well, if we get these robotic foundation models into people's hands, if the cost of hardware keeps dropping, maybe it will be like very practical for anybody to just play around with their own robot.

46:20I mean, you know, 15K is a bit expensive, but if it drops like, you know, another factor of two or four, maybe that'll be actually pretty practical. Awesome. Well, what's next? Where do you see it all going? Yeah. So there are a number of next steps that I'm pretty excited about. One of the things that I really want us to be able to do better, and I think that that's something that we should be able to talk about more in a few weeks is do a much better job of following complex instructions. So one of the things that's really cool about ChatGPT is you can actually give it a prompt that describes in detail almost like a job you want to do.

46:52So you don't just tell like, oh, you know, please write me an email to my boss. Like you would actually describe like, write me an email to my boss that describes how I want to raise and blah, blah, blah, blah, blah, you know, whatever. It's kind of a complete description of a task. And it would be really cool if we can do that with robots too. Instead of just telling it like fold the shirt or clean the table. You can tell it like, hey, I'm throwing a party. I already put the plates on the table, but there's some trash. So put away the trash, but leave the plates where they are and make sure the fork and the knife is next to the plate in a nice, tidy way.

47:20So something that kind of actually like describes what you actually want. And maybe the robot would do the task. Maybe he might even ask you like, hey, I didn't get that part. Like, are you sure you want me to like leave that plate there that doesn't look like it belongs? Like you can have a much more intricate interaction. And the interesting thing there is not just the interaction part, it's the ability to instruct the robot to really do the job. I think that that can also be a really interesting mechanism, not just for getting robots to do sophisticated tasks, but also getting robots to repurpose their skills.

47:51So if the robot can actually get this more intricate task and think about, hey, I've learned to do these particular behaviors, how do I adapt them to solve this new problem that I've been presented with? That's something that you could do with a lot of that semantic knowledge inside of VLMs, but it requires a little bit more processing. Maybe it requires a little bit of that test time compute. And that's something that we've been working on that we hope to be able to tell people more about in the near future. Do you make a distinction between complex instructions to instruct the robot to do a complex multi-step task and instructions, complex instructions that instruct the robot how to do a complex task?

48:31Yeah, that's a really interesting question um i would like to not have to make that distinction like the distinction is this is actually important in the sense of like the capability but i think you can have a system that does both of those and in particular like people actually do it in a very flexible way where people bring to bear their own knowledge and their own skills uh but they also benefit from the information that they're offering so if you already uh you know if you are a professional laundry folder like it's enough for someone to just tell you like hey like go do your job and you already know what to do.

49:02But if you are not experienced at that task, then maybe you'll still be able to do it if only somebody provides you with more instruction. And if you have a model that is trained end-to-end that can perform this kind of reasoning, put together the steps that it knows, and flexibly decide what to do, then it can do either of those depending on the situation, depending on its prior knowledge. Got it. Was there anything else on that list? Yeah. So other things that we're doing that I'm pretty excited about, we are trying obviously to push the boundaries on the generalization for these systems. So I'm very happy with how we've been able to demonstrate pretty sophisticated tasks, but it's still a challenge if you want those tasks to work with any object and in any environment.

49:43Generalization means a lot in this context. It's generalization to task, generalization to environment, generalization to object, generalization to platform. And to instruction as well. Yeah. So it's a very big space. And because it's so big, it's also kind of hard to say anything particularly definitive about it. But it's something that we're studying quite a lot. We're trying to see how different ways of collecting data, different ways of training models, different ways of transferring knowledge from internet scale pre-training can facilitate generalization. And I think something that's really exciting there is like those moments, like I mentioned, when we had that droid model and somebody was able to run it at a different location, students were able to run it at another university and just see like that spark of life where they just give it some task and just goes and does it.

50:26Maybe it does it poorly, maybe it's slow. But I think that's really special. And I think that the more we can enable that, that kind of aha moment where you just load up the model on your robot, tell it to do something, and it actually kind of gets it. Like that, that I think is really, really important. I think that if we're careful with transferring knowledge from the web, curating data in the right way, and setting up our model in the right way, then I think we can get a lot more of that. When I think about, you know, this is an example like this, where, you know, someone is using this model, presumably the same type of robot, but just another environment.

51:00And you think, oh, well, that should work. It's a software program. You put it someplace else. And I think back to, I think it was one of, it was like a Peter Abiel demo of like the robot trying to do knot tying and like just different colors of rope and stripes and stuff like that totally confounded the robot. Yeah. My very first robotics project, which was actually with Peter Abiel, we have to use the same background in every trial because that was the background that worked for the robot. So I'm really glad that we've gotten past that at this point. That's awesome. Well, I'm looking forward to keeping in touch and keeping up with the updates coming out of the work.

51:38Very cool stuff.

From the publisher

Today, we're joined by Sergey Levine, associate professor at UC Berkeley and co-founder of Physical Intelligence, to discuss π0 (pi-zero), a general-purpose robotic foundation model. We dig into the model architecture, which pairs a vision language model (VLM) with a diffusion-based action expert, and the model training "recipe," emphasizing the roles of pre-training and post-training with a diverse mixture of real-world data to ensure robust and intelligent robot learning. We review the data collection approach, which uses human operators and teleoperation rigs, the potential of synthetic data and reinforcement learning in enhancing robotic capabilities, and much more. We also introduce the team’s new FAST tokenizer, which opens the door to a fully Transformer-based model and significant improvements in learning and generalization. Finally, we cover the open-sourcing of π0 and future directions for their research.

The complete show notes for this episode can be found at https://twimlai.com/go/719.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
π0: A Foundation Model for Robotics with Sergey Levine - #719The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 53 min
Listen in VO