In short
Eye On A.I. - Episode #176: Sergey Levine: Decoding The Evolution of AI in Robotics
Podcast Overview Host: Craig S. Smith Guest: Sergey Levine, Associate Professor at UC Berkeley Theme: The evolution of AI in robotics, focusing on the latest advancements in AI control of robots through reinforcement learning and embodied AI.
---
Key Topics Discussed
- Introduction to Robotic AI
- Sergey Levine discusses the potential of home robots that can perform a variety of tasks.
- Emphasis on the importance of generalization in robotic manipulation skills.
- World Models and Language Models
- Imitation Learning: Robots learn from human demonstrations.
- World Models: Dynamics models representing environment behavior to aid reinforcement learning.
- Discussion on the significance of utilizing broad datasets for effective training.
- Challenges in Learning-Based Control
- Discussion of the difficulties in data acquisition for training robots in open-world environments.
- Importance of learning individual manipulation skills rather than just perceiving tasks.
- RTX Project Overview
- Collaboration between Google, Berkeley, and other universities aimed at generalizing controllers across different robot morphologies.
- Achievements in language-conditioned manipulation tasks, improving success rates significantly through pooled data.
- Model Architecture and Development
- Introduction of RT1 and RT2 models, employing transformers and vision-language models.
- Differences in model capabilities and architecture, focusing on versatility in tasks and better performance in complex queries.
- Future of Robotic Control Research
- Ongoing need for standardization and sharing of models across labs.
- Importance of having reusable models in robotics akin to those in NLP and computer vision.
- Hardware Development Impact
- Hardware constraints and their influence on the capabilities of AI-driven robotics.
- Exploration of how current hardware can meet the demands of complex tasks with adequate learning methodologies.
- Open-Source vs. Proprietary Debate
- Examination of the landscape of data accessibility in robotics.
- The balance between industry resources and academic research in advancing robotics.
- Global Research Landscape
- Comparison of robotics research developments in China and the U.S.
- Discussion on the presence of significant advancements in hardware and software from Chinese companies.
---
Key Takeaways
- Reinforcement Learning: The pivotal role in enabling robots to learn complex manipulation tasks.
- Generalization: Critical for effective home robotics, requiring extensive and diverse datasets.
- RTX Project: Demonstrated the benefits of pooled data from multiple research labs for improved model performance.
- Future Directions: Focus on creating adaptable, general-purpose robotic models that enhance productivity in various domains.
---
Conclusion The conversation with Sergey Levine provides a comprehensive insight into the evolving landscape of AI in robotics, highlighting the transformative potential of these technologies and the ongoing challenges faced in research and development. As advancements continue, the integration of AI and robotics promises to revolutionize our daily interactions with technology.
---
Stay Connected
- Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
---
Additional Notes For those interested in further details, the podcast episode includes a deeper dive into technical methodologies and future research directions. It is highly recommended for anyone wanting to understand the intersection of AI, robotics, and their implications for the future.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Imagine someday building a home robot that can perform a variety of tasks in your kitchen. then it becomes a lot more than just a perception problem. Then you really need to be able to learn the individual manipulation skills. You need to be able to generalize very broadly. Hi, I'm Craig Smith, and this is Eye on AI. Today, I talked to Sergey Levine, an associate professor at the University of California, Berkeley, who does research at the Robotic Artificial Intelligence and Learning Lab at the university and is pushing the boundaries of what AI control of robots can do. Sergey talked about some of his recent work in reinforcement learning and the aggregation of data sets from robots around the world to help train a model that can generalize across different kinds of robots.
0:53It's exciting research into embodied AI, bringing the transformative technology out of the computer and into the real world. I hope you find the conversation as fascinating as I did. So, Sergei, could you start by introducing yourself? So I'm an associate professor at UC Berkeley. I previously did my PhD at Stanford. And I also spend one day a week at robotics at Google, where I also work on robotic learning. My research concerns, of course, robotics, but also many other related techniques in machine learning, reinforcement learning. Lately, my group has also been doing work on various things related to reinforcement learning for language models, computational design, things like this, other aspects of decision making.
1:42You know, everyone's talking about world models and they're combining world models and language models. Are you working on world models at all? What's your view of that? But yeah, I guess there are a few things I could say about it. So generally, if we want to control robotic systems, there are a number of ways that machine learning can enable that. One very simple way is imitation learning. Imitation learning amounts to taking demonstrations, typically provided by a person controlling the system, and then imitating those demonstrations to try to produce an agent. And that can work for robots.
2:23It can work for lots of other things. Arguably, language models are just giant imitation learning machines because they're imitating how humans generate text. There are lots of other ways to do it. So a world model is essentially a dynamics model that represents how the environment will evolve in response to the agent's actions. And we can learn that from data, too. Typically, in reinforcement, it's referred to as model-based RL. So model-based RL means you train a model of how your environment behaves, and then you use that model to figure out how to act in the world. And it's a very old discipline, actually.
2:55In fact, the first learning-based control methods before model-free RL became such a popular thing were actually model-based RL methods. Some of the earliest neural network control methods actually used dynamics modeling. And again, there are a million different ways to instantiate this. you could instantiate dynamics models or world models by, for example, taking image observations and doing video prediction. You could also instantiate them by learning non-reconstructive representations, representations that sort of, roughly speaking, capture the state of the system without necessarily grounding it back into pixels and then predict that.
3:31So there's a lot of different approaches for doing this. Recently, I've been talking to Wave about their Gaia model, and you know i've seen the the videos but they they have that model built into a controller connected to a controller to operate a autonomous vehicle what's different about that structure architecture from what you're working on with reinforcement learning uh i don't think there's that much i could say about it because i don't i don't know how their system works um i mean i mean I've seen the public material, same as everyone else, but I don't really have any insight about the details there.
4:12Maybe one thing I would say, though, that most methods for learning-based control don't necessarily need to predict the raw pixels that the robot's camera will observe in the future. That is one way to do it, and there's a lot that can be done with that. But I think that the more significant distinction is actually how well we're able to utilize data to then produce more optimal decisions. And going through prediction is one way to do it. And you could predict pixels, which is what video prediction models do. You could also predict outcomes or rewards, which is what value functions do. At the end of the day, they are actually not all that different.
4:55and maybe a bigger distinction as to whether you can get a system that really works in the real world is what data it's trained on. For example, if you want robotic manipulation systems that will actually work in broad open world environments, you need to train them on data for broad open world environments. So a lot of what I'm actually concerned with in my research is how do we develop learning-based control techniques that can use large amounts of data and how do we figure out what data sets we can acquire to get really generalizable, in my case, often robotic manipulation skills, but also robotic navigation skills and things like that.
5:28Systems for manipulation for things like warehousing, oftentimes those problems can be, to a large extent, reduced to perception problems. So if you structure your environment in the right way, then as long as you're able to detect where the objects are, then you can use hand-designed strategies for tackling that. That tends not to work so well if you want to take the robotic system into more open-world environments. Like if you imagine someday building a home robot that can perform a variety of tasks in your kitchen, then it becomes a lot more than just a perception problem. Then you really need to be able to learn the individual manipulation skills.
6:03You need to be able to generalize very broadly. So maybe one thing that I could discuss here that could be relevant is a project that we actually did recently. this was actually a collaboration between Google, Berkeley, and several other universities on trying to see how could we get robotic controllers that actually generalize across different robot morphologies. And that's actually really important because if a lot of this comes down to data, then it's very difficult to get the kind of breadth and diversity of data from a single robot that can allow the kind of broad generalization that you want from a home robot.
6:40But if you can pool data from lots of different robots, then maybe you can actually get that kind of coverage. And furthermore, if you can actually do that and you get a system that generalizes across robots, then you get something really cool, which is in principle, someone could put together some new robotic system and then plug this kind of robot brain into it and immediately get something that can control that robot. Now, the work that we did so far on this wasn't really that concerned with building better models so much as just getting this diverse data set and just applying kind of standard techniques that we'd already developed in previous works.
7:09And that actually worked decently well. So this project was called RTX. And the idea there was that we got data from, in the end, it was 34 different research labs. Google was one of them. Berkeley was, actually, there were two labs at Berkeley that contributed data to this. And then we trained a model on this to control, to perform basically language condition manipulation tasks. So think like you give the robot an instruction, pick up the tomato and put it in the bowl, and the robot is supposed to do that. And then we took this model and we handed it off to the different labs that contributed data and had them tested in comparison to whatever model they had for their research, basically trained on their own system.
7:44And the multi-robot model actually was, on average, about 50 % better in terms of success rate. And that's actually pretty interesting because this was competing with whatever each lab's individual system was. And presumably, if they're good researchers, they build something that works pretty well in their system. Now, this was actually an imitation learning approach, language condition imitation. I think that whether it's imitation or prediction or world modeling, I mean, I think many of those techniques could be made to work. I think the high order bit here that I want to get across is that by actually getting this data set, you could actually get a system that you could plug into all these different robots and actually get good results out of it.
8:16Hmm. That's fascinating. That model was being trained on data sets from all the various participating labs? Yeah. So in these experiments, we were not testing whether it could generalize to a new robot. And that's a very exciting frontier for this, but that's still in the future. This was just trying to answer the question, if you include data from the other labs, does that one lab's robot get better? Now, of course, if you're sort of in the minority, if you're one of the groups that contributed a relatively small amount of data, you would expect to seek apparently more benefit from everyone else.
8:48Interestingly, actually, even the majority contributors saw a lot of benefits. So probably the largest data set at about 100 ,000 trials was from Google's own robot, the mobile base that we used in a lot of the robotics research there. And with that system, we were able to actually test it on various, we have this kind of test suite of difficult queries. They're actually meant to be queries that require synthesizing, you know, pre-trained knowledge from the web, as well as good instruction following ability and so on. So these are, they require spatial reasoning, things like that. And on the hardest test, we actually saw 3x improvement in performance over just using the Google data set.
9:28Now, that's actually, in my mind, pretty profound because the Google data set was very carefully curated, collected by basically professionals that were collecting robot data. And the fact that including all these additional sources of data from a long list of academic labs actually led to that much improvement really suggests that there is something kind of magical that happens when you combine enough data from enough different sources. Yeah, so for these experiments, we were actually passing the model around. Okay. The data set is public now. So anybody could take that data set and download it and train their own models.
9:59And we actually have an ongoing effort at UC Berkeley that my students are, for that initial experiment, it was just the model weights. That's fascinating. Just the model weights. So the architecture of that model was being replicated in each lab. They weren't using their own model. Correct. Yeah. So it was exactly the same model with exactly the same weights that had to drive all of the robots in all the locations. And that is actually, if you think about it, it's a very non-trivial thing, right? Because all the model gets to see is what the robot receives through the camera and has to figure out that, okay, now I'm driving a UR-10 industrial robot versus now I'm driving a low cost WITOX robot, or now I'm driving a FRANCA or the Google robot and adjust the controls according.
10:43When I was at the lab, as I recall, that you had robots that were networked, So the learnings from one were updating a central brain, and that then was controlling the other robots of each robot. Have you done broader experiments like that, very much like this? Yeah, I'm glad you asked that. So this was a lot of what we were trying to do over the last five years, actually. And in some ways, this multi-robot training effort, it's kind of, you know, partly it came about kind of an acknowledgement of the limitations of this kind of arm farm approach. So getting lots of robots in a room is great if you want to prototype, let's say, reinforcement learning algorithms.
11:32But if you really want broad generalization out of it, like they can't be all in the same room. So you really need to get much better coverage of the world. And by aggregating data from robots in many different sites, Now you can get much better coverage. Now this is still a prototype for what might be a larger system, because these are still data sets collected by researchers essentially doing science experiments. So you could imagine that in the future, the aggregation wouldn't be across different research labs, it would be across different deployed robots. Now that of course is a much more complex undertaking that requires more than just science.
12:04It also requires some kind of organizational effort, consensus from companies and so on. But that I think is kind of the real thing once that comes about. and you could imagine a future where the data streams from a variety of different deployed robots in lots of locations are agglomerated and then used to train one centralized robotic brain which can then be handed back to these robots to improve their performance. And the key thing that we want to de-risk with this project is just if you do this at any scale, you know even at the scale of academic labs, can you even get a policy that can drive all the different robots?
12:36Because if that's not possible then aggregating heterogeneous data wouldn't work and we would need to somehow figure out standardization. Standardization is hard. So what we know now is that we don't have to worry as much about standardization. We should just worry about scale. Yeah, this model then, the weights are being passed around, and they're controlling different form function robots, right? I mean, or were they all just variations? So in these experiments, the robots were all arms with parallel jaw grippers. We are right now experimenting with generalization across single arm and bimanual systems.
13:12At some point in the future, we'll also look at multi-fingered systems, things like that. So so far, truth in advertising, it's one arm with a parallel draw gripper. They're just different brands of arms. Now, they do vary a lot. So the small-scale kind of hobbyist Widow X arm is like maybe 50 centimeters long, kind of smaller with a weak gripper. The big DR-10 robot is an industrial robot meant for manufacturing, quite a bit bigger, beefier, has a stronger motor, stronger grip, or that sort of thing. So there's variability. They're still the same sort. Right. And the model that you're training on this aggregated data, is that the reinforcement?
13:48Can you describe that model? We actually trained two models. One was based on the RT1 model, which was developed at Google last year. And the RT1 model is basically, it's a transformer that reads in language instruction commands, an image and then it outputs discretized tokenized actions. So it's kind of almost like the most obvious way to design a transformer-based policy. The second model was the RT2 model, which is a more recent development, which actually uses a backbone from a pre-trained vision language model. So vision language models are trained to look at images and output responses to textual questions.
14:27So you give it a picture and you say, like, is there a dog in the picture? And it will produce some text to answer that. And then we took this vision language pre-trained backbone and then further fine-tuned on robot data to output robot actions in response to robot observations. So you can sort of think of like the VLM has a number of tasks that it can do. It can answer questions. It can produce captions. Now there's one more task, which is given a robot instruction output the actions for the robot. And now that's a much more powerful model, because it has all that internet knowledge baked into it from the vision language model pre-training.
14:57And that's the one that we've used for the more complex queries with spatial relations and things like that. Yeah. Is most of your work on the data side or on the model side? Well, it's really both. And to some extent, they also go hand in hand because depending on what your algorithms are able to handle, that will inform the kind of data that you need to get. For example, a lot of the more algorithmic work that my lab does these days is concerned with techniques for offline reinforcement learning. Offline reinforcement learning is basically a a way to take data and produce more optimal policies.
15:30So imitation learning methods, they take in data and they produce policies that reproduce the behavior in the data. Offline RL methods take in data and attempt to produce behaviors that are better than the average behavior in the data. So intuitively, you can think of it as using data to figure out what options are available to you and then selecting the best among those options. In fact, methods that use world models, as we discussed before, can be seen as offline RL methods, because typically the way they work is they train the world model on existing data, and then they use it to extract better control strategies than the typical thing that was demonstrated in the data set.
15:59But there are also other ways to build offline RL techniques that don't rely on world models, that rely on value function, things like that. Where do you think the research is heading? Because everything's moving so fast. Is there for robotic control, do you think that the research will settle on one architecture and then there will be different flavors of that architecture, but everyone will agree, yes, that this is the best way, and then it's a matter of training, generalizing across robots and networking data. Or do you think that there will be a series of models that will be used for various functions?
16:42Yeah, good question. So I'll give you an answer. It's a slightly aspirational answer, so maybe this is more like where I wish things were headed. I don't know if this is necessarily where things will head, But I think it's very important for robotics to kind of adopt a paradigm where we are in the habit of having reusable models, where just like in computer vision and NLP, if a researcher produces a good model, other robotics researchers should be able to use it. Now, that might seem like a very obvious thing, but this is not actually how robotics works today. Most robotic learning research, the artifact that is produced is not actually the model.
17:20It's the code or the paper or the insights. The models themselves are almost never portable, never mind across labs, even across different locations in the same lab, different times of day in the same lab, that sort of thing. and I think we really need to move that towards a setting where we have models that are trained on data sets that enable generalization across different locations and systems different objects that sort of thing that we can then give to other researchers other practitioners that will also run on their systems and once we've got we've got a good flow for doing that maybe using things like this RTX data set that has multiple robots maybe using some other days but something where we can and just get in the habit of doing that, then we can actually make more progress as a community towards shared generalizable systems.
18:08Now, until that happens, there's absolutely no question about whether people will use the same architecture or the same model. If they can't even share anything across, then that won't work. But once we can share something, and probably the key to that is a data set that enables that, then the community can figure it out. Maybe at that point, perhaps it'll be realistic to have a single pre-trained backbone, that, you know, like the LLAMA model in natural language processing, an analog to that in robotics, and then people can build on top of that. Or maybe there will be several such things. Maybe there will be a few big ones that kind of the big, well-equipped labs produce that others will then build on.
18:43But before we get to any of that, we need to just get in the habit of actually building models that others can run. The other side to robotics is just the hardware. And, I was talking to a guy the other day who was talking about where robotic control systems are heading. And he's not a roboticist or an AI researcher. But he was waxing very optimistic about there being household robots within three to five years. And that sounded unlikely to me because just the hardware alone is not, at least the hardware that I've seen, is nowhere near being able to do, you know, releasing it into an unstructured environment full of randomness.
19:35Do you think that the hardware is moving along with the AI or is it lagging? That's a good question. I think that a very important part of that question is just what kind of hardware we need. I think to a large extent, learning methods should actually lower the bar for the hardware that's necessary. Basically, the exercise you can do is you can get one of these little trash picker devices and see what kind of tasks you can do around the house with it. I mean, obviously, it's very limiting, so there's some things you couldn't do. But there's also a lot of things you could do with it. Certainly you can tidy up the floor, put things in different locations in the kitchen.
20:18It's actually kind of surprising how much a relatively primitive robotic system can accomplish. So there's very nice work out of Professor Chelsea Finn's group that I also helped with a little bit by a student named Tony Zhao, who developed a bimanual robotic system out of two low-cost robots from Charleston Robotics. So these are not even the fancy industrial arms. These are basically very fancy hobbyist robots. So they cost about, I think,$5 ,000 each. And most of his kind of cleverness in his research was in devising a very convenient teleoperation system, a teleoperation rig that he could hold with his hands and control this fairly cheap bimanual system.
21:06And he could demonstrate all sorts of very complex behaviors. You get this thing to like put a shoe on a foot, use tape to tape down a box, things like that. And the learning methodology that could produce the autonomous policy was well-designed, but not particularly profound. It's sort of used state-of-the-art transformer-based techniques, but didn't really have any particularly surprising innovation. The key to it was really building a really good teleoperation rig that allowed him to produce those behaviors and then very high-quality engineering to then get that down to a policy. So this is called the Aloha system.
21:43For those who are listening, I encourage you to check it out. And it probably gives some idea of what even very primitive hardware is capable of if it's equipped with the right data, the right kind of teleoperation rig to provide that data, and kind of good, bread and butter, modern machine learning techniques. Now, that's still not going to do everything around the house. But I suspect that for folks that watch these Aloha videos, it'll kind of maybe slightly change their mind in terms of the kind of hardware we really need for everyday tasks. So probably there is still some innovation, but it might be actually less than you think.
22:14That's interesting. And then the controller side, the AI side, the model side, I mean, if that is adequate, that hardware, how much more improvement is needed on the control side? That's a complex question because that's probably very heavily dependent on the required bar for robustness and the degree of generalization. So in some ways, it sort of parallels the autonomous driving story, right? Like if you wanted to build an autonomous car that could succeed in like 90 % of cases, well, that's probably something that we've had for over a decade. But if you want an autonomous car that will succeed, that will avoid catastrophic failures with enough robustness that you could just deploy it on any road in any city, just dealing with all those tail cases, that's still an open problem.
23:05And I think with home robots, it's going to be the same way that if you want to lop off the bulk of the things and the bulk of the situations, maybe that's not quite there yet. But I think that it's reasonable to imagine that we get there soon. But how long it takes to get that long tail fully figured out, that's a much more complex question. I think that one thing that's pretty interesting is the degree to which vision language models have progressed over the last, really over just like the last 12 months. And that's especially relevant for robotics because while the way that vision language models are typically used is more for traditional perception tasks, question answering, that sort of thing, the ability to reason about visual observations, perform inferences about spatial arrangements of objects, that sort of thing, that is something that is likely to translate into better robotic capability.
23:58And because generalization is one of those big challenges, I mentioned this long tail issue, I think there is a lot of reason to be optimistic about the potential for those models to eventually improve the robustness of robotic controllers as well. People are talking about combining language and vision, or I should say language and world models into agents that can reason, plan, and take action. That sounded to me very much like robotic control. I guess what I'm asking is, is that research and the people who are in robotic control research on different tracks? The answer is a little bit complicated, but maybe the short version is that yes, it's closely related to a lot of robotics problems.
24:42In fact, there's plenty of work in robotics on using language models for essentially constructing plans and then connecting those plans up to some kind of control mechanism that can bring them about. Now, probably this stuff started maybe like roughly two years ago. Probably one of the more well-known works in this area is the SACAN paper from Google, which used a language model to plan long horizon behaviors for robots. initially in this field one of the big challenges that people were concerned about is how to connect up the language model to perception and action because standard language models have to operate on symbolic representations of the world so you have to take those symbolic representations and somehow weld them on to rich sensory perception and complex actuation now initially the way that this was done was kind of along the lines of what you described by trying to construct some sort of a joint planning procedure that would figure out both a probable sequence of, you know, symbolic steps, essentially language, and the corresponding behaviors that would bring that about.
25:55There's actually a paper from one of my colleagues, one's called Grounded Decoding, which proposes a Bayesian filtering approach to doing exactly that. That said, something that we've seen over the last, maybe like six to nine months is that increasingly with vision language models becoming more powerful a very appealing alternative instead of doing this is to actually train models end-to-end to solve the entire problem now those models can still be doing planning if you have a vision language model that outputs text and also outputs actions you can do essentially the analog of chain of thought prompting you can tell it okay here's some complex problem and produce steps for solving that problem and then once you produce those steps then produce the actions.
Read the full transcript
26:37And that works. So you could tell a robot, okay, like make breakfast, and we'll say, okay, to make breakfast, I need to do this and this and this. And then for the first step of that, it'll try to output the actions. So that's a viable way to use visual language models. But then it's still, you would still end up with one model that does that. And that's very desirable, because if you have one model, then you don't need to deal with this problem of trying to somehow stuff visual observations into a symbolic representation to then pass into the language model. Basically, instead of designing that interface by hand, it emerges naturally through joint training of the whole thing.
27:08So this is actually the principle in which the RT2 model works. And one of the examples there that illustrates this kind of chain of thought style approach is we asked that we want to intentionally construct a scene where the correct behavior is a little bit non-obvious. So we had a scene that had some common household items and had some tools, the wrong kind of tool. So it's supposed to hammer and a nail. There's no hammer, but there's a rock. and we ask it, okay, you need to hammer in the nail. What should you do? And then it figures out that it should pick up the rock. It actually says rocks, and then it goes and executes the corresponding actions.
27:45So now that's very primitive planning, right? So it's kind of more somatic inference than planning, but these things are in their infancy. I think they'll progress a lot more over the next few years. Do you, in the last five years, which is about the time I think since I spoke to you, Has the progress in your field specifically mirrored the progress in generative AI? I think that progress in robotics always does tend to lag behind everything else, because when we figure out effective learning techniques, then it's always a longer journey to go from kind of conceptual method to small scale prototype to larger scale prototype.
28:34Because with generative models, well, you can harvest lots and lots of data off the web. So the lag between developing a method and then scaling it up to internet scale data is typically relatively short. With robots, that's usually not the case. So while certainly modern advances in generative of modeling have made a big impact on robotics. And there's particular very interesting adaptations of those techniques that combine with reinforcement learning planning and so on. I would say that so far, we have a lot of good indications of the potential for these things. But we don't have the kind of large scale prototypes that have been produced, for example, for diffusion models, for image generation, or for language models.
29:14And I think the key there is actually getting these kinds of reusable models with large and diverse data sets that would make it possible for us to produce these larger prototypes. Yeah. So what's next in your lab? Yeah. So one of the things we would like to do is provide the community with pre-trained models, now that we actually have a data set to work with, that can be easily adapted to a variety of downstream applications. So not just a model that can do anything, and that's maybe too ambitious of a goal, but at least a model that can be adapted to do anything. So if you could imagine, let's say, a model that is pre-trained to take in language, take in maybe goal observations, other forms of commands, and produce outputs for a variety of different robot embodiments, not with the goal necessarily of solving every problem, but providing a really good initialization so that somebody that has a particular specific robotic system with a particular desired formulation of their task, a particular objective, they could take this and with a much more modest amount of data adapt to their problem.
30:14And I think that now that we actually have good multi-robot data sets and fairly mature techniques in terms of how to train models with variable inputs and outputs. We're actually just about ready to do that. So our first prototype for this should be coming out very soon, but it's just going to be the first step. From there, a lot of what we have to investigate is what does the life cycle of such a system actually look like? What are the right techniques for efficiently fine-tuning robotic foundation models to particular domains to different morphologies different commands and so on and there's probably actually a lot of interesting questions to be answered there for example robots can collect data autonomously so could you for example have an autonomous fine-tuning procedure based off of one of these pre-trained models could you have a fine-tuning procedure that respects safety constraints things like that so there's a lot of interesting questions that we can answer once we have that base model all set up there's a lot of people are i've been talking to about the the proprietary open source debate with regard to generative AI.
31:14In robotics, is there an analogous situation where there are enterprises that have tremendous resources? I mean, the robotics are not as compute intensive, the models that you're talking about. Is that right? And so is it more equal at what's happening in industry and what's happening in research? Yeah, it's complicated. So certainly compute constraints are an issue, right? Especially once we go into vision language models, the most effective vision language models are actually the largest models out there. So the largest version of the ARC2 model, for example, is 500 billion parameters. So very much in the same ballpark as the largest models out there.
31:55Of course, you can do a lot of experiments on a much smaller scale, and that does make it somewhat more accessible. In terms of data, it's kind of interesting. There are definitely companies that have large numbers of robots deployed. Those are not necessarily the companies that have the most interesting data, though, because if they're deployed in a warehouse, it's mostly grasping of items. And maybe in some ways, the open data from researchers is actually more interesting. The picture changes somewhat if you go into mobility, things like autonomous driving. Like, yeah, there's going to be big industrial players with their own proprietary stuff.
32:28But even there, data sets constructed from dashboard-mounted cameras that are out there now are actually very large. Certainly not as large as what Tesla or Waymo has, but substantial. So I think you're right that some of the proprietary advantage is not as large. But maybe the more pessimistic take on it is that it's because no one has the data, so the companies don't have the data, because no one does. The control of an autonomous vehicle and the control of a robot arm or some other form factor, but are they different fields? I mean, when you're working on these models, are you thinking also about their application in autonomous vehicles?
33:17So traditionally, these are extremely different problems. But what we're increasingly seeing is a degree of consolidation in the sense that very similar building blocks can be reduced. So I think actual autonomous driving is maybe one of the tougher things because of all the constraints and regulations and all that stuff. But for small-scale mobile robots, think like drones, sidewalk robots, et cetera, we already have research projects where we've developed vision-based navigation policies for these things that use essentially the same exact architecture as what we use for the robotic manipulation problems.
33:50And a very natural next step is to actually combine not just have the same architecture, but literally the same model. So in principle, at this point, there isn't really any technological obstacle to doing that. Now, of course, there's a lot more to driving, let's say, an autonomous car than just avoiding obstacles and reaching the destination. You have to put in a lot more knowledge, constraints, all this other stuff. And that probably is rather specialized. But my hypothesis is that we'll probably actually see a lot of consolidation around the same basic building blocks for the core kind of perception, action kind of system within these things.
34:22And then maybe where they would differ is the kind of planning layer that sits on top of that and then directs it in terms of what to actually do in a given situation. In your work, there is a pull for academics because of the compute constraints and money, salary and that sort of thing to pull people into the industry. Are you working, straddling academia and industry? or you firmly? So I spend 20 % of my time working with Google DeepMind. I think that in terms of the degree to which industry research versus academic research in robotics is more or less appealing or progressive, I think probably it's a little more, I would say, tilted towards academia than, for example, natural language processing or vision.
35:15And maybe in part it's because there are a lot more of the kind of big questions to figure out before things actually have revenue, so to speak. Like, you know, you could build a language model or vision system that provides an actual business case today, whereas the analog and robotics is probably still a few years out. That said, I do think that it's, you know, there's a lot of rapid progress. And certainly a lot of students from my group are excited about starting companies based on technologies they're developing and things like that. So I think we'll see that kind of catching up in the near future.
35:45Yeah. And you think that, you know, this is the year that AI kind of hit the public sphere and people confuse robotics with AI all the time. Will the day come? I mean, obviously the day will come, but when do you think the day will come that there will be some commercial application or open source application that is adopted by the public that people will suddenly be talking about robots as opposed to AI? Yeah, that's a complex question because I think that if I had to guess, I would guess that a lot of what's needed beyond the core technologies is a pretty substantial upfront investment to sort of overcome the activation energy to get something like that to be practical.
36:42And that's not too unprecedented in the sense that more or less the same thing happened with language models, right? The core technology for next token prediction is pretty old. what was needed to develop the kinds of technologies and products that really capture the public imagination is a large investment of effort into engineering them really, really well, and curating, collecting, and assembling the right data sets to get them to work really well for a thing that basically anybody could grab and use. Partly there's a scientific question there, but a lot of it is really sort of organizational economics kind of questions.
37:19And the trouble with those things is that they're, I think they're hard to predict because they have more to do with like the point in time at which people decide that it's time to lay out those big resources to make it happen rather than just predicting when the technology will evolve. So, you know, the technology might evolve steadily, but then the inflection is really the resource allocation. So I don't know, I can't predict when that would happen. My, you know, if I had to bet, my bet would be like closer to five years than to 10, but I'm not sure. There's been this threat debate has created a lot of acrimony in the community.
37:56Do you have a view on that or is your territory removed enough that you don't engage in that? Yeah, it's a complicated question. I tend not to prefer not to get too much into discussions like that. Partly because I'm not really sure exactly how things will go. And I think that partly, perhaps as a roboticist, I maybe tend to be a little more pessimistic about where we're at in terms of overall AI systems. Kind of hard to imagine that an AI system that can't control a robot to do basic things that are easy for humans would be all that capable overall. But this stuff the stuff is hard to predict.
38:42And I think that maybe like the one constant in AI research is that people are continually surprised by the things that turn out to be easier than imagined and also things that turn out to be a lot harder than imagined. So if you go back a couple decades, it would be pretty shocking to think that, for example, artists and writers would feel threatened by AI systems long before the gardeners and cleaners would. But that is the world that we live in today and you know maybe that tells us to be a little humble about our predictions. Yeah, yeah that's right. And the governments around the world are very focused on regulation of generative AI specifically.
39:26Is there regulation, are governments looking at robotics or AI and robotics in the same way? And is there government support? There's been a lot of talk about providing compute resources to research and smaller companies so that that's not held within these big tech companies. Is there that kind of talk in robotics that the government should or could provide more resources to accelerate research. There's certainly plenty of talk about that. It's typically not, from what I've seen, not something that sort of tends to separate out robotics versus AI versus other things. There's certainly talk about it.
40:16I haven't seen a great deal of action yet, but I imagine it's something that moves slowly. So, yeah, I don't think I would say anything differently than any other AI researcher in that respect, And I don't think that from what I've seen so far, there's anything that kind of treats robotics in a particularly special way in that regard. But yes, this is a big issue. And it's probably something that we, you know, certainly in the United States, we need to think carefully about how we're going to maintain our technological edge and how we're going to allocate the resources that are necessary for that.
40:46And that leads me to another question, because I spent a lot of my life in China. Where is China in this research? Do you think they're ahead, behind? I'm not sure exactly. I mean, you know, one thing I will say is that I think that researchers from China, from Chinese universities, have been very successful across all areas of AI, including robotics. And certainly a lot of really interesting research in Greece and I do see coming out of China, for example. When we were doing a lot of our dataset collection work, we were actually very surprised to learn in the middle of it, that there was a really amazing data set that was released by some researchers from Shanghai that was comparable in size and scope and diversity to the one that we were collecting.
41:30And that was wonderful. They released it open source. I talked to them on the phone. They had really interesting thoughts about what they wanted to use it for. So I am seeing a lot of uptick in terms of the quality and the kind of results that are coming out. The other interesting thing is that actually a surprising number of advances in hardware have actually been enabled by companies out of China. For example, one of the most widely used platforms for quadrupedal locomotion research is produced by a company called Unitree from China. And I think a lot of the things that make that platform so appealing are that it's relatively simple, it's affordable, and it's designed in a way that makes it easy for researchers to get kind of into the guts of it.
42:17And that also, in my opinion, has actually been a very good thing because while we might be concerned about competition and all that, in the end, it's actually accelerating research here in the United States. So that's what I've seen so far. I mean, without trying to make any valid judgments on what's good or bad, it seems like there's a lot going on. That's it for this episode. I want to thank Sergey for his time. If you want to read a transcript of today's conversation, you can find one on our website, I on AI, that's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but AI is changing your world, so pay attention.
From the publisher
Join host Craig Smith on episode #176 of Eye on AI as he dives deep into the realm of robotic artificial intelligence with Sergey Levine, associate professor in the Department of Electrical Engineering and Computer Sciences at UC Berkeley.
In this episode, Sergey unveils the latest advancements in AI control of robots, exploring the implications of reinforcement learning and the concept of embodied AI.
Discover how Sergey's research is pushing the boundaries of AI, enabling robots to learn manipulation skills and generalize across diverse tasks, transforming the potential of home robots and beyond.
Sergey also shares insights into the RTX project, an ambitious collaboration designed to achieve remarkable generalization across different robot morphologies, enhancing robots' ability to perform language-conditioned manipulation tasks.
If you're fascinated by the intersection of AI, robotics, and the quest for creating adaptable, generalizable machines that promise to revolutionize our interaction with technology, this episode is a must-listen.
Remember to rate us on Apple Podcast and Spotify if this episode ignites your interest in the dynamic field of robotic AI and the visionary work of Sergey Levine.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction to Home Robots and AI in Robotics
(01:43) World Models and Language Models in Robotics
(04:01) The Challenge of Learning-Based Control and Data Utilization
(06:05) RTX Project: Generalizing Controllers Across Different Robots
(10:09) Uniformity in Model Architecture Across Labs
(13:50) Introduction of RT1 and RT2 Models for Robot Control
(16:06) The Future of Robotic Control Research and Architecture
(18:49) The Impact of Hardware Development on Robotics
(22:15) Advances in Controller and AI Model Development
(26:21) Planning and Acting with Vision Language Models
(31:38) The Proprietary vs. Open-Source Debate in Robotics
(36:23) The Future of Commercial and Open Source Robotics Applications
(40:59) The State of Robotics Research in China




