#173 Vincent Vanhoucke: How Is AI Helping Advance Robotics?

9 Mar 2024 · 33 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye on A.I. Podcast Episode #173 Summary

Episode Title

Vincent Vanhoucke: How Is AI Helping Advance Robotics?

Episode Overview In this episode, host Craig S. Smith interviews Vincent Vanhoucke, a senior director for robotics at Google DeepMind. The discussion centers on RT2, DeepMind's innovative robotic transformer system, which combines visual perception, natural language understanding, and action planning to advance robotics. Vanhoucke shares insights into the architecture of RT2, its capabilities, challenges, and the future of AI in robotics.

Key Highlights

  • Introduction to RT2: A multi-modal model capable of processing both language and visual inputs, allowing robots to understand and execute tasks with flexibility.
  • Vincent's Background:
  • 16 years at Google, initially in speech recognition and then pivoting to focus exclusively on robotics.
  • His foundational work in deep learning led to the establishment of Google Brain's robotics efforts.

RT2 Architecture Overview

  • Multimodal Approach:
  • RT2 merges capabilities of large language models (LLMs) and visual language models (VLMs) to create a single model for robotic control.
  • Trained on paired data of images and actions, enabling it to "speak" robot actions similarly to language.

Key Concepts in Robotic Control

  • Semantic Planning:
  • Instead of geometric paths, RT2 formulates plans in semantic terms (e.g., "make coffee") to simplify complex tasks.
  • Reinforcement Learning:
  • RT2 primarily uses supervised learning; however, reinforcement learning can enhance its performance by providing feedback based on task success.
  • Hallucinations in AI:
  • RT2 mitigates hallucinations by grounding actions in real-world sensory data, which helps maintain context and accuracy.

Future of Robotics with RT2

  • World Models:
  • The concept of using generative models to predict outcomes based on different actions, allowing robots to "imagine" potential futures based on their decisions.
  • Data Sharing Across Robots:
  • RT2 demonstrates that pooling data from different robots leads to better generalization and performance, similar to advancements seen in computer vision.

Open Source and Community Engagement

  • Open Sourcing RT1:
  • The smaller, portable RT1 model has been open-sourced to encourage community collaboration and development.
  • Diverse Hardware Integration:
  • RT2 aims to be hardware-agnostic, focusing on creating intelligent systems that can integrate with various robotic platforms.

Conclusion The episode underscores the transformative potential of AI-driven robotics with DeepMind’s RT2. As the technology progresses, collaboration and data sharing among different research institutions appear vital for further advancements. Vanhoucke expresses optimism about the future developments in robotics, driven by continuous improvements in AI models.

Additional Resources

  • Sponsor: Netsuite by Oracle
  • Offers insights into streamlining financial systems for businesses.
  • Stay Connected:
  • Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
  • Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

Important Timestamps

  • 00:00 - Preview and Introduction
  • 04:15 - Vincent's Background and DeepMind's Robotics
  • 06:04 - RT2 Architecture Overview
  • 09:06 - Semantic Planning in Robotics
  • 12:33 - Reinforcement Learning and Mitigating Model Hallucinations
  • 15:47 - World Models and Planning
  • 19:29 - Integrating Sensory Data for Planning
  • 24:08 - The Future of Robotic Control and RT2
  • 27:35 - RT2 Open Source and Hardware Integration

This episode provides deep insights into the intersection of AI and robotics, illuminating how current advances are reshaping our understanding and capabilities in robotic control systems.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00RT2 is a multi-modal model, so very similar to Gemini in spirit. and it can understand taken as an input both language and images and we've trained it on robotics data which is paired data of you know images and actions to basically be able to not just speak english but also speak robotese if you will basically we treat robot actions as merely another language that this chatbot can speak. And that has a lot of benefits from an architectural standpoint. That means the model can take advantage of the common sense understanding that a large language model has. A large language model knows that you can put a cup on the table, but putting a table on the cup doesn't make much sense.

0:59Hi, I'm Craig Smith, and this is Eye on AI. Hi. Household robots are getting closer to reality, and in this episode, I speak with Vincent Van Hook, a senior researcher at Google's DeepMind, about their work on robotic control using large language models and visual language models to create something that they call visual action models. We discuss the architecture of their robotic transformer system, RT2, which combines visual perception, natural language understanding, and robotic action planning into a single multimodal transformer model. Vincent explains how this approach allows for common sense reasoning, generalization across different robotic platforms, and the integration of world models for future action planning.

1:54I hope you find the conversation as amazing as I do. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud. Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application, development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle.

2:48So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, Databricks, Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Well, Vincent, I'm really, really delighted that you could join me. I have been talking to people about AI control of robotics for quite a while. I talked to Sergey Levine and Peter Abil, and I've spoken with Alex Kendall at Wave, who's working with the world models. But I was particularly taken by the RT2 visual language action model.

3:51And I wanted to understand that and understand how it compares to things that have gone before. But to start, could you introduce yourself, give your educational background, how you got to DeepMind? I know that you've only been on the robotics team for a year or so. Not quite. Not quite. I can give you a bit of a background. So I've been at Google about 16 years now. I started working on speech recognition originally, and then caught the sort of deep learning bug and started working on all areas of AI and machine learning. Worked a lot on computer vision. And then asked myself at some point, you know, imagine a future in which computer vision really, really works well.

4:50What are the consequences for the world? And it became very clear to me that the main consequence was that robotics would be changed forever and that actual robots operating in the real world would be a possibility. So about seven years ago, I pivoted to exclusively focusing on robotics. And I founded the robotics effort in Google Brain at Google and started sort of developing. What at the time was really a niche research agenda. How do you use deep learning and machine learning, the modern versions of machine learning to control robots? It was not an obvious path at the time. By now, it's a lot more well-established and there is a lot of people working on the topic.

5:47the last iteration of which is how do we use the latest kind of revolution in AI, the large model revolution, large language models or large multimodal models to benefit robotics. Right. And RT2 is robotics transformer. So these are transformer models. And it's, my understanding is that it's a combination of a large language model, a visual language model, and reinforcement learning. I mean, you talk about reward and, you know, thumbs up and thumbs down and that sort of thing. Can you walk us through that architecture and which are the new elements? I mean, the large model is, you know, coming from, yeah.

6:44The simplest version of it, in fact, is RT2 is a multimodal model, so very similar to Gemini in spirit. And it can understand, taken as an input, both language and images. and we've trained it on robotics data, which is paired data of images and actions to basically be able to not just speak English, but also speak robotese, if you will. Basically, we treat robot actions as merely another language that this chatbot can speak. and that has a lot of benefits from an architectural standpoint. That means the model can take advantage of the common sense understanding that a large language model has.

7:43A large language model knows that you can put a cup on the table, but putting a table on the cup doesn't make much sense or that if you have an open bottle, it might spill, but if it's closed, you're probably fine to pick it up. and it has also a very good sense of how to perceive the world at large right it's trained on all a lot of internet images and so it has semantic information about everything that the can you know can be perceived in the world and what we've seen is that by treating this as a unified model that does perception, semantic understanding, and action all at once, we are able to transfer those capabilities between the perception model and the actuation.

8:36And so suddenly the robots understand objects that it's never seen before. It understands concepts of making coffee that it never has encountered before in a way that transfers very well to the actuation. In RT2, there is actually no reinforcement learning involved. It's entirely trained using supervised data. So very similar to how you would train a large language model, essentially. Yeah.

9:09The reasoning, you need a fairly, I'm not sure the technical term, but you need a fairly long chain of reasoning for it to plan. Large language models themselves don't plan very well. So where does the planning come in and how do you refer back? One thing that's very interesting is that we can uplift the planning to happen essentially in semantic space. So if you want to make a plan for robots, if you, for example, you want your robot to make you a cup of coffee, the number of steps involved is very long. You have to sort of drive your robot in certain directions and then move the arm. It's got lots of degrees of freedom that needs to get right.

10:04Right. So you would think that it's a very long planning problem. But if you think about the problem at the semantic level, how do I make a coffee? You can ask your favorite chatbot, you know, and prompt it and say, you know, I'm a robot. How would I go about making coffee? It will give you some steps that are very plausible, which will to, you know, find the coffee machine, put the cup in the coffee machine, press the button. That's actually a very compact representation of what it means to make coffee. And so you can take that, use it as a semantic planner, like you would be as if you were walking a person through the steps required to the task.

10:51And then for each of the steps, then you can decompose that into actual actions. And that's what RT2 does, is taking this kind of low-level drive to the coffee machine and turn that into a low-level action that is understandable by the robot controller, essentially. And those actions, when it translates it into actions, because it's a transformer model, So it's breaking the actions into tokens. Is that right? Yeah, so each token represents basically a configuration of the action space of your robots. So it's basically literally the joint angles of your robots or the position of the gripper of the robot.

11:44Or you can imagine different ways of representing that action space. But at the end of the day, you can think of it as it's a program that instructs the robot. And those language models are really good at coding, right? We've seen that they're very good at producing sensible code. So this is just another instance of very specific code that's tailored to a robot that gets produced by the language model and can get translated into actions. And we do that in a closed loop so that every time the robot does an action, you know, we check the state of the world after that action and we replan and re-decide on what the next steps are.

12:33Yeah, well, that's where I thought reinforcement learning might come in. I saw you were talking to TechCrunch and you talked about giving the model a reward. Maybe I'm just confusing. No, no, no. It's totally correct. You can do that on top of it. So in the case of RT2, we did not do much reinforcement learning. But the next natural step is indeed having a reward, basically checking what the robot does, giving it a reward based on whether it succeeded or not and feeding that information back to the model in the general world of large language models you have supervised learning and then you have rlhf on top of it right it's it's it's the same kind of uh dichotomy where you you first fine-tune the model and then on top of that you use a reward model to make it even even better at the specific tasks that you care about.

13:37And how do you get around hallucinations? Because in both the VLM and the LLM, hallucinations, you can suppress them to a point through fine tuning, but you can't eliminate them. For example, the tokens you were talking about, the specific angles in the actuators in the robot arm, for example, are those encoded in or embedded in a vector database so that the models are retrieving something very specific rather than just a vector. generating them from their weights? No, but one thing that the robot has that a typical LM doesn't have is it has the real world in front of it, right? So if your robot is in your kitchen, it sees your kitchen.

14:43It doesn't hallucinate a kitchen. It's right there. The real world is really there to ground the actions and it's very it makes it a lot harder for it to hallucinate objects that would not exist but it's still technically possible that you know it will hallucinate that there is a cup of coffee on your table when there is not but we have this constant reinforcement of the robot you know seeing the environment it's in and repeatedly getting feedback about what is there, what are the consequences of its actions. And so it really mitigates this issue that other models have that are not grounded in any sort of specific real-world context where you actually have an anchor point that it's the real world.

15:35You cannot change that. You cannot hallucinate that. It's reality sort of is the reality check in a sense for the robot. Yeah. And that's why I've been interested in world models in Yen Le Koon's work and in Alex Kendall's work. because it's unlike LLMs that are grounded in knowledge encoded in text, which is one layer removed from reality. They perceive reality directly and then create a representation of that reality and then do planning at the representation level. Does that happen in your model? I mean, how does it differ from, for Gaia 1 with Waves AI, for example, my understanding is, you know, it's got input data coming from its cameras.

16:42it creates a representation and then it can plan in the representation space. And then there's a decoding to a control mechanism. But in your case, in the context of self-driving car, the world can be sort of, and I don't know what WAVE does specifically, but in general, it's a geometric problem where you have to avoid collision with pedestrians. You have the rules of the road. So you can distill the world into a geometric representation and plan in that geometric representation. In robotics at large, it tends to be a lot harder to come up with the simplified or distilled representation of the world in which you can plan.

17:32And so that's why we've kind of focused more on the semantic planning, right? So we're trying, again, taking the example of, you know, making a cup of coffee, I'm going to formulate a plan of how to make this cup of coffee, not as a geometric plan, like go to XYZ coordinate, but as a semantic plan, like go to the coffee machine, and then let the low level system understand this. But there is a place for world models. And in fact, the fact that we now have very good generative AI opens up the possibility of creating generative world models that enable basically the robots to imagine the future conditioned on the actions that they're taking.

18:24So you could imagine if you have a perfect video generative model, you can take your robots and say, imagine the future if I press this button or if I go in this place and have the video generative model generate an imagined future of what could happen if you do this. And in fact, we've done some experiments where we do generate these futures over a certain horizon and then look at the outcome and decide is this a better outcome or is this closer to our goal or not. and use that as a way to guide the planning process. It's literally the robot is dreaming of its future and or possible futures, depending on which actions gets taken and plans using that imagined future to make its decision.

19:24I think we're going to see a lot more of this in the future. Right. And that would be then combining a world model with the VLA. Is that right? And just forgive me, I'm a journalist, but that's part of the role is to bring it down to layman's understanding. So you've got a robot with a camera and touch sensors and other things, the visual and all of the sensory data comes in, and then the VLM interprets that and recognizes the objects and recognizes the robot's position in that and encodes that in tokens. And then the LLM is the interface with the human. The LLM is the interface with the human, but it's a lot more than that.

20:38And that's one thing that is fascinating about those large language models is that you think they're about language, but they're really about, a lot of it is about common sense reasoning. They have a lot of information about the world. They know about the human world indirectly, essentially through reading all the textual data that's available on the web. And that common sense reasoning was one of the hardest things to get access to or to be able to manipulate in AI in general. And for robotics in particular, understanding this common sense is really, really critical if you want to do any form of planning in a real world environment.

21:29So the LLM is really this repository of knowledge about the human world that the robot can draw from and plan against. So it's not just language, it's also about understanding. And that's what makes it very exciting from that standpoint. Yeah. So when you give the system, the VLA, an instruction in natural language, the LLM breaks it down into subtasks to execute that language. And then what's the link then to the VLM, to telling the robot how to move its actuators to affect the action? The VLM and the LLMs are all connected through those token representations. and so the vision, if you will, is also an input to the system and you take an image and you turn it into a stream of tokens.

22:39You take the command from the person and take another stream of tokens. In fact, you just concatenate them all into one long string of things and then the transformer takes those tokens and turns them into tokens that are in roboties, essentially, that basically are action tokens that get translated by the hardware into actions. So it's all trained together. It's all one big model, one big transformer model that handles both the perception, the reasoning, and the action at the same time. And the fact that this is a joint model really means that the robot can reason about the semantics of things, the vision aspects and its own embodiment, its own proprioception, if you will, in one shot.

23:33Very much like what human would do. You don't necessarily segregate your understanding of the vision from your understanding of the semantics of the world. It's one common model. And what's interesting is that that model can be trained on text data, image data from the web, robot action data from all robot experiences. And all of that gets merged into something that is more than the sum of its parts, if you will. yeah and and the point of this is to create a robot brain so to speak that can generalize uh in in eventually almost any uh environment or or situation you I mean you've collected a lot of data globally from different labs.

24:32I was talking to Sergey and then Peter Chen at Kovarian. They were talking about networking with robots in different form factors in different environments globally, and all of that data coming back to train a model. Are you doing that? Yeah, so this is this RTX project. One thing that was very surprising when we started working on this RT2 project is that we found that, you know, you could imagine that every robot is different, has different embodiments, different degrees of freedom or different configuration, and that you would want to train a model per robot because they're very different. What we found is that the different robot languages, if you will, that you had to produce are mere dialects of each other and that you don't really need to train a separate model for every robot and that if you train a model that goes across robots, you actually get some benefit from each of them.

25:45Like there is positive transfer between the different robots. So you add one additional robot to the mix with its own experiences, its own data, its own embodiment, and everybody benefits. And so with the RTX project, we try to stress that hypothesis by going to 20 plus different institutions, asking them, hey, let's pull together all our data. And, you know, don't curate anything. Just take the data you have. It's not, you know, the cameras are in different positions. That's okay. The robots are different. That's okay. Your tasks have nothing to do with each other. It's okay. Just pull everything together.

26:25Let's train a big model on all this data. And then we sense the model back to the institutions and ask them to evaluate so that the evaluation would be fair and we wouldn't have a hand in it. And they got uniformly gains from using this bigger, more general model. So this is super exciting for everybody in robotics because in computer vision, we saw that collecting large amounts of data from various sources really improved the vision model. It wasn't obvious that you could do that in the context of robotics. Now, we have at least one proof point that pulling data and resources together could improve things.

27:11I think that will change how people think about robotics data. It will enable more collaboration between the different labs. I think it will have a very positive virtuous cycle on the community. And we really want to push this forward because I think that could really have a very nice potential for everybody involved. Yeah, I mean, it's exciting to watch. It's the most exciting thing I've seen in robotic control to date. Is it open source, RT2? Or how can people work on this? RT1 is open source. This is a model that's smaller and more portable and that we've just open sourced the training code for it.

28:04and we have also checkpoints and models for it so the community can leverage it and make derivatives of it and play with it. And we're going to continue sort of open sourcing a lot of those models because I feel like we're in the very early days of this revolution and we really want all the community to build up from those models. To Boston Dynamics, who, at least from the public view, has the most advanced hardware, are you able to integrate this with hardware systems like that that are much more versatile? and, you know. So Boston Dynamics is not part of Google anymore. So we don't have access to their robots.

29:06They're a fantastic team and they are very impressive robots. We have a number of robots that we work with. To some extent, we try to work with a very diverse set of robots because we have this hypothesis that the diversity of experiences and of embodiments is really the key to unlocking the potential of those kinds of models. So we try to be very hardware agnostic to some extent and bring in a lot of partners into the mix and a lot of various diverse form factors. We're really focused on the brain aspects and bringing the intelligence to bear. Okay, and last question. What's the next? Are we going to see another announcement in three to six months?

29:57Where are you going with this? The entire field is very greenfield right now. There is a lot of directions to explore. And you see what I think you can expect is that the progress is going to mirror the progress you see in large multimodal models like Gemini. And those are evolving rapidly. And the capabilities that are enabled by those models are evolving rapidly. What we've shown with RT2 is that robotics can track with those improvements because we're basically a thin layer on top of the advances that those models provide. And so we're going to expect a lot more in that direction in the future.

30:43Okay. Okay, Vincent, I really appreciate the time. It's one of the most fascinating areas I'm reading about. If we could have another call at some point, I would love it. But I'll let you go now. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs.

31:34OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, Databricks, Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I all run together. That's it for today's episode. If you want to read a transcript of the conversation, you can find one, as always, on our website, IonAI, that's E-Y-E hyphen O-N dot A-I.

32:28In the meantime, remember, the singularity may not be near, but A-I is changing your world, so pay attention.

From the publisher

Join host Craig Smith for episode #173 of Eye on AI, featuring a groundbreaking discussion with Vincent Vanhoucke, distinguished Scientist, and Senior Director for Robotics at Google DeepMind. 

In this episode, Vincent dives into the cutting-edge world of robotic control with RT2, DeepMind's innovative robotic transformer system. Learn how RT2's integration of visual perception, natural language understanding, and action planning is pushing the boundaries of robotics, enabling machines to reason and act with unprecedented flexibility and intelligence. 

Discover the challenges, technological marvels, and future prospects of robotics as we explore how DeepMind is shaping the future of AI-driven automation.

Tune in to uncover the potential of RT2 in transforming robotic capabilities and the implications for diverse applications across industries. 

Don't forget to rate us on Apple Podcast and Spotify if you find the episode insightful!


This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.

Download NetSuite's popular KPI Checklist, designed to give you consistently excellent performance - absolutely free at https://netsuite.com/EYEONAI


Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI


(00:00) Preview and Introduction 
(04:15) Vincent's Background and DeepMind's Robotics
(06:04) RT2 Architecture Overview
(09:06) Semantic Planning in Robotics
(12:33) Reinforcement Learning and Mitigating Model Hallucinations
(15:47) World Models and Planning
(19:29) Integrating Sensory Data for Planning
(24:08) The Future of Robotic Control and RT2
(27:35) RT2 Open Source and Hardware Integration

More from Eye On A.I.

All 266 episodes
#173 Vincent Vanhoucke: How Is AI Helping Advance Robotics?Eye On A.I. · 33 min
Listen in VO