Introducing Gemini Robotics 2

2 Aug 2026 · 39 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Google DeepMind’s “Gemini Robotics 2” launch, focusing on an embodied reasoning model (ER 2.0) plus “action models” to give robots whole-body intelligence, dexterous manipulation, and multi-robot collaboration. They argue robots are still 5–10 years from everyday use due to unsolved physical AGI problems, especially dexterity.

Guests

Logan Kilpatrick (Google DeepMind, host). Carolina, Stewart, Kanishka, and Jay (DeepMind team members in the Gemini Robotics Lab; they discuss model architecture, dexterity, data, and deployment).

Key claims

ER 2.0 enables robots to understand where all body parts are in 3D and reason about complex tasks; dexterity remains the hardest unsolved area (contact-rich, ~20+ hand degrees of freedom). Progress is limited by lack of “internet of physical interaction” data; teleoperation and wearable/egocentric data have embodiment gaps. Deployment likely starts in semi-structured industrial settings before homes.

Notable examples

“Please clean my garage” (put away items, avoid obstacles, handle fallen objects); dexterous demos packing lunch (grapes into Ziploc), unscrewing a bulb in a sphere, tying knots/trash-bag tie-off; multi-robot parallel work; cross-embodiment transfer across robots like “Franca Duo” (two arms). Availability: ER 2.0 via AI Studio/Gemini Enterprise Agents Platform API; action models via deep partners and a trusted tester on-device program.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenge of Robotics

0:00 to 0:38

Exploring the complexities and timelines of integrating robots into daily life.

“So actually I was get asked this question like when do you feel robots is going to enter our daily life?”

Introducing Gemini Robotics 2

0:44 to 2:15

Overview of the new embodied reasoning model and its capabilities.

“Robotics Lab with Carolina, Stewart, Kanishka, and Jay.”

Capabilities of Gemini Robotics 2 Models

2:15 to 4:49

Discussion on whole body intelligence, dexterity, and multi-robot collaboration.

“Do you want to talk through sort of the different models, the suite of models that are becoming available?”

The Evolution of Robotic Techniques

4:49 to 7:37

How general-purpose robotics has evolved and the breakthroughs it has seen.

“I have a home robot vacuum and then I have another robot that I have ordered.”

Gemini Robotics and AI Integration

7:37 to 9:20

Exploring the integration of Gemini's multimodal world understanding into robotics.

“And so last year what we did was we brought all of the power in Gemini's multimodal world understanding combined with those techniques in order to bring essentially Gemini's intelligence to robots.”

Challenges in Dexterous Manipulation

9:20 to 11:40

Examining the complexities and unsolved problems in dexterous manipulation for robots.

“What's the place to actually hill climb to get to the place where you start to see robots folding laundry successfully in most cases?”

Data Collection and Quality in Robotics

11:40 to 14:00

Discussing the importance of data in improving robotics capabilities and the challenges faced.

“Maybe there's not richness in like the labeled format that we actually need to make progress.”

Challenges in Data Collection for Robotics

14:00 to 14:48

Understanding the difficulties in collecting physical data for AI models.

General Purpose Robotics vs. Specialized Tasks

14:48 to 17:08

Exploring the balance between general purpose and specific robotic capabilities.

“like our data sets are so tiny compared to the other digital token sets that, yeah, they don't play well yet.”

Deployment Challenges for Robotics

17:08 to 19:16

Discussing the complexities of deploying robots in real-world environments.

“And you have to check, I think we always found it's the opposite word.”
Show all 17 chapters

Future of Robotics and Dexterous Manipulation

19:16 to 22:20

Looking ahead at the advancements needed for robots to integrate into daily life.

“We are going to see over the next two years if that changes a lot.”

The Path to Physical AGI

22:20 to 27:00

Examining the relationship between digital AGI and physical robotic intelligence.

“So maybe there's like one or two of these and then yeah, then the tale of the things that need to be solved will get to and yeah, we'll have robots around us and maybe the cities will start looking different.”

Demonstrating Dexterous Capabilities of GR2

27:00 to 28:00

Showcasing the advanced dexterity features of the Gemini Robotics 2 model.

“I think that even in the past release, we're starting to see the very first times where the embodied reasoning model can actually watch the robot do something and have opinions about it.”

Understanding Robot Dexterity

28:00 to 30:18

Learn about the challenges and tasks involved in programming robot dexterity.

“And the cool thing is that we train GR to control the whole body and the hands with the same recipe.”

Cross-Embodiment Transfer in Robotics

30:18 to 32:07

Discover how robots can transfer learned tasks across different embodiments.

“And so in this example, this is like a partner hardware that we've generalized the ER model to be able to work on that specific set of hands in that context.”

Programming Robots with Data-Driven Approaches

32:07 to 36:29

Explore the evolution of programming robots using data-driven models.

“And fundamentally, if you're actually doing a task where you're organizing things, I mean, 90 % of that task is not about exactly how you move your hands.”

Embodied Reasoning Model and Its Applications

36:29 to 39:00

Learn about the Embodied Reasoning model and its capabilities across various tasks.

“Yeah, I mean, these models are better in a few ways.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Robotics is incredibly hard. Yeah, a good t-shirt. Robotics is incredibly hard. Yes. So actually I was get asked this question like when do you feel robots is going to enter our daily life? So if you ask me like three years ago, I would say probably beyond my lifetime. Interesting. If you ask me two years ago, I said maybe 10 years. So if you ask me now, I think that it's between five to 10 years. So you can see the speed of, you know, evolution of this technology is amazingly fast.

0:37How's it going everyone? My name is Logan Kilpatrick. I'm part of the Google DeepMind team. Welcome back to Release Notes. Today we're here in the Gemini Robotics Lab with Carolina, Stewart, Kanishka, and Jay. We're talking about the new embodied reasoning model and actually the whole suite of Gemini Robotics launches that are sort of coming out. So I'm super excited. I have a million questions. We're talking off camera, but Carolina that maybe you can kick us off with sort of just like the headline for this moment. And then we can actually take a step back after that and like talk about the arc maybe across all the work that it's taken to get here.

1:11And we'll get to see cool robots, which I'm excited about. Yeah, definitely. What we're building here is the intelligence layer to power any robot to do a broad range of useful tasks. And we've been building towards this moment for a very long time. Actually, our team has been working towards general purpose robotics since the inception of the team. And we had a long history of always thinking, how do we solve the problem from first principles in a way that doesn't just take shortcuts and tries to solve the entire problem? And so the entire problem really means infusing a robot with like human level intelligence, right?

1:46So that it can understand the environment as you and I can, so that it can reason about what it means to complete a task. And then it can actually take action with all the dexterity that us humans have and take for granted. So that's the general goal. I love that. And the actual release that we're talking about today is the new package of Gemini Robotics 2 models, starting with the ER model, but a bunch of other stuff as well. Do you want to talk through sort of the different models, the suite of models that are becoming available? Yeah. So Gemini Robotics 2 essentially is bringing whole body intelligence to robots.

2:24so what that means is that we're enabling a model that can understand where the entire every part of the robot is in space and you can reason about doing more complicated tasks so like imagine that you get in your closet and you're trying to clean it up and put away all your clothes and your shoes like that requires you to move your body in all kinds of ways avoid obstacles reach things high pick things from the floor this is something that we couldn't really do before in a way that understands what's going on. So we literally like put the robot in a garage and ask it like, can you please clean up this garage?

2:57Right? Is that the real prompt in those examples? It's literally just like, please clean my garage. Yeah, you'll get to see it. Inspired by our own needs, I think. I love it. So then the robot needs to reason about what does it mean to clean up a garage? It needs to like think about, oh, I'm going to put all the cleaning supplies in the same place. It needs to be able to like put things high up high. If something falls on the ground, It has to understand that and then go pick it up and then put it away. So that's one of the big things that we're bringing in this release. The second thing that we worked significantly on and we still have a lot more work to do is dexterity.

3:32So, again, us humans take for granted this like really dexterous hands that we have. But pretty much everything you do every day requires dexterity. And that means like folding things, opening doors, just picking anything that will fall out of your hands. and so dexterity is an area that we worked on a lot and if you just again if you just get yourself in the morning coffee like that is actually a pretty dexterous task yeah and one thing to know is that like because we're building this intelligence layer across robots is we definitely work a lot on making this work really well humanoids but we also bring other robots that we're controlling so the one that has this right behind you is actually our franca duo and it has two Franca arms.

4:15And we also control with the exact same model, this robot in order to do pretty dexterous things with those grippers, like pack things neatly. So that's another form of dexterity. And then the third thing that we bring in is what we call multi-robot collaboration, which is different robots have different capabilities, right? And what we're doing here is that you're bringing the robot, the intelligence to know what it needs to do to complete a task, but also the understanding that it can call all the robots to accelerate the task or to do things in parallel and do things faster. So those are roughly the three things that we're doing in this release.

4:50I'm excited for this. I have a home robot vacuum and then I have another robot that I have ordered. Hopefully it'll be powered by Gemini Robotics too. That doesn't vacuum, but does a bunch of other stuff. And so I feel like this is actually going to be, I didn't think about this today, but I was thinking to myself, I was like, can I get the one to sort of carry the other one around and like, you know, go deploy the robot to go do certain tasks. But But ideally, they could just communicate as robots together and sort of get the work done, which is super interesting. Maybe we can also talk about the arc to sort of get here and sort of obviously the second Gemini robotics model.

5:24But is there like other things that are like worth? Obviously, Google's been doing a lot of robotic stuff. DeepMind, maybe I actually don't know, maybe also had a bunch of robotic stuff, but can talk about any of sort of the research and the arc to get sort of this launch moment. Yeah, I mean, it's been a long path. I think many of us have been here also for a while and have seen all these steps through. But yeah, I mean, we've always from the beginning, like I said, we're really thinking about general purpose robotics before I think the field really realized that that was feasible. And so we've done a few iterations where we've brought sort of techniques that now are table sticks for the community.

6:01So, for example, we first introduced like reinforcement learning to a learning simulation, how to control whole robots. Right. So you see a lot of robots today like dancing and doing acrobatics. Those robots are actually using those techniques in order to move the robots in a way that feels stable and that can mimic a particular sequence. We also shown what's possible when it comes to bringing LLMs and BLMs when it comes to planning for robots. Before that effort, basically robots did not understand semantics. They did not understand our world. They did not know what you meant when you said, bring me a cup.

6:36Yeah. Right. You actually had to say, no, bring me that object at position X, Y, Z in space. So that was one clear breakthrough. We also introduced transformers to robotics. And sort of that shifted the feel into this era of data driven robotics, where you now had to just like collect a lot of data and teach the robot how to do many different tasks. and then introduce even the concept of VLA, which is a new type of foundation model called Vision Language Action Model that essentially enables robots to understand natural language and visual input and then directly control them in a way that is general.

7:10So this VLA type of foundation model is again adopted in the community. And then we've also shown what's possible when it comes to dexterity. I think many of us did not really think that it was possible to tie shoelaces, for example, in our careers. And that's something that today people tie shoelaces, they fold their laundry. These are things that we show that was possible. And I think all of that has sort of come together towards like this Gemini robotics models that we introduced last year. And so last year what we did was we brought all of the power in Gemini's multimodal world understanding combined with those techniques in order to bring essentially Gemini's intelligence to robots.

7:47and all it is is that we enable Gemini to also think about moving robots so we add actions as a modality in Gemini and so that essentially enables Gemini to understand when you ask it to like turn around this it understands what it means to to move around the bottle or if you ask it to do something more complex like pick up all the things that are pink then it understands what that means so that's essentially what Gemini Robotics is and then we continuously improve it And Jimena Robotics 2 is a pretty big step function with respect to our previous one. I love that. I think our team recently thinks about this problem from a frontier model perspective.

8:26So really kind of leveraging. You don't have to solve robotics from scratch. There's all this world understanding in these big models. So the idea is how do you latch on to that understanding and then use the robot for useful things. So you've had this history of using data as a scaling paradigm. But now it's more like these frontier intelligent things. is how do you hook up this like physical thing to those models and then kind of, you know, bootstrap robotics from those? And what ends up being the limitation in practice? Because I'm thinking, for example, like I've seen the demos of, you know, the robots dancing and all this crazy stuff and doing backflips, which I can do.

9:00But then you also see for like AI agents, the sort of like meme canonical demo is like booking travel. And I feel like for robots, it's like doing folding laundry. But yet it feels like also maybe the robots can't actually fold laundry. And maybe that's not true. maybe Gemini Robotics, to sort of close that gap. But I'm curious, there's so much understanding and intelligence baked into the model. What's the place to actually hill climb to get to the place where you start to see robots folding laundry successfully in most cases? Or actually, are we already there? And maybe I'm not fully calibrated on model progress in this regard.

9:36So there is a lot of progress in the last 10 years. So one thing that you mentioned this locomotion or whole body control has reached a new height where you see the humanoid robots that can you know backflip and doing very agile motions. Also our team sort of pioneered a research called reinforcement learning and simptural transfer that makes all those things happen. So it seems that the locomotion is nearly a solved problem. What's remaining is actually a very hard problem is dextrose manipulation, which is that how can you use a hand or grippers to interact with all these objects in order to accomplish the task in your daily life.

10:11The reason that's very complicated is because it's very contact-rich. So you need to think about a lot of contact points on the object in order to move them in a desirable way. And to control your hand, it has over 20 degrees of freedoms. So you need to coordinate all these joints and muscles in order to do those tasks. And all these things are way, way harder than locomotion problems. You only need to control yourself on usually a flat ground or slightly perturbed ground. So dexterous manipulation with all these different objects in real world is really an unsolved problem for robotics for now.

10:47Interesting. Where is like the quality gain come from these days? Is it like we just get more data or it's like new techniques or just like scaling up general purpose models? Like how do we actually, it's an unsolved problem, but like where do we make progress on actually solving it? I think data is a big part of it. Like as Jay mentioned, we're missing this internet of physical interaction data. As I open this cup, there's a sequence of I make a move and the environment moves in response to it. So this kind of sequence of interaction with the physical world, there's no internet of this. So I think data is a key part to this physical AGI component that we haven't unlocked yet.

11:25And that is an open question. How do you collect it? Data quality really matters here. So I think, yeah, just how do we collect the scaled digital version of physical interactions is like is an open thing and we're basically like hill climbing that as like one of the big levers in like unlocking physical agi do we not have the ability to and i'm guessing actually one of the threads that were i'm always talking to folks about is sort of like the back to this building on the world model there's like interrupt between main gemini and sort of some of these like domain specific cases like can we not can you not take like a you know a video of someone unscrewing the water bottle and sort of like intuit some of this data and like get some of it out of those types of, and it feels like there's a very, there's a richness in that type of data.

12:06Maybe there's not richness in like the labeled format that we actually need to make progress. But have we gotten closer to being able to like actually leverage some of the existing non-robotics data to do these tasks? I think you can, if you look at the data progression in the last couple of years, so people usually talk about this data pyramid. On the top of that is, you know, teleoperation data. These are the data where you move some controllers and the robot is going to move accordingly. So this is called teleoperation. These data are very useful because you get all the signals, basically how you should move the robot in different scenarios to train the robot.

12:42But these data are not very scalable because teleoperation is very costly. You mean the human, they are in the robot in the loop and so on. And people say, maybe we need something more scalable. And people think about wearable device. There is something called the Yumi, which is in the academic world where people build these wearable grippers so that humans can collect the data without the robot in the loop. Pretend you're a robot in your house. Yes. I've seen some of these videos. It's interesting. Then those data becomes really scalable. But the problem is human and robots are different. There is this embodiment gap you need to cross.

13:19And down beneath it, maybe the widest base is really egocentric human data. Basically, human, you take a video of human doing things, then hopefully we can learn from those. But again, you don't know how much actuation or how much muscle force you do with each movement of humans. So you are missing a very important label of the actions. At the same time, as I mentioned, that human robots are different. So it's the hardest data to leverage. Of course, we are making progress on leveraging all the entire pyramid of data, but we're not quite there yet. yeah i'm thinking about like in the context of like mainline gemini for some of these use cases where like the models don't really work you start to see like a thousand high quality trajectories like makes a massive difference in like overall quality of the model like do you see that type of like again without getting into the specifics like is it low-hanging fruit free hill climbing everywhere or is it like really you actually need like the the reason tell it because i'm thinking i'm like you know it's google we could we could go get a thousand tele-operated you know examples of like people going in collecting that data but you're saying we need like it would be like the scales of millions of millions in order to get any interesting this is what Germany is pre-trained on the internet right so there's a there's a lot of like nice biases there for language and vision stuff yeah but whenever we add this physical thing into the model like it doesn't play well with the pre-trained stuff so in fact that's when that's one of the reasons why we have our Germany Robotics model is because it is hard to upstream something without killing all the other generalization properties of the model so this physical thing like our data sets are so tiny compared to the other digital token sets that, yeah, they don't play well yet.

14:54So either we scale these up and they start playing well, or there's some other way we can connect them. But it's still an unsolved problem. That's why you don't see these frontier models, like, you know, directly controlling robots. Yeah. You inherit things like natural language understanding, visual understanding. Like, I don't have to pick up a bottle that is black and white and different shapes, like all of that, which we now take for granted. That's true. I didn't even look about that. But before you had to collect every single object. So we do get some generalization, but I think the big part that is missing is really understanding motion.

15:24And motion is not something that you inherited from Gemini today, right? Motion is something that we have to teach Gemini. Or understand force. Or force. These are the things that I think Gemini is not trained on. Yeah. I'm also super curious, like there's obviously like such a distribution of like the actual robotic use cases. And I'm curious for us and for Gemini Robotics, like has there been a focus? Is it like we want to enable home robotic use cases and we see traction? Is it like, obviously, there's a huge amount of industrial automation stuff happening? Is there, is the, I know, Carolyn, you said like, we want general purpose robots.

15:57And so theoretically, you could do all those things. But I'm curious, actually, if there's been like, is it sort of like jagged as far as like progress or capability or like things that are actually working more today versus versus not? Yeah, there is, I think, two answers to that question. I would say in the capability front. We certainly believe that you want to be able to solve a broad range of tasks. And actually by going narrow, you're going to build a policy that is going to be a lot more brittle. It might work better in that environment, but the minute you change anything about the environment, these are going to start to go wrong.

16:28So our approach is certainly like don't compromise on you're trying to solve general purpose tasks and understand general purpose motion and handle a broad range of different objects. I would say in terms of deployment, I think of it a bit differently, right i think it's much more likely that this uh robots are going to be useful in environments like industrial environments that are semi-structured that has safety pretty much uh under control and that you can get a lot of real world experience of what it means to like launch these models and then from there i imagine that we would go to things like retail and other environments that are also starting to get into human-centric spaces but are there less vulnerable than in the middle of like your house with your kids and your pets you're you're raining on my q4 2026 home robot orders right now i'm like i'm waiting i've written on all my chores for 2026 as soon as i start getting some of these robots delivered so it's not yeah i mean there's plenty of people out there that think that they're gonna go home first and i think there's like the appealing thing there is that home first requires that diversity that generalization so like it forces that problem it makes it very front and center.

17:34And you have to check, I think we always found it's the opposite word. Like if you do train on this like one narrow thing, as yeah, it becomes worse. So like having it collect on many different diverse cases, it just helps the general intelligence of it. So yeah, from a strategy perspective and like a learning perspective, it makes sense to go super broad first. Yeah. I mean, then there could be the super fans like yourself that decide to have the robot at home, even if it's like a little early. I hope it doesn't break my stuff. Yeah, I think it is interesting to see like what people's because I assume like these robots will be delivered to people and actually I think it'll be super interesting for all of us like just see what's people's reaction like if it works 90 % of the time and 10 % of the time it's you know cracking a wine glass or something like that like are you happy with that is if it's a three dollar wine glass maybe I have no idea part of this challenge of like scaling up data I'm curious like why actually back to this home example and maybe there's a bunch of industrial examples the tension and you see those sort of opposite extreme of this in the context of like mainline gemini where you know there's billions of people using etc etc we sort of have a flywheel and getting signal from the real world like why don't we see more like real live deployments of robots today in like actually to help us scale getting a lot of this data and in some of these in ways it's just like it's not scalable and i'm thinking back to and maybe this is not right but like you see obviously waymo was doing this for a long time and sort of had cars driving around not being used.

19:00There's actually a ton of other self-driving car startups still doing similar things today and sort of in operation, theoretically collecting data. I don't know what they're actually doing. And maybe I'm wrong about this, but like it feels like that isn't the case of like what is happening in robotics today, at least like visibly to an external observer. And I'm curious why. Do you want to know, Zerla? We are going to see over the next two years if that changes a lot. But I think right now, when we talk to partners about their own experiences, we kind of hear two things consistently. One is while the policies are showing incredibly cool generalizing capabilities, it's actually difficult to get them to be narrowly successful, but broad enough that they can handle all the little things that kind of vary and break over the course of an entire day.

19:42And I think you look back to autonomy, this is a really serious challenge there as well, where you can kind of give a really impressive demo, but you're not ready to remove the safety driver for a long time. And one of the distinctions between autonomy and robotics is sort of this teleoperation option. It's like, and you can just like, I'll just put a safety driver in. And if the car gets stuck, they'll just take over and drive. For a lot of the robots, you know, we saw this with Aloha. We could do that. We could actually get somebody right there. And if the robot got stuck, they could like take over and fix it.

20:07But as the robots get bigger, get more capable, gain more degrees of freedom, which we said actually want in a deployment, it's harder and harder for somebody to jump in and help. And so I think we're seeing this challenge right now where we're struggling to kind of get that bootstrapped system where it's good enough that you're ready to deploy it. And then also we're trying to figure out what is this analogy where if it does get stuck, how do I both fix it quickly and learn from that moment? And so there's little patterns forming. But right now, I think everyone is really focused on like, how do I get this kind of baseline performance?

20:38And so I think there's a big rush in the industry and I get that baseline capability. Yeah. And this is part of what we're doing with our partners, actually. So as you notice, there's many different robot types here and none of them were built by us, actually. So the way we work is that we have these deep partners that we work with in order to accelerate both the AI capabilities and the hardware. And so we're deeply connecting with each other to see what's missing to get to that deployment. And our goals together is to accelerate that deployment as fast as possible and to bring it to real world applications so that we can learn whether what we're learning is actually useful and valuable and where to spend more time next.

21:14But that's absolutely the goal that we have with our partners is how to bring this to useful applications as soon as possible and how to learn from it, how to deploy them in a safe manner so that we can get as much information as possible early on. Yeah. Back to this like two year time horizon, potentially of where we see things changing. I always make those comments to people that like, if you 10 years ago, were sort of like put in a very short term time machine and landed in 2026, if you looked outside minus anywhere where there's Waymos deployed, and you sort of like looked into a city, like, essentially, physically, everything looks the same, like you wouldn't be able to like, maybe you'd spot a new phone that somebody had or something like that.

21:51But like, more or less, the physical world around us looks the same. and they've missed the fact that like we actually have these like extremely intelligent we've like almost not not actually solved intelligence but like made a huge amount of progress on solving intelligence and so it's going to be very interesting to see the physical world around us start to change as like you get these new autonomous systems sort of like doing things in the real world i'm actually curious like assuming that we get some breakthroughs the sort of dexterous manipulation problem is solved like is it like basically a manufacturing problem then at that point to like actually just like scares are still like 50 other things that need to land if we were to like somehow you know we could make dexterous manipulation work really really well everything out like navigation is all like all the other bits of the story have been yeah i mean i think first of all dexterous manipulation is the hardest stuff so we would all be very happy when we when that is solved but my guess is that if you want to get to the point that i think we all dream up which is like you walk around and there's robots doing different things regardless of where you are they're like helping society in a useful way and not just in industrial settings but like i think the second problem we're going to hit is like the human centric aspect of it it's like understanding humans being able to be useful to humans and safe to humans in all that context and that is something that we're also making progress towards i think for them for robots to be in everyday spaces there's all kinds of safety aspects that need to also be solved all kinds of security privacy aspects that need to be solved that are completely parallel and very different to dexterity but i certainly think if we crack the dexterity in a general way it would definitely just like blow up the opportunity of what's possible with robotics i think you will start seeing robots around once you have that problem like we go to robotics conferences and we get a little peek in the future and it is crazy like you see these robots like walking around giving demos so it it feels like you know like star wars the future so i feel like we're a few ways away from that and this dexterous manipulation is like one big chunk of that.

23:52So maybe there's like one or two of these and then yeah, then the tale of the things that need to be solved will get to and yeah, we'll have robots around us and maybe the cities will start looking different. Yeah. I think robotics is incredibly hard. Yeah. That's a good t-shirt. Robotics is incredibly hard. Yes. So I was asked this question, like, when do you feel robots are going to enter our daily life? Yeah. So if you ask me like three years ago, I would say probably beyond my lifetime. Interesting. If you ask me two years ago, I said maybe 10 years. Is Waymo considered a robot or no? No, no, no.

24:29It's a general purpose robot. General purpose robot. Enter our daily lives. So if you ask me now, I think that it's between five to 10 years. So you can see the speed of evolution of this technology is amazingly fast. But there are still a lot of things that we need to solve. I'm curious, actually, though, like five to 10 years away framing. Also, if you sort of like stack that up on framing, are we going to claim general purpose intelligence if you sort of don't have this embodied characteristic or are these two things like completely separate? You have a bias group here. Physically, you can't do AGI until you solve physical AGI.

25:04So if I walked up a robot and say, do anything that I could do, I would expect it to be able to do it. So I think that definition maybe sometimes get lost, but we always live it every day. So like I think robotics we call this like the moral paradox where things that are really easy for humans are very difficult for a boss. Like the A is kind of, you know, pass the bar exam and code up like, you know, all these like operating systems, but they can't cook you eggs or like flip a, you know, burger. So there's some paradox that were like physical AGI, I think is a part of AGI, at least for me, but it is in some ways more fundamentally different than the digital agents.

25:39So it'll take, I think, a bit more work to get that. I think that will end after the digital AGI thing has happened. So I think that's why you had the 2 plus 3 to get to 5. Yeah. And I think it is very possible that getting to digital AGI, and I'm sure it will actually dramatically accelerate the speed to our physical AGI. Not only on the intelligent aspect, but you can also use this in order to build better robots like on the hardware side. Because we haven't talked much about sensors, but everything that we're using today is primarily ignoring all of the sensors that you have in your hand so today like we're just look using vision and that's basically it and the position of the hand in order to determine whether you have picked up this glass but when i pick it up i can feel it in you know all over my hand and that's something that we are not even scratching the surface on today and we think that if you want to be able to do everything a human can you definitely are going to need to have more sensing capabilities than what we have today with robots.

26:44So there's an aspect of also the hardware catching up to get into the level that is capable of achieving human-level behaviors and manipulation. Yeah, skin is an unsolved hardware problem. Yeah, yeah, I can imagine that being true. You talked about this kind of recursive loops. I think that even in the past release, we're starting to see the very first times where the embodied reasoning model can actually watch the robot do something and have opinions about it. And over time, that actually starts to form a real loop. You're like, I think you should go collect a little bit of different data. Or I think you should actually...

27:15And increasingly, the researchers are asking, like, can you please provide an interface by which the higher-level model can actually give, like, guiding instructions to lower-level models? It's actually painful to watch sometimes because the high-level model's like, no, no, just grab it a little higher. And so we're starting to see these more and more. And so I do think as more and more of the kind of core capabilities of, like, Gemini spatial reasoning and things we do with embodied reasoning, start to really get closer to AGI, you will get some of those feedback loops that start to form. Yeah, that's super interesting.

Read the full transcript

27:44Well, let's look at a demo, maybe of sort of the dexterous hands, because I want to see it come to life. Sure, yeah. So here we're looking at GR2, and it's controlling these very high degree of freedom hands. I think there are 20 different joints that it can control per hand. And the cool thing is that we train GR to control the whole body and the hands with the same recipe. So there's nothing special about the hand. It's just like we collected much more diverse, rich, dexterous data, and the model is able to perform the tasks. So let's take a look at some of these tasks.

28:19In order to be useful, a robot needs the dexterity that we take for granted. You probably don't think about how to drive 22 separate joints when you operate your hand, but that's what we're asking these AI models to do. In Gemini Robotics 2, we came up with a set of tasks in order to test and develop dexterity.

28:43So we're asking the robot to pack lunch by putting the grapes into the Ziploc bag. So this requires a lot of precision, but also a lot of coordination. Now the really hard part is getting the Ziploc closed. Very nice job, Apollo! Hello. Hey, Apollo. Can you unscrew the bulb? The bulb is actually in a sphere, right? So those contacts need to be very precise. For it to be engaged with the fingertips, there's actually a lot of motions and a lot of dexterity that's involved when you do that. You have to do it. That was the easy zip lock back too. I'm like, I can't even do the regular ones. That's actually very complicated.

29:20Thank you, Apollo. We're advancing what we can do even with parallel groupers. need to have dexterity, precision, and 3D space understanding. This is a robot that you have behind you. That we are trying to solve. We can move the kit around, we can move the tools around. The robot is going to be able to understand how to reorient the objects in space and then precisely put them in. The robot's task is to tie a knot, tie off the trash bag. Multi-fingered hands are a key ingredient for this kind of intricate knot tying dexterity. Come on robot, you got this. I still have not figured out how to do it.

30:04I've never seen a part like that before. We are pushing our understanding of how robots may interact with complex objects in the real world, like trash bags or like hazardous waste. it would be great if we could send a robot to do that rather than have humans put themselves at risk.

30:29Very cool. And so in this example, this is like a partner hardware that we've generalized the ER model to be able to work on that specific set of hands in that context. So this is the VLA, the action model. The action model. Yeah, so that one is trained to then use these high dexterity hands. So we did collect teleoperation data to see how the task can be done, and then that data helps the model understand how to control these robots. And does the dexterous hand sort of use case generalize? Or is that also something that, as you see across all the different hardware robotic partners, the hands are all different, or the degrees of freedom are different, and that's why it makes it complicated?

31:15So I think that hands are a good place to push the limits of dexterity, but we are seeing some really cool signs of cross embodiment transfer. In our GR 1.5 release, we talked about this more explicitly, where yeah, we are seeing transfer between like the gripper task and like the hands task. So these models, when we train it with all the data, like we don't train per embodiment models, GR 2 is trained on, you know, many, many robots. And we do see these signs of life, like it understands basic concepts and it can transfer that action from one robot to the other. and this is actually really important because i think robots will continue to evolve all the time and we see it even in with you know all the robots that we have every year like they evolve in some interesting way right like and the hands is one of those areas that is like very ripe for a lot of acceleration over the next year so we fully expect the hands to be changing constantly so i think it is really important to enable models that can work across all these different embodiments And fundamentally, if you're actually doing a task where you're organizing things, I mean, 90 % of that task is not about exactly how you move your hands.

32:16It's about understanding where you're putting things. And then the last 10 % is about exactly how you move your hand to achieve it. And so a lot of that transfers between robots. Of course, there is some limitations, like a gripper can only grasp things this way. And a hand could actually do something more complex. but there's a lot of semantics that are shared between them. I'm curious actually if like models having code quality is sort of like at all correlated with some of these use cases and maybe my mental model is like off on this but you imagine like you can like deterministically program robots in certain cases and so could you like is there, I'm curious if we do anything around that or if that's actually like a use case that's helped Like assuming we get like, you know, super intelligence at code, whatever, in the next couple of years because we're hell climbing it.

33:05Does that somehow like help, you know, you could like almost deterministically like program sequences of things that the robots are doing? Or is that not? I can give one example. So I think one place that can really come into play is in simulation. So on the real robot, it's very difficult to write deterministic code. They can actually take in just a raw set of pixels from a bunch of different cameras and actually give you like thoughtful, correct joint angles. So you can play some games with inverse kinematics, but it's really hard. In the simulator, you often have access to privileged information.

33:33So you actually secretly know exactly how far away this lid is from my fingers. And so if you can get to a point where you're really starting to build confidence, where the simulator is actually either a source of data or a place you want to evaluate a policy, now you can basically use that code to guide yourself in much more precisely because you actually have access to really correct resolution, really precise information, rather than forcing that code to kind of interpret sort of the messiness of the real world. Yeah, yeah. No, that makes sense. Interesting. Even for accelerating the research loop, right?

34:03Like we're using agents today, right? To be able to like run experiments, see what works, see what didn't work, plot the differences, detect that something is going the wrong direction early and then change parameters. All of that is already happening, right? So in that sense, yeah. It's exciting. So those will not be simple code. So I think that we have tried to write code to control robots for decades. Yeah, so because those code, if it's rule-based, it's very hard to generalize to all kinds of environments. This is why we are switching to this very data-driven paradigm. But in theory, that all the neural networks, you know, training, data-driven optimization, they're still within the code space.

34:39We still write code to generate all these things. So I think eventually it's possible, but it requires a lot of auto-research to make that happen. Yeah, I was thinking about like very like routines almost of like very explicit, it. I guess maybe the action space is like too unconstrained to make that happen. But like, could I, you know, I don't know. Just thinking of examples where like it might, you could do something useful if you could like write code to do some of these use cases deterministically. But now I hear what you're saying that it doesn't, it doesn't generalize well. Well, so how can people actually start getting access to the model?

35:13I feel like it's, we've made a bunch of progress. Like what's the availability story? Where can people actually start using it? Yeah, we're really excited. So Embodied Reasoning 2.0 or 2 is going to come out. That'll be directly available via AI Studio. And I think I can say this, but soon it'll be available via the Gemini Enterprise. Agents Platform. Agents Platform. But Gemini Robotics 2 will have the Embodied Reasoning model available directly via an API, and that will be general accessible. And we're really encouraged people to use that. And then the action models themselves are also going to be available.

35:47we work directly with our deep partners. So they'll be the ones who use sort of the biggest, strongest versions. We also have an on-device version of that that our trusted testers can use. And so people are welcome to join. We currently have a wait list, but we're trying to do more about it, our trusted tester program. And then they can actually get access to a on-device deployable version of the action model where they can actually fine tune that model directly on either their tasks or their robots and actually try it out in practice. Yeah, that's awesome. I'm excited. What any any advice? And maybe I'm maybe I'm misremembering this, but I feel like we do have a bunch of customers who use the ER models who actually like aren't robotics companies and they just happen to be in one of these like domains for like video audio spatial understanding or something like that.

36:28I don't know if that's like a suggested path for for folks, but I assume it's like those domains of like video spatial understanding where like the ER model is like better on a bunch of these core benchmarks. Yeah, I mean, these models are better in a few ways. Like the ER model in particular, this release is a lot better at video understanding. So before, it's always been very good and state-of-the-art at spatial understanding. So 2D and 3D bounding box understanding where objects are in 3D space. Now it can understand videos. It understands also the semantics of a task. So if you ask it, at what point should I stop pouring my coffee?

37:03Or am I done closing this Ziploc bag? It actually understands how far along you are in that progress. And so that's extremely useful whether you're doing any kind of like video understanding capability or for robotics, exactly. If you use it as your agent, then now it can be the agent that understands how far along you are and decides to like switch to a different task or decide that you're done. And then these models are also significantly safer. This is our safest model yet. and it's safer not only in the regular way in which all of our Gemini models are safe in terms of content safety, but it's also safer because it understands the likelihood that a model is going to be completing a task.

37:45So it also helps you understand that if you give an instruction that is very ambiguous, for example, then it will ask proactively the human, oh, that instruction is very ambiguous. Like, what do you mean? And so it helps with proactive clarification. And then the last one is that we're also making it really strong at detecting humans and humans' proximity to robots, which is very important when you're talking about collaborative robots that are in human-centric spaces. So those are just a few areas. We're also introducing a new safety benchmark that we have open source. It's called Asimov Agentic.

38:18Asimov is our benchmark, and it essentially has a large set of examples, real world examples, where you have to make a decision about what the robot would do next or what you should do next based on this situation. So it's a lot about semantic physical understanding. It's a lot about common sense that robots would need to have if they're going to be operating and doing lots of tasks around us. Very cool. This was an awesome conversation. I was super interested to hear about the launches. I'm very excited for folks to get their hands on the models. It's cool to also come to y 'all's space. I feel like there's, it's very much more interesting than the normal Google offices.

38:56So I'm glad to be a guest and see all the cool hard work that y 'all are doing. So congrats on the launch. Very excited. Thanks everyone for watching this episode of Release Notes. We'll see you in the next one.

From the publisher

Carolina Parada, Stuart Bowers, Kanishka Rao, and Jie Tan from Google DeepMind join host Logan Kilpatrick inside the Gemini Robotics Lab to introduce Gemini Robotics 2, Google DeepMind's new suite of models bringing whole body intelligence, dexterity, and multi-robot collaboration to general purpose robotics. Their conversation covers the arc from reinforcement learning to vision language action models and the open challenge of physical AGI.

Watch on YouTube: https://www.youtube.com/watch?v=-rYFDefcq3k

More from Google AI: Release Notes

All 30 episodes
Introducing Gemini Robotics 2Google AI: Release Notes · 39 min
Listen in VO