Intelligent Robots in 2026: Are We There Yet? with Nikita Rudin - #760

8 Jan 2026 · 1 h 7 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

TWIML AI Podcast Episode #760: Intelligent Robots in 2026: Are We There Yet? with Nikita Rudin

Episode Summary In this episode, host Sam Charrington speaks with Nikita Rudin, co-founder and CEO of Flexion Robotics, about the current state and future potential of robotics technology. They discuss the challenges of achieving fully autonomous robots and delve into the advancements in robotic locomotion, the sim-to-real gap, and the interplay between various training methods in AI.

Key Topics Covered

Current State of Robotics

  • Lack of Value Generation: Rudin asserts that no humanoid robot currently generates significant value, indicating robots may perform tasks poorly or not as intended.
  • Progress in Locomotion: Discussion on how reinforcement learning and simulation have accelerated robot locomotion capabilities, yet challenges remain.

Advancements in Robot Learning

  • Reinforcement Learning (RL): Explains the role of reinforcement learning in training robots, particularly in locomotion and navigation.
  • Sim-to-Real Gap: The gap between simulated and real-world performance is exacerbated when visual inputs are introduced, complicating the training process.
  • End-to-End Models vs. Modular Approaches: Debate surrounding whether an end-to-end model is preferable over a modular system, with Rudin advocating for a pragmatic, modular approach for the time being.

Training Strategies

  • Real-to-Sim Concept: Introduces the notion of refining simulation parameters with real-world data to enhance training fidelity.
  • Hierarchical Approach: Flexion Robotics employs a hierarchical strategy that combines Vision-Language Models (VLMs) for high-level task orchestration and low-level tracking for precise motor control.
  • Combining Imitation Learning and RL: Discussed the potential of combining these techniques to enhance training efficiency and effectiveness in achieving complex behaviors.

Future Predictions

  • Robots in 2026: Rudin predicts that true value-generating humanoid robots will emerge by late 2025 or early 2026, starting in industrial settings before moving to consumer spaces.
  • Industrial vs. Consumer Robots: Anticipates a gradual transition where industrial applications will precede widespread consumer deployment.

Key Takeaways

  • Complexity of Robot Locomotion: Current locomotion capabilities remain limited, as robots struggle to navigate in diverse environments without substantial pre-planning.
  • Need for Modular Systems: While end-to-end neural networks have potential, the current complexity of tasks necessitates a modular approach that combines various learning techniques.
  • Challenges with Visual Inputs: Adding visual input to robots complicates their operation due to increased noise in the data and challenges in sim-to-real transfer.
  • Importance of Practical Training: Effective training for robots requires significant data, often through imitation of human actions, highlighting the importance of practical implementation in learning algorithms.

Conclusion Nikita Rudin emphasizes that while the field of robotics is rapidly advancing, significant challenges remain before fully autonomous, reliable robots can operate in the real world. The discussion underscores the importance of blending various learning strategies and realistic training environments to bridge the gap between robotic potential and practical application.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Value of Humanoid Robots

0:00 to 0:34

Discussion on the current lack of value generated by humanoid robots.

“My hot take on that, and I'll be happy to be proven wrong, but I think there is not a single humanoid robot today that generates value.”

Nikita Rudin's Background and PhD Work

0:51 to 3:43

Nikita shares insights from his PhD research on legged robot locomotion.

“You've been working in this space for quite a while.”

Challenges in Robot Locomotion

3:43 to 4:43

Exploration of the challenges and advancements in robot locomotion technology.

“The robot dog is going, maybe opening some doors.”

The Role of Perception in Robotics

4:43 to 6:46

Discussion on how perception impacts robot locomotion and the sim-to-real gap.

“can it do it or not, how reliable it is, my take is that locomotion is not solved.”

Navigating Complex Terrain with Robots

6:46 to 9:35

Insights into how robots navigate and make decisions in complex environments.

“The typical thing we refer to is the so-called sim-to-real gap.”

Reinforcement Learning in Robotics

9:35 to 14:02

Discussion on the use of reinforcement learning for training robots and locomotion models.

“It sounds like what you're saying is that a more modular approach can still be a pragmatic way to overcome the challenges of end-to-end training.”

Challenges of Humanoid Robot Movement

14:02 to 15:40

Learn about the specific challenges in tuning humanoid robots versus quadrupeds.

“the URDF that describes the robot, retraining it, then you're ready to deploy.”

Robot Demos: Excitement and Realities

15:40 to 18:10

Explore the excitement around robot demos and the reality of their capabilities.

“We've been working on robots for a really long time, but it seems like over the past year, we're seeing advancements via video demos coming very quickly.”

The Reality of Humanoid Robots in Homes

18:10 to 19:54

Discuss the readiness of humanoid robots for home environments and industry focus.

“for a humanoid robot that is, you know, ready for the home, quote-unquote?”

Value Generation in Humanoid Robotics

19:54 to 21:28

Understand the current limitations of humanoid robots in generating real value.

“But you have a little bit more control over what's happening.”
Show all 31 chapters

Teleoperation and Data Collection in Robotics

21:28 to 22:29

Learn about the role of teleoperation in training robots and collecting data.

“Meaning there might be a robot that does something fairly close to what it's supposed to do in a factory or in a warehouse, but it's not the exact task.”

Supervised Learning and Action Models in Robotics

22:29 to 25:03

Explore how supervised learning is applied in robotics and the evolution of action models.

“or is that more a supervised learning type of an approach?”

Generalization Challenges in Robotics

25:03 to 26:45

Discuss the challenges of generalization in robotic training approaches.

“I think Ross Tedrick had an amazing talk at Stanford.”

Sim-to-Real Gap and its Implications

26:45 to 28:00

Understand the complexities involved in bridging the sim-to-real gap in robotics.

“I would guess it's because the robot cannot ignore the visual input.”

Understanding Sim-to-Real Challenges

28:00 to 30:00

Explore the complexities of translating simulations into real-world robot actions.

“You know, we've been making, you know, this has been a known issue for many, many years.”

Adapting Robots for New Tasks

30:00 to 32:20

Learn how teams adapt robots for new environments and tasks with evolving technology.

“which sounds very computationally expensive.”

Breaking Down Complex Tasks

32:20 to 34:30

Discover how robots can be trained to perform complex tasks through simple primitives.

“It should be less than one day of work if we optimize some of our processes.”

Combining Models for Better Performance

34:30 to 37:30

Understand the integration of separate models to enhance robot capabilities.

“But what's interesting is that that part is basically solved with a VLM.”

Challenges in Onboard Computing

37:30 to 42:01

Examine the limitations and requirements of computing power in robotic systems.

“So you see some interpolation between those different models.”

The Role of Large Models in Robotics

42:01 to 43:44

Learn how larger models enhance robot reasoning and performance in various settings.

“So some of these models, just to clarify, some of these models can fit on the robot.”

Building Simulation Environments for RL

43:45 to 45:30

Understand the importance of simulation environments and custom RL algorithms in robotics.

“Like the simulator is like the platform and the simulation environment is like the thing that you create about your scenario?”

Combining Imitation Learning and RL

45:31 to 47:35

Explore strategies for integrating imitation learning with reinforcement learning to improve robot training.

“I can talk about two different ways to combine them.”

Challenges of RL in Real-World Scenarios

47:36 to 49:23

Discuss the complexities of applying reinforcement learning in real-world tasks and the limitations of simulation.

“How do you think about comparing and contrasting those approaches?”

Human Learning vs. Robot Learning

49:24 to 51:19

Compare human learning processes with robotic learning, focusing on efficiency and complexity.

“RL in real is better, you know, in what ways would it be better for that particular scenario?”

Reward Functions in Reinforcement Learning

51:20 to 54:04

Delve into the complexities of designing reward functions for robotic tasks and their significance in training.

“if you think like if you're doing some tasks with your hands, the amount of information you're getting from all the nerves and all your skin, also from your muscles that are tired, etc.”

Future Predictions in Robotics

54:05 to 56:00

Hear insights on the future of robotics and potential advancements expected in the coming years.

“So we are training, for example, we're mostly using some variant of PPO, which is an actor-critic algorithm, which means that you're training both an actor and a critic.”

Guiding Robot Training with Demonstrations

56:00 to 56:48

Explore how to guide robot learning through human demonstrations.

“Once you've described the perfect task that you want, how do you guide the policy, the training process towards that?”

Predictions for Robotics in 2026

56:48 to 57:53

Hear expert predictions on the future of robotics and humanoid robots.

“or, you know, several years, whatever horizon you'd like to offer in terms of robotics?”

Diversity in Robot Designs

57:53 to 59:16

Discuss the variety of robotic strategies and designs in development.

“And presumably you see that happening first in industrial settings and then consumer?”

Form Factors: Humanoid vs. Alternative Robots

59:16 to 1:01:34

Evaluate the effectiveness of humanoid robots compared to other designs.

“there are a few companies in the U.S., maybe one in Europe, and probably 50 in China building that exact thing.”

Robotics Kits for Learning and Exploration

1:01:34 to 1:04:39

Discover accessible robotics kits for enthusiasts and learners.

“wheeled platforms get stuck a very common thing that is easy to imagine is if the floor is not perfectly flat If you have something, you have cables or, of course, stairs, your wheel platform is stuck.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00My hot take on that, and I'll be happy to be proven wrong, but I think there is not a single humanoid robot today that generates value. meaning there might be a robot that does something fairly close to what it's supposed to do in a factory or in a warehouse but it's not the exact task so in the end it's not generating value because it's not doing the actual thing it's supposed to do.

0:34all right everyone welcome to another episode of the terminal ai podcast i am your host sam charrington today i'm joined by nikita rudin nikita is co-founder and ceo of flexion robotics before we get going be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Nikita, welcome to the podcast. Thank you. Very excited to be here. I'm excited to have you on the show and I'm looking forward to digging into our topic for the conversation, which is really digging into the gap between where we are today with robotics and where we need to be to fulfill the vision of the technology.

1:11You've been working in this space for quite a while. You did your PhD at ETH Zurich and spent some time at NVIDIA. Why don't you share a little bit about your PhD and the focus of your research? So when I started, we were trying to use simulation with reinforcement learning to teach a legged robot very simple things, like just walking on flat ground. And when the robot could take a few steps, that was already a big success. And the core focus was to reduce the training time needed to achieve that. And when you say legged, like a quadruped? Like a quadruped, a big legged robot dog. We're not using Boston Atlantic Spots.

1:53We're using Antibiotics Animal. Antibiotics is a Swiss startup that was a spinoff from our lab. Very similar to a spot, but it's made in Switzerland. We're really trying to reduce the training time needed to achieve that. So before I started, there were some results for reinforcement learning for such quadrupeds, but it would take weeks of computation to achieve anything. And using GPUs and massively parallel simulators, we managed to reduce that to just a few minutes. Actually, we had a demo on stage at some conference where we were running training live on a laptop while I was holding the robot, and every 15 seconds, the laptop would send the latest policy to the robot.

2:34And you could literally see how it went from just falling over to taking a first step. And then after three or four minutes, it would be able to walk around the stage. That was a pretty cool visual demo for everyone to see exactly how the learning process happens. And from there, my PhD was pushing the agility of that robot using the similar techniques. It was still training neural networks and simulations and transferring them to the real world. But the inputs got more complicated. The tasks got more complicated. So by the end, we could go to a search and rescue facility here in Switzerland. So you have to imagine collapsed buildings, a lot of mud, moss, gravel, big rocks, terrain that is very hard to navigate, even for a human.

3:19And we would just tell the robot to go from point A to point B and would use its whole body. So it would use the knees to climb on top of big, big rocks and then jump over gaps. and all again, all autonomously, all end-to-end using images and the state of the robot to plan its next actions. And telling that story about deploying this robot in a search and rescue context, envisioning the demo, I've seen similar things. The robot dog is going, maybe opening some doors. Maybe that wasn't part of your demo, but I've seen similar demos of the robot dog climbing hills and crossing rubble. and I think those demos like, you know, attempt in many cases to land the idea that, hey, you know, like flagging the ground, we're done here.

4:05Like talk a little bit about the distance between, you know, what you were able to accomplish with that demo and like what you think needs to be done to deploy one of these robot dogs in a real search and rescue scenario, for example. It's a great question. And I had this debate so many times, even with my colleagues at NVIDIA and ETH. The general question is, is locomotion solved or not? And we already had this debate five years ago when the robot could just barely walk on the flat ground and people were saying, yeah, locomotion is solved. We don't need to focus on it anymore. My take was always, no, until the robot can really go anywhere a human can go, and you don't even need to think about can it do it or not, how reliable it is, my take is that locomotion is not solved.

4:52so where we are today is that anything that is blind can be very very robust blind meaning that the robot does not perceive its terrain it's just reacting that makes the training much easier because you just have to throw a lot of things at it it always tries to be stable especially for a quadruped it's fairly easy to remain stable and for example you can even walk upstairs it will just hit the first step realize there is something and then climb up. Doing it with perception. So now you want it to actually react. You don't want it to hit the first step. You want it to place the feet much more carefully.

5:31That's harder. Today, end of 2025, we're a stable where we can do that. We can have fairly good policies that plan the sequence of actions based on the perceptive inputs. Because we're training everything in simulation, this means that we need to be much more careful in how we simulate those sensors as well. I need to put a lot of effort into creating complicated trains that can be seen by the cameras and also all sorts of noise models to simulate all the disturbances and defects of those images. So what I'm hearing you say is that the addition of the additional information that you would think would help actually makes it more difficult because it introduces a lot of noise.

6:15And whereas previously the robot could kind of stumble through the terrain, now it's trying to incorporate this visual input to plan, but it ends up making it. What actually happens when it is trying to do this? Do you see it stuttering or does it just not work? Is it just hard to train from a model perspective? What happens? So, I mean, in the end, what happens is that the final behavior is better if you do everything right. But that's a big if. The typical thing we refer to is the so-called sim-to-real gap. If you're training things in simulation and deploy them in real life, it's not the same.

6:59and things that work really well in simulation might not work at all in real life. That sim-to-real gap is much larger once you have perception in the loop. Once you're simulating either depth images or RGB images, it's even worse. So that makes the job of the researcher, of the engineer harder. You need to cross that sim-to-real gap for both the physics of the robot and the perceptive inputs. You have to simulate the sensors carefully. But if you do it right, then the final behavior is actually much better because you can see that the robot is not simply reacting to whatever is happening under its feet, but actually planning accordingly in advance.

7:36Is this problem of locomotion solved then with vision, at least for, you know, let's take quadrupeds as an example, or are there still kind of outstanding issues? Is there a generality gap? How do you think about it? It's interesting. It really depends where you draw the line on locomotion. Once the robot can cross complicated terrain, the next step goes more into navigation, which means where should it go? Should it climb that thing or should it avoid it? For now, everything I've described so far was mostly using geometry, so there are no semantics. But now if you imagine the robot walking behind you.

8:23Meaning you give it a point A and a point B, and it's going to take the straight line and kind of plow through whatever is between here and there, as opposed to think about which way to go? More or less. If there is a huge wall, it might avoid it. But then mostly it will try to climb on whatever is in front of it. Okay. Now if you imagine the robot walking behind me in the office, you don't really want it to climb on every single desk. You want to avoid some things, but also walk on others, right? So if there are stairs, you want to take the stairs, but you don't want to hit plants or whatever else, which means that suddenly you have to add semantics to the policy.

9:00And what does it mean to add semantics to the policy? Once again, if you go from simulation to reality, it makes the simterial gap bigger because suddenly you have to simulate all these offices, all these different objects. Plus, you probably need to give it images, not just depth images. You need to give it RGB images, which means you need to simulate it in a photoreal way. Or the other option, which is probably the more correct one in the short term, is to split the problem. So you train the robot to be very good at walking on anything, but you don't give it some adding information. And you train another thing on top, which will be the, you can call it the planner or a higher level policy that will steer it around.

9:41And now, historically, in the conversations I've had with roboticists, this has been a big debate, whether we should be using end-to-end deep learning models that can figure all this stuff out or using a more modular approach. It sounds like what you're saying is that a more modular approach can still be a pragmatic way to overcome the challenges of end-to-end training. Yes, it's interesting because my whole PhD was about going more and more end-to-end, specifically for locomotion. And still, I'm here arguing that we should not do everything end-to-end. I think at some point we'll get there. But in the short-medium term, as you said, the more pragmatic approach is to split the problem and use different techniques for different parts of the problem.

10:29We're splitting the problem into kind of a locomotion model and a planner model. Does the planner, like what, what's the objective for the planner when you're trying to, you know, the, you know, English language objective is I want this thing to intelligently choose the, like the best path. But what is best path? How do you define that? How do you create an objective around that? Is it the path that maybe it's the path that uses the, that's most power efficient if you're in a robot, or maybe it's the path that, you know, gets you there faster or at least distance. Like, how do you balance all that?

11:11There are different ways to do this. If you still choose the RL route, when we're doing that, you can train these planners with reinforcement learning. And typically you have to define this reward function. It would include things like don't hit anything, so avoid objects. And don't move too fast, because that's one thing that makes robots seem very dangerous. These reinforcement learning policies will try to optimize everything so they will go very, very quickly to the goal. This is not really what you want with a robot that operates around humans. Typically, we're actually trying to slow them down as much as possible.

11:48And yeah, that's mostly it. And then it depends really on what is in front of the robot. There might be things that it's okay to work on others that it's not. There is also another approach. Once you split the problem in half, you could train the locomotion with pure reinforcement learning and simulation, but you could train the planner with other data. For example, videos of humans walking around the office. You can extract the trajectories from that. And then you don't really need reinforcement learning anymore. You train behavior cloning, imitation learning policy that will just steer the robot just like a human would walk.

12:21And thus far, speaking about how a human would walk, thus far we've been talking primarily about quadrupeds. How much of all of this translates from quadruped to humanoid robots? All of it transfers. That is the magic of reinforcement learning, that these policies don't really care if they're controlling a quadruped or a humanoid. And this was the big switch from my PhD to flexion, where we're working mostly on humanoids. And we've seen that the exact same techniques transfer. There is one interesting thing that happens with humanoids. That makes sense to me for a planner, but it's less intuitive for a locomotion model.

13:03And maybe I'm thinking of the locomotion model that maybe the locomotion model itself that I'm thinking of is like split into multiple components, but let me be more specific. I'm including in locomotion, like, you know, the outputs that control stepper motors and all that kind of stuff. Is that part of what is trained in the locomotion model? And then I would think that you would need to at least like tune it or do something else to, if you're going to change the form factor of your robot. No, this is a very good point. When I say it transfers, the general techniques transfer. The models themselves don't.

13:50So for sure, you need to retrain a new policy, a new controller for the new robot. But if all the simulation pipelines are general enough, it can be as easy as changing the input file, the URDF that describes the robot, retraining it, then you're ready to deploy. The tuning part is an interesting one because it is a little bit harder for humanoids compared to quadrupeds. And I don't think it's really related to the fact that it has two or four legs or anything like that. Personally, I think it's mostly related to the fact that we have very specific expectations of how a humanoid robot should walk.

14:27Whereas a quadruped will, if it walks slightly differently from a dog, it's completely fine. But a humanoid, if it doesn't move the arms in the right way, if it bends the knees too much or walks a little bit sideways, humans have a very strong reaction to that. I saw a tweet kind of touching on the same idea just the other day. And it was essentially this humanoid robot that was kind of locomoting like a quadruped. Like it was like on its back with its arms like this and like moving really quickly. and the main thrust of the tweet was that humanoid robots move like humanoid robots because that's our expectation but that there are potentially other even with that form factor there are potentially other more efficient ways for these things to move but they just seem wrong to us.

15:27Yeah, that's true. But if we wanted robots operating around humans, we have to create some trust. So we have to take the less optimal route if it makes humans feel a bit more comfortable. We've been working on robots for a really long time, but it seems like over the past year, we're seeing advancements via video demos coming very quickly. I kind of asked this question before, but I want to have you kind of parse through how to think about you know these videos like you know we were seeing humanoid robots walking at the beginning of the year and now they're running you know now they're doing dishes and all these things like what goes into creating a demo like that um and what are the limitations of what it says about what the robot is capable of first of all i would say that it's really exciting that the whole ecosystem is moving so quickly.

16:27It feels like every single day there is a new video of a new robot doing something, and that's really, really good. We see a lot of progress both in hardware, but also how it's on the AI side as well, what these robots are capable of doing. Having said that, typically when I personally see a demo of a robot doing something, my approach is to think about what is the absolute easiest way to achieve that, and that is typically how it's done. So we are showing, as an ecosystem, we're showing a vision of what robots should be doing, but the behind the scenes is sometimes a little bit different. For example, if you have a robot standing and doing some sort of manipulation on a table, folding sheets or something like that, typically one of two things is true.

17:17Either there is someone hiding behind the robot, behind a curtain, somewhere in another room, teleoperating the robot, so it's not really autonomous. That's, let's say, one-third of the cases. And the other two-thirds are where the robot is actually autonomous. But to get there, 100 people had to teleoperate 100 robots to collect a lot of data, hours and hours, in some cases thousands of hours, of robots doing very, very similar things in probably the same environment. then collect data, train a policy that can imitate that data, and then they're able to deploy robots autonomously, which is not exactly what you would see when you watch that video because it feels like the robot can just adapt to anything and come to your home and do everything you're doing.

18:05On that front, we're just not there yet, but we'll get there soon. Remind me the name of the company that started taking pre-orders for a humanoid robot that is, you know, ready for the home, quote-unquote? You're referring to it to 1X. 1X. I looked at that and, you know, I've had a lot of conversations along these lines about kind of where robots are.

18:33And looking at that, I questioned like, okay, are we like really a lot further than I think we are? Or, you know, is something else happening here and this is, you know, the early buyers are going to be beta testers and it may or may not work once it gets to the house. Do you have any takes on not that specific company necessarily, but like the readiness of humanoid robots for the home? Also, OneX is not the only one. There are a few more who announced similar things. I mean, in a way, it's good. But it's very, very ambitious to sell robots next year into people's homes. I think they have a big challenge ahead of them.

19:17So let's see how far they can get next year. But usually these companies are fairly honest that it is a better of a better, an alpha program. It will be just for early adopters. It will take a few more years before you can really buy these robots and send them into homes. That's partially why as a company we're focusing more on industrial use cases. There are a lot of other challenges with industrial risk because now suddenly performance is really important. You need to be very fast. You cannot slow everything down, at least in most cases.

19:54But you have a little bit more control over what's happening. So, for example, if we want to deploy 10 robots in a new warehouse, it's easier for us to send an engineer for one or two days to check that everything is in order, as the operators they should if they don't fine-tune a few things. And then let robots work, which is not something you can do in everyone's home. And are the tasks in the industrial setting, I'm imagining they're more repetitive, more consistent, less variation than, you know, run to the fridge and grab me a Coke. Yeah, that's right. and we get to decide which tasks we tackle and which ones we leave for the future.

20:42So we can start with simpler things. A lot of it is moving objects around, bringing objects from point A to point B, moving boxes, opening boxes, taking items out or putting objects into boxes, putting the box in a truck, sending it further. And this seems really within reach or next year, maybe the year after that. But even though you might have seen a video of a robot doing that, that doesn't mean that it's, you know, ready now and people are doing it now without, you know, with it not being in a development phase. Is that fair? Yeah, that's fair. My hot take on that, and I'll be happy to be proven wrong, but I think there is not a single humanoid robot today that actually generates value.

21:28Meaning there might be a robot that does something fairly close to what it's supposed to do in a factory or in a warehouse, but it's not the exact task. So in the end, it's not generating value because it's not doing the actual thing it's supposed to do. Meaning it's doing some variant of the thing or it's like there's a handler that's fixing up, cleaning up after the robot as it makes a mess across the... Exactly, yeah. And typically it would have more handlers than you had people before. So you could argue the value is negative. But once again, we'll fix that. We talked about how you create these demos and the idea that there's either real-time teleoperation or many, many people doing teleoperation to collect training data.

22:19Talk a little bit about then, after that training data is collected via teleoperation, what the approach is for training. Is that data then used as part of RL, or is that more a supervised learning type of an approach? Typically, it is supervised. So you record the data as images of the camera and how the commands that the teleoperator sent to the robot, which typically is like how should you move your hands in space and how should you move your fingers. So that is recorded. and then a big transformer is trained to produce the same actions from the same pictures, the same images. Now, what's interesting is the whole field shifted a little bit from just training these transformers from scratch to using vision and language encoders that were pre-trained on internet scale data.

23:20But VLMs, off-the-shelf VLMs? You take a VLM, you remove the output, and you train a new part of the network on top, and then you call that a VLA. It's a vision language action model, where the vision language part was pre-trained before, and the action part is trained from scratch. Got it. So as opposed to predicting a next language token, you're now predicting an action token, which is then translated into, you know, a separate motor motion or something like that. Yeah, exactly. Kind of compare and contrast that approach with what folks were doing before. Are we doing that because it's cool or are we doing that because it, you know, how much is having a pre-trained model to start with like save us from the generic transformers?

24:17It's a very good question. The general thought is that it helps with generalization. Since the language and vision encoders were trained on Internet scale data, they're supposed to generalize. A typical case was that if you don't do that, you would train a robot during the day. Then if the lights go down at night, it won't be able to perform anymore. Even though there are still lights. Everything should just work. a human wouldn't even see the difference, but because the image embedding changes a bit, the policy doesn't perform anymore. I believe that this gets better with the pre-trained encoders.

25:01To be completely honest, this is overall the generalization capabilities that still need to be proven. I think Ross Tedrick had an amazing talk at Stanford. We were talking about their efforts at the Toyota Research Institute where they were comparing training policies on a very specific task with little data versus training more journalists with a lot, a lot of data. And they were seeing some signs of generalization, but I don't want to quote him directly, but it seemed like it's not fully understood yet how much of the generalization is coming from that. Now, Matt, those two little data, lots of data to the Transformer VLM discussion, and the VLM would be the little data and the transformer was a lot of data because we're assuming the VLM was pre-trained or was it reversed?

25:56No, it's actually reversed. If you include the pre-training as data that you get for free, then you have a massive amount of data for pre-training and then you can add less data for fine-tuning this action head. I guess it doesn't matter which one is which because the results were somewhat inconclusive is what I'm hearing. We need to go in more detail and I'm quoting other people here. So it's a bit hard to say. I think it makes a lot of sense to have pre-trained visual encoders, language encoders, because you don't want to relearn language every single time you want the robot to do something.

26:36Like language is language. And by the way, we have amazing VLMs now, so might as well use them. there's more of a question of this action head should you train it on a lot of random data or just on the thing that you want the robot to do in the end and this is still an open question One question that that raises for me is I just had a conversation where we were talking about how with VLMs generally they kind of ignore a lot of the visual information and really rely more heavily on the language information And it seems like in a robotic scenario that, you know, that would be even more harmful to what you're trying to accomplish.

27:23Do you run into that as a challenge? I've heard the same thing. I haven't seen it in the VLA case. I would guess it's because the robot cannot ignore the visual input. it's the main source of information. What tends to happen, however, is that they ignore the language inputs. If you train the robot to always do the same thing, I don't know, if you have a box and you have an object inside, it always has to take it out, it will completely ignore the language. It will just do the same thing. It will try to guess from the image what it's supposed to do. You mentioned the sim-to-real gap. You know, we've been making, you know, this has been a known issue for many, many years.

28:10We have been making good progress on closing that gap. Talk a little bit about your experience, what is required today to create a model in SEM and have it run in real. Do you have to, are you doing things kind of explicitly or specifically to address kind of real world or is it just the models are better, the process is better and you don't really think about that anymore and it just kind of works? You need to do a lot of things very explicitly. The challenge is that to cross the sim-to-real gap, you need to have a very deep understanding of both worlds, of the simulation and how it works and of the real world, which means that if you want to have a robot that walks around as it should in sim, you need to go very deep.

29:04You need to know exactly what's happening between a command that the policy outputs and then all the way down to torque in the motors. And there are typically 10 different layers of transformations, even just on software, of how we go from high-level command to actual current in the motor. And it's very tempting to ignore that. But by understanding every single layer and knowing all the different transformations, then you can properly simulate it. and this really unlocks better performance. So that suggests the level of simulation that you're doing isn't like you pull up your sim environment and get generic humanoid robot and you're going to train some model and deploy it to something else.

Read the full transcript

29:50It's like you have a digital twin of your humanoid robot in a high-fidelity simulation environment and you're training to a very fine level of detail, which sounds very computationally expensive. No, so we're much closer to what you described first. So we have a very generic simulation environment, but there are some very specific things that are important. One clear example is what are the torque and velocity limits of a motor? You cannot expect it to do something that is not possible on the real robot. So you need to add those limits. And there are a few more things like that, like what kind of delay can you expect between a command.

30:29So you need to identify a few of those parameters. And we are actually doing this usually in what we call a real to sim process. So we take the real robot, we literally hang it in the air, we let it shake a little bit, collect data of all the different motors, and then we know which are those important effects that we need to identify. We identify them and add them to the simulator. But simulation speed is the most important thing. So you cannot afford to simulate all those different effects, currents, magnetic fields, etc. You need to abstract all of it away. That sounded very hard and expensive.

31:07My personal take is that you still need to understand them even though you're not simulating them. Got it. And so is the, it sounds then that it sounds then like the result of that process is not a general model that you could deploy to any humanoid robot, but one that is specific to the humanoid robot for which you took the reel to send, you know, those key parameters. but that because you're able to abstract it out to these some handful of or several handfuls of key parameters, it's relatively easy to do new robots. Yeah, and this was also the surprising part of one of the key learnings of this year is switching robots is fairly easy as long as the hardware performs reasonably well.

32:04So now as a company, we work with a few different suppliers of robots and a few different partners as well with whom we're working closely. We've deployed controllers on, let's say, between five and ten different robots. And we see now that making a new robot walk is a few days of work. And it should be less. It should be less than one day of work if we optimize some of our processes. Now, bringing them to a new task, this is a bit more challenging. this requires more engineering today and this is what we're focusing on one of our key metrics is how much human effort is involved in bringing a new robot to a new task a new robot is very easy today, a new task is something we're working on and in this context like how we've kind of talked a little bit about this but how specific is a task meaning like is a task pick-in-place or is a task robot in this warehouse picking off of this line and placing into these bins?

33:10That's a great question. More like pick-in-place. But there is an interesting concept there, which is we are trying to leverage the information contained in large VLMs to orchestrate and break down complex tasks into clear subtasks. Even though that's not what we're focusing on today, cooking is a great example, a great metaphor. If you wanted to, cooking, yes, if you wanted to train the robot to cook every single meal on the planet and you would say each meal is its own task, you would never finish that, right? Like the set of tasks is huge. But what you could do is you give the recipe to a VLM.

33:58you also give it images of what robot sees if the recipe says cut a cucumber the VLM would say grab the knife grab the cucumber and do this sort of motion to cut it and then you can break it down into much simpler primitives like cutting things, holding a pan putting it down somewhere, filling a glass of water, pouring it things like that and suddenly the set of these primitives is not infinite anymore the challenge is that now you need a higher level intelligence that will orchestrate all these primitives. But what's interesting is that that part is basically solved with a VLM. It's not 100 % there, but it's moving much faster than the actual physical interaction of doing all these motions.

34:46So can you elaborate on that? The orchestration is solved, and if so, how is it solved? and what's the relationship between that orchestration and what we talk about is like the reasoning capabilities of these large models. Is it the same thing or related? It's similar. So maybe one way to describe this is on our website, we have two videos. We have a video of a robot walking in a forest and picking up trash. This is mostly there to showcase what's possible and also play a little bit on our Swiss angle, using our nature. We have another video where the robot is doing the same thing in our office.

35:29And in that second video, it's 100 % autonomous. So you give it a text prompt. I think we're saying something like, pick up the toys in front of you and drop them in the basket at the end. And we're using an off-the-shelf VLM for that. The way this works is we're giving the images to the VLM and we're allowing it to do tool use to call specific skills of the robot. So the VLM would say, oh, I see a toy there on the ground. Let's walk to the toy. And this let's walk to the toy is a skill that is actually triggered and executed by the robot. Once we're there, it will trigger, pick up the toy, then go to the basket and drop it off.

36:10And so by having a few of those primitives, which are walk to things, I mean, the walking is locomotion, as we discussed, is itself very complicated. So you can walk on stairs, can walk on a bunch of different complex terrain. But by having like walking, picking things up from the ground and then dropping them somewhere else, we can recombine it in many, many different ways without any retraining, just by prompting an off-the-shelf VLM. When I hear tool use, I hear like separate process or module or model. is that the is that the case and like how do you think about you know then I think of like if you know if we got a bunch of tools you've got a bunch of these you know separate models or modules like are they actually does this architecture imply that they are in fact separate and you know trained separately or are they more universal somehow?

37:13Another great question so that in our case specifically today, they are separate, but we are actively working on merging them together into one single model, a more general model. And the hope with that, again, is that you see some generalization. So you see some interpolation between those different models.

37:37And the way we would do that is actually by still training those different modalities primitive separately, and then using them as data generators to collect a massive amount of data in simulation to then train one of those larger VLAs across the whole data set such that it can perform everything. We're seeing early results of that in our company. Things are going in that direction, but we still need to prove that this actually leads to the generalization we were talking about before. And hearing you describe that, You know, often when I hear like these kind of student teacher types of approaches or, you know, I think of like distillation and trying to get to smaller models, you're not necessarily trying to do that.

38:23But talk a little bit about the hardware capabilities from a model inference perspective and like where we are in terms of, you know, model size, that kind of thing. In our plan, if we go through that whole process, we train, let's say, 50 of those primitives. We distill everything. we are still developing a hierarchical pipeline where you have three models interacting with each other. It would start with a relatively large, let's say, VLM that would be allowed to use reasoning. And the output of that one would be that VLM would go from a very abstract task to clear subtasks such that if you have a robot here and I tell it, go pick something up in the fridge and would say, turn around, go through the door, open the fridge, grab the thing, close the fridge, et cetera, et cetera, clear instructions.

39:16Then those clear instructions would go to that VLA that we were describing before, where if it receives as instruction, open the fridge, it will plan a kinematic motion for the arm to grab the fridge handle and open it just a few seconds into the future. And finally, we have what we call a whole body tracker, which will receive this plan of how the hand should move and how the whole body should move. And then we'll control the motors of that specific robot to execute that motion. Okay. Now, the size and frequencies of these models are very different, and that's why we think it makes sense to have three of them.

39:53The final one, the whole body tracker, is a very simple, very small model. Typically, it's a very small transformer. To be completely honest, it doesn't even need to be a transformer. But today, everything should be a transformer. and that can run very easily. We typically run them at 50 hertz, 50 times per second, even on the CPU of the onboard computer of the robot, not even the GPU, because it takes more time to send it to GPU and get it back.

40:23Then I'll skip the VLA for now, going to the VLM. That one is typically fairly hard to run onboard on the robot. For now, it's running off-board either in our office, in a server rack or even in the cloud, which creates some challenges. Once you want to deploy 100 robots in a warehouse, either you have an amazingly good internet connection or you have to install server racks in that warehouse. So we are hoping that compute keeps progressing such as we can finally feed those things on the robots themselves, typically on a Jetson. And then the VLA, this is where compute is the most limiting today because we cannot really put it off board it still needs to run fairly fast let's say 10 times per second with with minimal delay so it needs to be on board and they also typically use diffusion which means that you don't infer it just once at this part of the network is inferred is inferred multiple steps so this is where compute is the most critical the on-board compute of drawbot is the most critical yeah i hadn't thought about you know in the case of the um the one except we were talking about these other robots in the home that like they're essentially like their brains are in the cloud i assume that you know somehow we were able to get these models small enough to run locally which seemed like a lot but um just from a latency perspective uh that i don't know it's it's hard to imagine that being particularly tenable and consistent as well.

42:03So some of these models, just to clarify, some of these models can fit on the robot. But what we're seeing today is if you want to do some of this more abstract reasoning, so you give it a very abstract task and then it has to orchestrate something for multiple minutes, there you would really benefit from larger models. Yeah, and I would imagine that you would want even more abstract models in the home, you know, for consumer tasks than you would require in an industrial setting. Is that true? Yes, I would say that's true because in an industrial setting, if the task is repetitive, you can more or less pre-compute those very abstract instructions or go from very abstract instructions to clear instructions.

42:54In a home, if you have a human just telling you, something. For sure, you need the scale of large models. Are you using off-the-shelf RL environments or is part of what you're creating the simulation environment for creating these models? It is a big step of what we're creating. So we are not building simulators ourselves. We're using existing simulators, including ones from NVIDIA. We also test and experiment with many others. But one of our key know-how is how to properly build those simulation environments, and we have our own custom RL algorithms on top to benefit from that as much as possible.

43:35Got it. So the simulator and the simulation environment are distinct. Is the simulation environment, when you say that, is that the configuration of the environment in the simulator? Like the simulator is like the platform and the simulation environment is like the thing that you create about your scenario? Or is there... Are there just those two levels of abstraction or are there three levels of abstraction I guess is... I guess you would add the RL algorithm as a third component that interacts with both. But that's it. The simulator itself is basically a physics engine and a renderer. And then you have to put a robot in there.

44:20If you wanted to walk on stairs, you have to create stairs. But you cannot just ask the robot to randomly figure out how to walk on very complex stairs. So you have to create a whole what we call a curriculum of difficulties. So you would start with very small stairs and progressively make them harder. And the same is true for all sorts of tasks. When we're training a robot to open a door, we have to create a simulated version of the door. Then we have to help the robots. We have to figure out exactly all these training processes on top of just the scenario itself. And so as we've talked about this, you've kind of

44:59positioned RL and imitation as these two alternatives. But is it also possible to use imitation in conjunction with RL to kind of bootstrap learning and like, you know, help the robot get over these, you know, figure out stairs more quickly? Like, is that still a research problem or is that something that, you know, we're able to do in practice now? That's a good question. It's still a research problem, but we are seeing good signs of life of it, I would say. I can talk about two different ways to combine them. One way is use a few demonstrations to help the RL process. And this is something we're doing very actively.

45:46So if you have a human showing, like doing the task, you can extract just a little bit of information from that to help the exploration process of the reinforcement learning, such that the robot is not, you know, just randomly shaking and trying to figure out everything from scratch, but you're guiding it a little bit towards the right solution. And a completely different way to approach imitation learning plus RL is what we're seeing a bit more in other companies and in academia, which is doing more imitation learning for pre-training and then adding some flavor of RL on top to try to improve the behavior after the fact.

46:26A big question there is if the imitation learning pre-training was done without any simulator, do you suddenly need to add a simulation again to do the fine tuning or not? And I think this is still a very, very open question. You know, when I think back to, you know, some of the earliest conversations I had in robotics, like, you know, these are folks like Peter Rabiel. And I remember, I think, you know, this was like pre-VLM stuff and maybe even pre-Transformer. I don't remember. But like, you know, some of that earliest work was like they'd have hundreds of robots, you know real robots at google like and they were starting an experiment with rl i believe you know but like it was the robots were rling and it was very expensive it used a lot of you know you needed to have a hundred you know time on a hundred robots um you know but they didn't have to deal with this rl to sim gap um you know you're more focused on you know simulation and other ways to close that gap, but there are still proponents of RL in real life.

47:38How do you think about comparing and contrasting those approaches? There are definitely people who are big proponents of RL in real life, or another way to put it, who are against simulation at all. In some cases, there are good reasons for that, because some things are fundamentally hard to simulate. I can talk about a specific task we're focusing on, which involves the robot manipulating cardboard boxes. The robot has to walk, pick up the box, bring it somewhere, put it on a table, open it, take what is inside out of the box, and put it on a shelf, for example.

48:22In that case, most of the... If you break it down into subtasks, most of them are very well simulatable, except one very specific piece of it, which is opening the box. Because you can imagine there is maybe tape on the box. You have to take a knife, cut through the tape to be able to open it. And it's possible, but it's still a lot of effort to simulate the interaction of the tape with the cardboard and exactly how the knife cuts through it. So what we are trying to do there is identify those specific cases where simulation is still limited and use real data, but only for those very specific cases.

48:59and then mix it with simulated data of everything else. And we think this is how we basically get the best of both worlds. We get as much simulation as possible, and as simulators develop, they'll take more and more of the whole set of tasks. But while there is a gap, we'll reuse real data for those specific cases. You know, for the folks that say that, you know, RL in real is better, you know, in what ways would it be better for that particular scenario? It sounds like you would just go through a lot of boxes, but you still, like, unless you're giving your robot a utility knife, like, it seems like it's the same problem in real, right?

49:44I agree with you. I think it is much harder to do it in real life compared to simulation, especially with reinforcement learning. What you could do is do a little bit of imitation learning for that specific case. And that's much easier than letting the robot learn everything from scratch. I would guess the only argument for real life, for pure real life RL, would be that you don't need to deal with simulation, which can be very, very hard, especially if you don't have expertise in-house in terms of how to create those simulated environments and how to tune the simulators to behave nicely. When everything is in real life, well, you already have the perfect simulator in a way.

50:28But on the other hand, you have very expensive hardware and then any failure and then you reset is way more expensive than in simulation. Yeah, I think it's clear why it's compelling and also why it is aspirational. Like the idea that as humans, we don't simulate the world to learn things. We explore in the world and we learn that way. And so I'd want my robot to be able to do that. But we're nowhere near the sample efficiency in robots as we are in humans. So you would end up breaking a lot of robots and or boxes to get there. And then you've only solved one task. I think another interesting point is that the human reward signal is extremely complicated.

51:23if you think like if you're doing some tasks with your hands, the amount of information you're getting from all the nerves and all your skin, also from your muscles that are tired, etc. is extremely complicated. And we don't have that information at all with a robot. Typically, tactile sensing is very primitive. So if the robot is slowly damaging itself, you wouldn't know until a motor breaks. So if you're doing reinforcement learning in real life, getting rid of behaviors that would damage the motors simply won't work because you'll get maybe one event per week the reward is not there. And this is where simulation helps once again because in simulation we have perfect information about everything.

52:07We can design reward functions that will avoid breaking motors damaging the mechanics of the robot for example. Based on your earlier point about incorporating vision reducing performance or making it more difficult to convert on a model, you know, one approach to that is, oh, let's just add skin or let's add additional sensors. But, you know, any additional sensors or each additional sensor rather increases the burden, the computational burden on these models. Absolutely. Plus, mechanically, it makes everything more brittle. So with cameras, I would say we are there today. We can add cameras to our robots.

52:52They're very cheap. They're very reliable. Tactile sensing is just not there today. And thinking about crafting that reward function, talk a little bit about that process. That is key in any type of RL, is figuring out what that objective function is, a reward function.

53:16And how standardized are they for a given task? Or do they vary very widely and require a lot of hand tuning? And maybe as a secondary question on the language side or in coding agents, there's a lot of talk now about trying to incorporate value functions that allow that provides signal for positive behavior before the end objective. And I'm curious if value functions are a practical thing in robotics today as well, or a conversation. Is it research or is that something that we're using today? Absolutely. Value functions are part of the RL algorithm itself. So we are training, for example, we're mostly using some variant of PPO, which is an actor-critic algorithm, which means that you're training both an actor and a critic.

54:20The critic is basically a value function. It's not used at deployment. It's only used to help the training of the actor during the training process. Now we're seeing some research into how it could actually be used even at deployment. I think this is more on the research side. It's not proven yet. And then about reward tuning, it is a big topic. We have 35 people in a company and quite a few of these people still spend hours tuning rewards. And typically this is obviously referred to in a negative way. You don't want people tuning rewards. But I think you have to make a distinction between two types of tuning.

55:08There are general rewards that just simply come from the task itself. So if we think again about how a robot is, just locomotion, how robots are walking, you would say the task is just, you know, go from point A to point B. But in reality, it's a bit more complicated. You want the robot to go from point A to point B. You don't want it to use too much energy to do that. You don't want it to hit the ground too hard. You probably don't want it to slip everywhere. You don't want the arms doing completely crazy motions around it. So by the time you describe even in text what you actually want, you already have, I don't know, maybe 15 lines.

55:45And so that translates to 15 different reward functions that you have to come up with in Tune. And I think my personal take is that that part of Tune is fine. There is another kind of Tune that tends to happen a lot in RL, which is related to exploration. Once you've described the perfect task that you want, how do you guide the policy, the training process towards that? For example, if you're on a robot that opens a door, you might need to tell it, put your hand close to the handle, then close your fingers, then pull on the door. And these are really things that don't scale across tasks. And this is something we're trying to avoid as much as possible.

56:27And this is why we're working on other techniques where we can use one or two demonstrations from a human to help the learning process instead of all these manually tuned reward functions. You know, we're at the end of the year and this is kind of a natural time for folks to make predictions. Do you have any predictions for the upcoming year or, you know, several years, whatever horizon you'd like to offer in terms of robotics? Like, how do you think about the future? You know, I said before that I think there isn't a single humanoid robot providing value in the world today. My prediction is that will change around the end of next year, maybe beginning of 2027.

57:11So we won't, again, prediction, hard to say exactly what's going to happen, but I don't think we'll get this, you know, chat GPT moment where suddenly you have a billion robots everywhere just because you need to build the hardware. It doesn't scale like getting access to chat GPT, right? But I would predict that around the end of next year, we'll start seeing robots doing actual work. It will be just a few here and there. And then in 2027, 2028, we'll scale both the numbers of robots per task, but also the set of different tasks that these robots can do. Which means that in the coming years, we'll go from, I don't know, hundreds of robots to thousands and then very quickly to millions, tens of millions, et cetera.

57:55And presumably you see that happening first in industrial settings and then consumer? This is my prediction, yes. We'll see that first in industrial, then consumer, so at home. And I hope that after that, we can go to more, to crazier applications. Like we should really send humanoid robots to Mars to build colonies before humans land there. When you think about the currently available robots bots that you have, you know, seen and or worked with, like what are they all kind of the same? Like are all the, you know, the dogs the same, all the human noise are roughly the same, or do you see, you know, big differences between them from a hardware perspective?

58:46And if so, like, are there, you know, ones that are particularly exciting for you now? That's a great question. There, there probably three or four different strategies you can take in terms of how you're designing your human and robot, specifically what kind of actuators you use, what kind of gearboxes. And since there are many companies exploring that space, everything is happening in parallel. No matter which of those strategies you take, there are a few companies in the U.S., maybe one in Europe, and probably 50 in China building that exact thing. So the competition is really fierce. One part where hardware is not there yet today is on the end effectors, on the hands.

59:35It's still debatable if you actually need very dexterous hands. I think one of the big reasons why many companies develop hands with high dexterity, so like more than 20 degrees of freedom in a hand, is because they're using imitation learning. They're imitating humans, which means that you need to be able to imitate everything a human does. Once you go the RL route, you can learn other kinds of behaviors with much simpler grippers. But the debate is still open on that. Again, we've been talking about dogs and humanoids. But there's this broader question, which is, is humanoid the best form factor for a robot?

1:00:15Like, should we be making robots with two arms, two legs? Do you think that that's the way to go? or are we kind of anchored on this because it's our form but there are better forms that you've seen or think about? That's another great question. So I worked on more than 25 robots I think by now. So any number of legs and arms you can imagine from zero to probably four, five, six legs. And there is room for all sorts of robots in the world. Honestly, I mostly use, and as a company we use the word humanoid for the lack of a better word. What we mean by that is not the human form factor, but human capabilities.

1:01:01So very basically, we want robots that can go where humans go and can manipulate their environments in a similar way to how humans do that. So probably you would need at least two arms with some sort of ender factor to interact with the environment. and then whether you have legs or wheels both are fine both have their own applications so you can go a long way especially in industry with a wheeled platform and we're working with robots like that as well having said that it's surprising how quickly wheeled platforms get stuck a very common thing that is easy to imagine is if the floor is not perfectly flat If you have something, you have cables or, of course, stairs, your wheel platform is stuck.

1:01:48But another very important part is also the footprint. With those wheel platforms, you have two choices. Either you make them very large, and then they're stable by default because they have a very large footprint, but they don't fit through tight spaces anymore. And very quickly, especially in slightly older industrial settings, you have tight spaces. And the alternative is maybe some gyroscopic Segway-like thing? That could be one. Typically what we see is that you just have a small platform, which means that you have to be very, very careful how you move the torso on top, because if you lean too far, it just falls over.

1:02:27I haven't seen the gyroscopic platform yet. Okay. I still think a three-legged robot is probably the coolest robot I've seen so far, but it's maybe not the most applicable for industrial tasks on Earth. And then maybe one more question. There are, you know, quite a few now like robotics kits or, you know, robots are getting more accessible for folks that are interested in the space and want to play, but, you know, don't have access to a humanoid robot like or don't have any robots. It's like, you know, what are some cool things that someone can, you know, order now, you know, maybe get by the holidays or soon thereafter, meaning not like pre-order for 2027 and start playing around?

1:03:12Like what, you know, if you were, you know, advising someone who is like excited about getting their hands dirty, like, you know, what would you tell them to start doing? there's an amazing community around Hugging Face and their little robot project where they have very cheap robot arms and they can help you learn about the whole teleoperation data collection training and deployment pipeline with those arms so that's a really good way to learn about that for the locomotion and reinforcement learning aspect it's a little bit harder because you probably want a robot with legs which also means that the robot should be able to fall and stand up without completely breaking.

1:03:55I think the best bet there is the Chinese quadrupeds that are really getting fairly cheap. It's still multiple thousands of dollars, but let's say it's affordable for a university or a school or if you really want to go much deeper into that at home as well. and you get your quadruped and you unbox it, what can you do with it? Or where do you start with trying to do some experiments with it? So when you unbox it, typically it can already do quite a lot. So it will be able to walk. In some cases, they even have things like slam pipelines. So it can do some navigation, can avoid obstacles, things like that.

1:04:46But then the challenge is that you want to get rid of all that software and basically recreate it from scratch. And again, there are many communities online. There are many GitHub reposts that help you get started. The unit we go to is probably the most standard platform, so we started there. And there are people who open source already everything from training to deploying these policies on those robots. Okay. Cool. Awesome. Well, Nikita, thanks so much for jumping on and sharing a bit about what you're up to. Very cool stuff. Thank you so much. I really enjoyed this. Thank you.

1:06:05you

From the publisher

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion Robotics to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today.

The complete show notes for this episode can be found at https://twimlai.com/go/760.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Intelligent Robots in 2026: Are We There Yet? with Nikita Rudin - #760The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 7 min
Listen in VO