Bridging the Sim2real Gap in Robotics with Marius Memmel - #695

30 Jul 2024 · 57 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Summary: Bridging the Sim2Real Gap in Robotics with Marius Memmel - #695

Podcast Information

  • Podcast Title: The TWIML AI Podcast
  • Episode Title: Bridging the Sim2real Gap in Robotics with Marius Memmel
  • Host: Sam Charrington
  • Guest: Marius Memmel, PhD student at the University of Washington
  • Air Date: [Link to Episode](https://twimlai.com/go/695)

---

Episode Overview In this episode, Sam Charrington speaks with Marius Memmel about his research focusing on sim-to-real transfer approaches to develop autonomous robotic agents capable of functioning in unstructured environments, like homes. They discuss Marius's recent papers, ASID and URDFormer, which address the challenges of sim-to-real gaps in robotics and innovative frameworks for improving robot learning.

---

Key Topics Discussed

  1. Introduction to Marius Memmel's Background
  2. Non-traditional background with a bachelor's in business information systems.
  3. Pursued a master's in computer science, leading to a fascination with AI and robotics.
  4. Current research at the University of Washington under Professor Abhishek Gupta and Professor Dieter Fox.
  1. The Challenge of Robotics in Unstructured Environments
  2. Traditional robotics often relies on structured environments (e.g., factories).
  3. Real-world settings (e.g., kitchens) present numerous challenges due to clutter, varied object dynamics, and the unpredictability of household tasks.
  4. Complex tasks like stacking cups involve understanding the unique characteristics of each object.
  1. Sim-to-Real Transfer
  2. Sim-to-Real Gap: The difficulty of transferring knowledge from simulated environments to real-world applications.
  3. Simulation offers a way to generate high-quality, cost-effective data for training models, but discrepancies between simulated and real-world dynamics pose challenges.
  4. Importance of developing autonomous simulation models to avoid manual tuning.
  1. Introduction of ASID Framework
  2. ASID (Autonomous Simulation Identification) enables robots to autonomously create and refine their simulation environments.
  3. The framework is divided into exploration (learning to gather data about the real world) and exploitation (using that data to improve task performance).
  4. Introduces the concept of Fisher Information to measure how sensitive trajectories are to different physical parameters.
  1. URDFormer Framework
  2. Focuses on the automatic reconstruction of kinematic structures, enabling robots to learn from their environments.
  3. Utilizes a combination of RGB images and depth information to create URDF (Unified Robot Description Format) documents.
  4. Leverages synthetic data generated by procedural methods and text-to-image generation (e.g., Stable Diffusion) to train models for URDF predictions.
  1. Application and Future Directions
  2. Potential for combining ASID and URDFormer to allow robots to autonomously construct simulations for various environments.
  3. Explores the integration of Large Language Models (LLMs) for better initialization of simulation data.
  4. Discusses the need for creating high-quality data sets without overwhelming computational demands as robotics technology evolves.

---

Key Takeaways

  • Robotics must transition from structured environments to successfully navigate the complexities of real-world settings.
  • Sim-to-real methodologies like ASID and URDFormer bridge the gap by automating the simulation process and improving learning efficiency.
  • Fisher Information provides a novel metric for assessing trajectory sensitivity and guiding exploration strategies in robotic learning.
  • Future advancements will focus on creating better data generation techniques and exploring synergies between simulation and real-world learning.

---

Notable Quotes

  • "If we think about how we would approach such a system, classic robotics would probably try something like task and motion planning." — Marius Memmel
  • "Constructing simulations on the fly is a valid path... informed by the real world is a promising path." — Marius Memmel

---

Conclusion The discussion highlights the innovative approaches being developed in robotics that aim to reduce the sim-to-real gap, ultimately allowing robots to function effectively in everyday environments. Marius Memmel’s research showcases the importance of automation in simulation construction and underscores the future possibilities for intelligent robotic systems.

---

For complete show notes, visit

[TWIML AI Podcast Episode #695](https://twimlai.com/go/695).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Like in your kitchen, you want the robot to, I don't know, do the dishes and like kind of clean up. your kitchen is very cluttered it's very unstructured and let's say like you wanted the robot to like take a cup and like put it on another cup then it becomes very tricky because like all cups are kind of different some have like a different center of mass or like a different friction basically and then it becomes really important to kind of like understand like how all of these these items like behave so your robot can react to them and like find a good way to place them without like dropping it every time

0:45all right everyone welcome to another episode of the twiML ai podcast i am your host sam charrington and today i'm joined by marius memel marius is a phd student at the university of washington before we get going be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Marius, welcome to the podcast. Thanks, Sam. Thanks for having me. I'm excited to jump into our conversation. We're going to be focused on your research into building AI for robotic agents and Sim2Real as an approach to doing that. Before we dive in, I'd love to have you share a little bit about your background and how you came to study in the field.

1:26Yeah, so my background is actually fairly non-traditional. I got my bachelor's in business information systems in Germany. So I actually have like more of a business kind of like flavor to it. And then I decided to go back to grad school actually to get my master's in computer science from the TO Darmstadt, also in Germany. And then that's the first time I kind of like got into touch with research and like the whole AI community. And I found it like very, very interesting, very exciting. so I decided to pursue a PhD and I initially worked a lot in like computer vision because that's like an easy way to get started but found my way more and more towards like the control and the more like robotic side yeah and finally I kind of like applied to grad school and I ended up at the University of Washington where I'm part of the weird lab led by Professor Abhishek Gupta and the RSE lab led by Professor Dieter Fox.

2:26And yeah, I've been working on robotics ever since. Yeah, that's kind of how I ended up here. Awesome, awesome. Yeah, I interviewed Abhishek, actually it was a while ago, 2021, talking about his research into applying reinforcement learning to real world robotics. And he was a PhD student at the time at Berkeley. So super awesome to have a chance to now speak to one of his students at U of W. Let's maybe get started by having you share a little bit about the way you think about your research agenda broadly and what you're finding most interesting and how you're approaching this idea of robotic agents.

3:11Yeah, so I guess the best way to start with this is what do we want out of robotics and what is the holy grail or the dream robot that everybody wants to have? and in my opinion that's like an agent or a robot that can like function autonomously ideally with no human intervention and we wanted to solve like a variety of tasks so like a lot of the robotics tasks we care about are pick and place so like take the cup and put it somewhere but actually there's like this small niche or there's like a lot of tasks that are more dynamic than just like placing something somewhere because when you think about it um like in your kitchen you want the robot to i don't know do the dishes and like kind of clean up and then pick and place becomes a lot harder because your environment or like your kitchen is very cluttered it's very unstructured um and let's say like you wanted the robot to like take a cup and like put it on another cup um to kind of like be most space efficient then it becomes very tricky because like all cups are kind of different some have like a different center of mass or like a different different friction basically um maybe if you have like pots and pans you want to like stack them like you can't just like put them all on on the table that's kind of like wasting a lot of space and then it becomes really important to kind of like understand like how all of these these items like behave so your robot can kind of like react to them and like kind of like find a good way to place them without like dropping it um every time and if we think about like how we would like approach such a such a system or like how we build such a system um classic robotics would probably try something like a task and motion planning approach uh where you have like information about the objects you have your robot you have a kinematics model and then you start making like higher level plan so like if the task is um like clean up the kitchen or like do the dishes you would start with okay we have these five items uh let's pick up the mug open the cabinet and then put the mug in the cabinet and so on and so forth then you would run a motion planner to like make the robot execute these actions while this is like this works for like very structured environments it doesn't generalize very well.

5:32And it requires a lot of information about where are the objects and how do I do a grasp? And a lot of what we call privileged information that we don't really have in most real-world settings. So that's where the whole robot learning aspect comes in. There is this hope for generalization. So you teach the robot how to pick up a couple of different mugs, And then it can generalize to a new scene, a new environment and new kinds of mugs and like still perform the task. What I've heard there so far is a bit of a classic. The real world is a lot more complex than the traditional robotics models tend to focus on.

6:20So, you know, pick in place, you've got like this perfectly stacked row of boxes and you want to pick one up and move it to another place. and the real world when you're dealing with you know a bunch of misshapen and or uh randomly shaped objects like cups and mugs and you want to stack them up can be a lot more complex and so those traditional approaches don't transfer very well to the real world setting is that the core idea exactly i i would say like the classic approaches do a fairly good job in like a like a most structured environment like warehouses but like once you start moving to people's homes like you can't assume everything's structured so you have to like deal with all like the uncertainty and the clutter and so the approach that you've taken with this kind of falls under the broad category of you know what's known as sim to real which is kind of incorporating simulation into solving the problem.

7:21Talk a little bit about Sim2Real broadly. Yeah, so I guess the big problem of robot learning is where do we get the data from? We know from NLP and computer vision that getting the data is the most important part, and then you're able to train very powerful models. And robotics data is very expensive. You can collect it in the real world, which is expensive. It can break the robot, it can break the environment, and it's just limited. Breaking a lot of glasses and plates trying to do the dishes. Exactly. And I guess RL is the extreme where you put your robot in your kitchen and you say, okay, now try to pick up this glass.

8:01And you just let it run for 10 hours and you just hope that it doesn't break it in the process. So simulation kind of offers this cheap way of getting data, very inexpensive, high quality, and a lot of data. So you basically build a simulation. And you use that as your data engine. And then you can use that data to train any kind of model. You can use reinforcement learning. You can use behavior cloning. Basically, yeah, most of the robot learning techniques. And then you take that model that you trained on the simulation data and you put it back into the real world. And then you hope that it can solve the task.

8:42The problem here is that there's a gap. So you have the simulation on the one side and you have the real world on the other side. and unless you spend a lot of time constructing that simulator to represent the real world it's probably not going to transfer because there's like it either like the observations look differently um you have like a mismatch in the dynamics a lot of all simulators basically make approximations to like the physics so like there's always like a kind of like a mismatch between the two so that's kind of like where where the challenges uh or like where the challenge of sim to real comes in like how how do we kind of like get this simulation as close to the real world as possible so like the traditional way or like what most people do is they just build that simulator by hand but that's like not a very scalable and not a very autonomous approach because if you want the robot to like clean up your your dishes you don't want to walk in and like build a perfectly accurate simulator and then train the robot in simulation you want the robot to kind of like do this autonomously.

9:44And that's kind of like where our inspiration for like ACID came from. Yeah, one question that came up for me in the ACID paper is a pretty basic one. And that is, how do you select the granularity of the task? Meaning your ultimate goal is for the robot to do the dishes. And there's a bunch of atomic subtasks there, maybe grasping and stacking and washing and all these other things in the acid paper and generally in the field, you're kind of choosing something that you think is kind of atomic there. But that varies from project to project. In acid, you're not really focusing atomically on grasping.

10:33You're grasping and moving and stacking. How do you think about defining the task? And do you think, more importantly, that this idea of allowing the agent to kind of construct its own simulation environment will allow for some flexibility in the way that the task is ultimately defined? Yeah, that's a very good question. So for the most part, I think it's like an open question and people are still researching on how to pick the right horizon and how do we combine different so-called skills. And I would say in this project, we kind of focused on the longest horizon that we could comfortably deal with without adding in additional methods.

11:21so like especially if you do like reinforcement learning or in our case like the final policies are basically um parameterized primitives we just like picked something that like would work out of the box so we could focus on like um kind of like narrowing that simtoria gap versus kind of like like opening like this whole box of like how do we get longer horizons out of it got it got it okay but maybe like a different way to look at it is like our method like would give you a simulator um that kind of like represents the world and like the objects you care about so from there on you could actually like use any kind of like long horizon method you can do skill learning in that simulator and then transfer those skills because like now like the idea is basically now you have that accurate simulator and you can use it for whatever purpose for whatever task as long as it involves these parameters and these objects you care about.

12:20So broadly speaking, then you have this agent, you wanted to solve some task in the real world. You've defined it however you define it based on this horizon that you're comfortable with. And then the task is to come up with a simulator that can help you build a model of the real world that you can use to generate data. well, the simulator helps you generate data that helps you create a model that your robot can use to define a policy, right? Yeah, I think like, yeah. And in terms of like Sim2Real, you're absolutely right. Okay, cool. So you've got, or we've established the role of, you know, simulation and generating data that helps you create a model and a policy.

13:04And then the, you know, core issue is like, how do you create this simulation? And even when you do create this simulation, like how do you close the gap between, you know, what the simulation, the data that the simulation creates for you or the model that results in the data from that simulation and performance in the real world? And that is where ACID comes in. Exactly. All right. So tell us a little bit then about ACID and how it changes the approach. Yeah. So it's like a sim-to-real to sim-to-real approach. And that one's split into an exploration phase and an exploitation phase. So in the exploration phase, you really want to learn how to explore.

13:49And you want to explore the real world to get a better simulator. And then using that better simulator, you're entering the exploitation stage where you do sim to real, but to solve the task. And so talk a little bit about how the exploration in the real world complements and builds on the exploration that you've done in the simulation phase? We use simulation to synthesize this exploration behavior and distill it down into a policy so we can then roll it out in the real world. And the way to think about this is that in simulation, we train on, like we do domain randomization and we train over a distribution of parameters.

14:34So we train on a bunch of different center of masses. and then we roll it out in the real world and now because it's trained to identify a variety of center of masses it can like figure out how to like push that rod in the real world or like how to kind of like excite the object a little bit uh to get that kind of information and then we record we record that data basically we should think about this as as kind of two distinct exploration steps. One is in the simulation and this is like an upfront step that is unrelated to the task at hand.

15:18It's unrelated to a specific instance of the task, but it's related to kind of this class of tasks of stacking the rod. You've got this simulation set up. The robot is or the model is presented with a different types of rods, type meaning center of mass configuration. And it essentially has to learn how to manipulate the rod to determine where the center of mass is. And that's the model that it's learning. And then you go to the real part of exploration where you have a specific configuration, a specific rod with a specific center of mass and then the robot needs to use what it learned in simulated exploration with lots of different configurations to narrow in on the dynamics of the real world situation yeah so the the burden is actually a little bit lower on that policy like the only thing you want from the policy is to kind of like push that rod and then you can use that data to run system identification to actually like you run an optimization based technique to find that parameter so like you're only using that policy to get useful data that you can then use to run optimization to find the actual parameters we talked about kind of this two-step sim reel sim reel you did a simulation you got the reel push your we're then in the real we push the rod we get some data then the next part of this is going back into simulation is that where the optimization happens or is that optimization yeah yeah so what you're trying to do you have like these real world observations and you have a simulator so now if you replay the actions you took in the real world you get a corresponding trajectory in the simulator and now you just like change the parameters in the simulator until you perfectly match your real world rollout and at the point where those match you have that exact physics parameter got it and so in the in the real world of exploration part you're just pushing the the robot's just pushing the rod once or is the robot playing with the rod and collecting multiple How much data is the robot trying to collect or how much data do you need from that?

17:56So in our cases, we actually got away with a single rollout. But you still sometimes you see the robot like not do a single push, but like to kind of like a second one or like kind of like push it to one side. So it kind of like learns how to do like kind of like directed behavior in that sense. And so then the robot kind of takes the data points that it's collected back in the simulation. And in simulation, it can do as much manipulation of the rod as it wants to or needs to until it identifies what that parameter is. Exactly. and then you have a policy that can be deployed to the robot and kind of zero shot you know base with this policy pick up the rod stack it and it works a lot better than other approaches have worked yeah so if you want to do a sim to real you kind of like start with a simulator and then you kind of like hand tune it and that hand tuning might be fine for like quasi-aesthetic task where you don't care a lot about physics.

19:02But as soon as the tasks get more dynamic, it's fairly tricky. Let's say you're in a scene or you have set up your dynamic task. How do you measure things like friction or center of mass? The only approach is you either collect a bunch of data and you run system identification or you just hand-tune the parameters until it results in the behavior you're expecting. and then like the problem with system identification is what kind of data do you need to identify the system properly like what kind of data do you need to like find out friction parameters or like the different um like different center of masses and stuff like that other approaches would be something like domain randomization but like without prior interaction you don't know where that center of mass is.

19:54So you either pick it up at a random location and you place it randomly, or you just converge to a single one, which is not very likely to be the actual center of mass. And because our system gets this real-world interaction data, it knows where to pick it up to solve the task. We propose this training objective that gives you a policy that seeks out these trajectories that kind of like show you something or give you some information about the underlying physics parameters. So you can take that data and like just plug it into system identification. So we're trying to like eliminate this like manual process of like hand tuning the simulator until it's right and make it like more autonomous and like let the robot figure out how to explore and then identify the system.

20:45You mentioned that one of the key elements of that is the objective function that you're targeting. Can you talk a little bit about that and why it's important? Yeah, so without diving too much into the details, basically the objective is inspired by the Fisher information and especially the Fisher information with respect to the physics parameters. So let's say you have a system or like a simulator and your parameters are something like mass and friction. And then you roll out a single trajectory in that simulator and the Fisher information describes, or in this case, the Fisher information matrix describes how sensitive that one trajectory is to those parameters.

21:36So maybe to provide a simple example, in the paper we have this rod that has a different center of mass. And depending on the center of mass, it behaves differently when you push it. So if the center of mass is all the way to the right and you push it on the left side it kind of rotates around the center of mass. So if you now have this pushing trajectory and you move that center of mass the final state of this rod will change depending on the center of mass. So this would be like an example of a trajectory of like a high fissure information. and if you like kind of reformulate this into an objective you can apply like an algorithm like a policy search algorithm like ppo to learn how to find those trajectories from data in the the videos that you just referred to you're starting with kind of a robotic arm that is on a you know a table a surface and then you've got one of these rods that has been placed on that surface and you know from the beginning the robot you know knows nothing about this environment except that there's a rod there or you describe what the the the initial setup is like what does the robot know and i was also curious there's a visual component uh in that environment or is there some other uh sensing uh modality for you know where the rod is positioned and all that kind of stuff yeah so that's um that's actually that was our initial goal was to have like a full pipeline to just like do everything um like reconstructing the geometry figuring out what's in the scene and then figuring out the dynamics but we had to like we had to realize really quickly that this is like the holy grail that would be the holy grail like simulation construction and it's actually a lot harder to just like do everything uh in like a single single project and there's like a lot more things that the community has to work on to make this happen.

23:42So in our approach, the whole like the policy runs from state. We have a tracker for the rod and we start with like a simulation that captures the geometry and the kinematic structure. So it has the robot, it has the table and it has the rod. And then the robot basically knows there is a rod. I know the orientation of the rod but now I have to figure out what the center of mass or how I can find the center of mass of the rod and that's where the robot starts and then through interaction it knows that task as well that you have predefined the key system identification parameters meaning it knows that the center of mass is one of the things or the thing that it needs to identify yeah so the policy doesn't know that um but in order to compute the fisher information with respect to those physics parameters you have to like define an initial set um all physics parameters that you care about okay and elaborate on that um distinction the policy doesn't know that but you need to know that in order to compute um yeah so the the reward is the reward depends on the physics parameters you choose and the policy then kind of like gets that reward but there but like so in some of i guess um to elaborate there's like some approaches that do um like adaptation so you give the policy like some latent representation of the current state or like of your current um physics parameters but in our case like the policy doesn't know that but But then that also means that our policy is specific to that set of parameters you choose.

25:35Got it. Got it. So you've got, again, you've got this setup. The policy doesn't know anything about the specifics of the physics parameters, but in the training process earns reward for doing things. And that reward is based on the physics parameters. Should we interpret Fisher information as something analogous to like a Shannon information kind of thing? Like if the robot is taking actions that allows it to learn more, then it achieves a higher reward? Or is it more about, you know, accuracy with regard to the task? So you can think of the Fisher information matrix as like, like you can, you can decompose it or like show that it actually boils down to learning an unbiased estimator that takes in some trajectory data and predicts the true parameters.

26:44So like if you maximize the objective, you're collecting data that allows you to better estimate the underlying parameters. And like if you, if you write out the whole equation, you'll actually see that the fish information matrix contains the gradient of the dynamics, which in this case is the simulator, with respect to the parameters. So basically, if the robot does something, or shows behavior, or produces a trajectory, and that trajectory changes a lot when you change the physics parameter, then you receive a high reward. So maybe illustrate this on an example. like in case of the rod like you put the the rod on the table and if the robot doesn't like push the rod on either side and just like does nothing you can change the center of mass and the scene would look the same like you could roll out the same trajectory a hundred times it would always look the same because the center of mass doesn't change but if the robot now kind of like interacts with the rod like starts to push it now you run the same trajectory but you change the center of mass, you get a different result.

27:53And that's basically what this Fisher information captures. It's like, how much would the trajectory change if we change the parameters? But we just still roll out the same trajectory. So it's like, how sensitive is your trajectory to those parameters? I guess, is it a metric applied to an action or applied to a current state of a policy? Maybe that's another way to ask the question. yeah so we can actually compute it on like a state action pair um so given given a state you apply an action and then how much how much does like this the next state change if we change that parameter but we apply the same action and then yeah we have to like we have to like take the gradient of the simulator and then we can kind of compute it but it's like a it's an a state like a single single time step reward that you get out of this and then you just like sum it up over the full trajectory is the application of fisher information in this context is that novel or is that something that's frequently done uh in problems like this yeah so there's been like uh past research or like on using the fisher information for system and education purposes but um our method is like the the first one that kind of like puts it into like the same to real loop um because like i guess the biggest problem of this is that like you need a lot of data to like make this work like training the initial policy on that reward is like super expensive and we're kind of like the first to show that you can use simulation to train such a policy and then transfer that to the real world.

29:40And yeah, and like the cool thing about this is that like one of the insights that we made is if you only have to transfer an exploration policy, that's a lot easier than like just transferring the task policy. Because like tasks are like the downstream tasks are usually more complex, like pick up like the rod and like balance it on like this tower, versus like the exploration behavior is like actually fairly simple. it's just like reach for the rod and like just like hit it on one of the sides right and that's very likely to transfer versus like picking up the rod and placing it somewhere like you might miss the grasp you might like not be able to perfectly balance it you might like drop it off the table uh like all these sorts of things can happen and the intuition there is that in the latter case with the task policy um the you know the robot is doing something and it's trying things, but it doesn't know why it fails and it doesn't have any kind of guiding principle.

30:42But if you are kind of teaching it this exploration policy, then it is kind of learning a key aspect of how the system works, or it's learning how to identify a key parameter in the system and it can use that information to then solve the task fairly repeatably yeah yeah i guess like it uses the information in terms of like it updates this internal simulator to then have like a better representation of the real world and like this whole exploration step is only kind of like trying to figure out what what to do or what like what's the best behavior to get like good data to actually find out what this parameter does or like what the actual value is and then you can do like a almost a zero shot transfer of your final policy because now your simulator is like so accurate and like represents the real world that the gap is like really really small before we started recording one of the things that you mentioned is that the fisher information is kind of underappreciated as a general purpose reward function that doesn't necessarily require the entire framework of asset as you've presented it can you elaborate on on that a little bit yeah so asset like in general is like this this whole framework of like having an exploration stage versus and having like this exploitation phase but the objective we propose is like much more than that in our case it's just a means to an end to kind of like get this exploration bit uh figured out but what it actually is it's like it's a it's like very reusable it's just like a reward function and it's actually fairly easy to tune because the only thing you have to do is like like tune um um a normalization normalization constant uh to kind of like um like other like sometimes the reward just explodes which makes it hard for your algorithm to learn.

32:47So you kind of have to like normalize it depending on your use case. But you can really use it for like any case where like some form of exploration would be useful. For example, if you have like an adaptation method and you're struggling with like getting good data to adapt, you might just like want to like throw in a little bit of that like Fisher information objective and like teach the policy how to find informative data that you can then like that your policy or your algorithm then can use. And I think it's a really, really powerful method, especially in any SimTrial scenario where you, in some form or another, care about parameters and identifying the situation or the scene you're in.

33:32Got it. So if you've got some kind of framework that has some parameter that is some parameter that kind of drives everything, then Fisher information can be used to teach the agent how to best identify that parameter using with its actions. Yeah, yeah, that's a good way to put it. and then can you talk a little bit about how this approach uh compares to um approaches that try to integrate exploration and exploitation you've you know said that this performs better but you know what have other people done and you know why do you think it performs better yeah so i think there's like probably like two two directions or like directions of like related work one would be like the adaptation literature where you or like something like a rabbit motor adaptation where you pre-train or like you train your agent in simulation over a distribution of parameters basically domain randomization but you give the agent like the ability to identify these parameters on the fly and adapt to them so the agent in the real world would then like start rolling out its policy and from an observation of history or like of the the states that observed in that environment it would figure out like is in case of like a quadruped it's like is the ground like more rugged is it like more slippery and then adapt its actions based on what it has seen before and then the other direction is something like like an iterative like a very iterative approach where you have a simulation and you have the real world and you iteratively roll out the like your policy in the real world you go back to simulation you adjust the simulator and you keep doing that and that like that is very similar to our approach and i guess in the instantiation that we present it's just a single a single rollout of this this iterative process but the crucial the crucial difference is that we have a specific exploration policy and those methods kind of like just use their task policy so you roll out your task policy and then you use that as your signal to update your simulator and that kind of works for for a bunch of like different tasks but when it comes to or like when your task is special in a matter that like you only have a single a single shot in the real world because otherwise like the behavior might be dangerous or you might break something then our approach is clearly superior because like let's say let's say we consider the rod like if you would do like this iterative approach of your task policy you would train a pick up a pick and place policy in simulation and then you roll it out so now what's going to happen is on the first try you're probably not going to get it right so you pick up the rod on the wrong end you place it on the tower and the rod falls so in case of the rod that might not be not not be super crucial but like when you go to your kitchen and like the robot starts picking up the mug and it drops the mug then like the whole the task is over like there's no way to recover from it right so that's where you want like a more like a safer like a more um simple exploration strategy and not like directly execute the task right yeah it seems like the common idea in both of these uh in comparing with both of these alternative approaches is that um by uh learning your parameter and simulation it's less expensive than trying to do it in the real world exactly like if you if you would try these like reinforcement learning based approaches in the real world um it becomes very tricky like it's very expensive and it still is.

37:34There's some cases where it works, but if you put the robot in your kitchen environment, it's just not feasible from a deployment perspective. Yeah, yeah, yeah. Another paper that I wanted to ask you to talk a little bit about is the URD former paper. Can you talk a little bit about the setting there and how it relates to the types of things that we just discussed? Yeah, so the overarching theme with both Acid and UR Deformer is basically that of simulator construction and constructing it autonomously and on the fly. Because that's kind of the limiting factor that holds them to reel back. It's like we don't have the time and the capacity to just build simulators for everything.

38:26we have to find a way to like automatically kind of like outsource that and while ACID kind of like deals with like a mismatch in dynamics so it finds out like the first six parameters your RDFORMER tries to reconstruct the kinematic structure or like the geometry of the scene so as I mentioned earlier like in ACID we kind of like pre-built part of the simulation which is basically objects, tables, etc and your RDFORMER technically can just do that from a single RGB observation. So the whole idea is we take an observation of the environment, ideally an image, and then we use that to reconstruct the scene in simulation.

39:11And we can then use that simulator to kind of like bootstrap policy learning and do controller synthesis. You say ideally an image. I thought I saw something in the URD former paper that suggested that there was depth information as well. Is that the case? Yeah, we use depth information. I guess I said ideally because one could imagine that in the future you have some depth estimation model on top of it. But yeah, that's correct. We use RGB in depth. Got it. So ideally, you won't need the depth information. But today, you're using RGB in depth. Yeah. Got it. And so talk a little bit about URD and what that is.

39:57It sounds like it is a kind of a specification for a simulator, for example. Yeah, so URDF stands for Unified Robotic Description Format. And it's like a common way to kind of like represent the kinematic structure. so a lot of simulators like like pi bullet for example used urdfs um to specify the kinematic chain of the robot and like the environment um and all the objects and you could like in like you can convert it to different formats but basically it like it's a kinematic tree that describes the scene um and if you're like able to predict that document it's usually like like a text document similar to like an xml format then you can you basically reconstructed the scene or like you found a representation that you can load into your simulator that then like kind of constructs the scene got it so if i can take an image of you know using the example from the paper a chest of drawers and pass it through some process that produces a urdf document then now I can construct a simulator that can be used by you know to train a robot to open and close doors for drawers on that chest for example yeah and the important bit is that it's like you can use it to learn or like you train a robot to open and close drawers on that specific instance and that's what makes it so powerful like you send the robot in your room it takes an image of the scene, it reconstructs the kinematic structure, and then you can train a policy that is specific to that environment, which is a lot easier to learn than open any cabinet.

41:50That just requires way too much data. Yeah, got it. And the URD former process or paper incorporates text-to-image generation as part of the the way it generates data to build out this model. Can you talk about the way that that is incorporated? Yeah. So like the big problem or like, I guess the model at the basis of this whole project is like an inverse model where you're trying to go from an RGB image of the real world to like this URDF. But that data just doesn't exist. Like unless you hire like a bunch of like designers that reconstruct or model the real world to generate your data, you're going to have a hard time training that model.

42:40So the insight we actually made here is that we can formulate this as an inverse model and we can use synthetically generated data to train that model. And that's where this image generation comes in. We procedurally generate a bunch of simulation environments basically gives us the urdf plus like a simulator we can like spin up and and take like views from different directions right and then we use stable diffusion uh to i would say like re-skin um these images so you give it something that looks like a simulator and you tell it oh make this look like a modern kitchen and it will generate um that's also why we we need depth to kind of like um just like like have the structure in place and like force the model to just like not generate arbitrary kitchen images and we can then use that to generate our data set so we're basically generating a bunch of scenes and simulation and use stable diffusion because stable diffusion is trained on a lot of like real world images we use that to generate realistic looking data to train our model and is constraining stable diffusion based on depth information was that already a solved problem or is that part of what you did in the paper um no there's like um there's models out there it's uh depth guided stable diffusion no calm that already know that yeah okay uh so So are you starting with a small set of URDFs or are you generating from nothing the URDFs that you then kind of skin with stable diffusion?

44:37We have like procedural generation. So you can write like a bunch of rules and then basically iterate all those rules. And then you kind of like, yeah, like you start with like a cabinet and then you add another to the right. and then you add like an oven or something and like you can like iterate over those rules and basically generate um a lot of like these urds okay and your cabinet generator might take in the number of drawers or something like that and so you can kind of build up these worlds based on rules and then you've got um this you know maybe think of it as like a wireframe or like a basic shape or environment.

45:19And then with some additional text description, you can have StableDiffusion generate a bunch of different versions of real-world representations that you can then use for your inverse problem. Yeah, exactly. And this procedural generation is a common method. like people use it to generate a lot of simulation data the big problem with it is like it's constrained by those rules so it's really hard to get um like the real world data distribution where like like different kitchen layouts like it's really hard to capture that in like predefined rules so our method's basically using these procedural generation methods to get the data for you already former and then what we what we can do is we we just like crawl a bunch of like real world kitchen scenes or real world living rooms and then transport those to the simulation and now you have you have the exact real world data distribution of how kitchens look like and not just procedurally generated ones where it's like when you're trying to capture the real world distribution with some like handwritten rules but it's not that they're not they will never be like exactly the same like how actual kitchens look like right do you have to worry about you know kind of unnatural artifacts of the generation process like i don't know what the you know the kitchen version of you know hands with six fingers are uh in stable diffusion but do you just assume that you know given enough data um you know that's noise and you don't need to worry about it or you know do you do you do some filtering pre-processing of images or how do you think about that yeah um the equivalent of six fingers in the real world for cabinets is definitely knobs and handles um it happened because they're so small it happens a lot but like stable diffusion will just like like make it look like a very modern kitchen where like there's just no handle um and then that is really hard um so yeah we we do a bunch of filtering um i think we ended up in some cases even like just generating textures and like reskinning them to like make sure the underlying um like properties like handles are preserved so then you're able to you know the ultimate goal here is kind of bootstrapping yourself to this urd former you know generator that can take images uh and spit out urdf documents and how many of these uh synthetic images did you need to create in order to produce a reliable urdf former yeah that's a good question i think we were on the order of like hundred thousands around around about that um but like basically it's a like it's limited by compute so like you can just like spin up that that pipeline and generate as many as you want um and like use any kind of procedural generation technique that you want um yeah so it's not very limited in that sense and the the model that you use for the urd former is that it's like a transformer based model uh yeah it's it's transformer based uh we have a pre-processing step uh for some of the cabinets where like we use grounded dino um to do bounding box detection um and then feed like those features into that transformer because it's a lot easier to reconstruct if you know like where handles are in the image or like where um like doors are basically uh so the model doesn't like have to learn how to like segment things and then um predict the urdf but like basically you give it like some additional information that you get from like a like a v11 basically well in the context of language models i don't know if you run into the same things but for example when you try to get language models to spit out structured documents like you know json you know it can be difficult and oftentimes folks will use preprocessors or linters or various things to try to enforce structure.

49:42Did you run into similar challenges trying to ensure that the documents that you generated were semantically correct? And did you just throw out ones that weren't or did you apply some set of rules to enforce that structure? How did you deal with that? Yeah, that's a good point. We actually tried to use, I think, GPT-4 came out in the middle of the project. And we just tried it out of the box to see. Because we didn't think about language as an output modality. So we tried GPT-4 and it did a horrible job at predicting URDFs for pretty much anything. like it it maybe got right that like oh a drawer has like a handle and kind of like a drawer but like that's as good as it got so we actually settled for like predicting kind of a tree structure so like we kind of format the output in a way that resembles a valid URDF and then basically just predict like sizes and type of the different meshes.

Read the full transcript

50:56And this way we kind of like make sure that the output URDF is actually valid and can be used in the simulator. And I should have asked this earlier, but how robust or descriptive is the URDF? For example, I'm thinking about a cabinet. It's got a handle. That handle has a shape. uh you know but then once you you know interact with that handle you know the drawer has a depth the drawer you know depending on its slides has some you know let's simplify it to like friction um and you know but you know there's various kind of dynamics is that does urd specify all of that or is it just about kind of the exterior shape and then the you know the dynamics are inferred somewhere else in the simulator.

51:53Yeah, no, you can basically define all of those. So a scene is deconstructed in different meshes. In that case, it would be one side of the drawer or a handle. And then you can kind of, depending on how you define your document, you define the position with respect to the other parts. And then you have different joints. So you have sliding joints. You have rotational joints. You have free joints. And then you can also define these physics parameters, I guess, depending on the downstream simulator you're using. And you already form it doesn't give you those. But what we ended up doing is we just initialize with an approximately good value.

52:37And then during policy training, you can just do domain randomization on top of it. And it actually turns out that for drawers and doors, it's not that big of a deal because it's almost quasi-static and you don't really have to rip it out. And the robot can kind of deal with higher and lower friction. So in that case, it actually works quite well. Interesting, interesting. And then, so these are separate papers. Are they, how do they come together? Or have you done a project that kind of uses them end-to-end as part of a single exploration? Yeah, we haven't combined them yet. So I was the lead on the asset paper, but on the RDFORMER, actually Zoe Chen and Aaron Walsman kind of like we're driving the most of the effort.

53:33I think we have yet to combine the two ideas. I think there's a lot of opportunity for sending a robot in your home and like making it figure out how to open a door, put stuff in the drawer, stuff like that. Yeah, so far we have like, the projects exist as separate entities. Zoe has actually taken a robot into her bedroom to like run UR Deformer in like a real world kind of setting. But yeah, we are thinking about like combining it to kind of like show that if you combine like dynamics and kind of like geometric reconstruction, that you're like one step closer to constructing simulators on the fly and then having like sim to real basically as an autonomous option for like robot learning and solving those kind of tasks.

54:24Yes, I'd love to have you share a little bit about your perspective on where this is all going. This is kind of a fast moving time, both in AI broadly as well as robotics. We recently saw like pretty impressive demos like figure and other things. And it seems like every few weeks, there's a pretty cool looking demo of robotics. Like, how do you see your work playing into that and where it's all heading? Yeah, so I think what like, especially these two projects have shown us and some other projects too, is that we can like constructing simulations on the fly is kind of like a valid path. And especially constructing information or like constructing simulation informed by the real world is a promising path.

55:09I think the interesting part is how can we leverage these giant VLMs to maybe initialize the simulation with different sorts of priors and maybe to use them to bridge the gap a little bit better. And then if we have these simulators, how should we actually use that data? Because let's say In natural language processing and computer vision, it's common understanding that data quality can be an issue. And depending on how you choose your data, it gives you better models with better capabilities. Because in robotics, we not yet have enough data. We pretty much use all the data we have. But I think with better simulators, we will run into this issue of having almost too much data.

56:01and if we don't want to train models for three weeks that can pick up 10 different mugs, we have to figure out a way how to generate better data and how to train on that data, basically. Yeah. Awesome. Well, Marius, thanks so much for jumping on and sharing a bit about what you've been working on. Very interesting stuff. Thanks for having me and thanks for the interesting conversation. Awesome. Thank you. Thank you.

From the publisher

Today, we're joined by Marius Memmel, a PhD student at the University of Washington, to discuss his research on sim-to-real transfer approaches for developing autonomous robotic agents in unstructured environments. Our conversation focuses on his recent ASID and URDFormer papers. We explore the complexities presented by real-world settings like a cluttered kitchen, data acquisition challenges for training robust models, the importance of simulation, and the challenge of bridging the sim2real gap in robotics. Marius introduces ASID, a framework designed to enable robots to autonomously generate and refine simulation models to improve sim-to-real transfer. We discuss the role of Fisher information as a metric for trajectory sensitivity to physical parameters and the importance of exploration and exploitation phases in robot learning. Additionally, we cover URDFormer, a transformer-based model that generates URDF documents for scene and object reconstruction to create realistic simulation environments.

The complete show notes for this episode can be found at https://twimlai.com/go/695.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Bridging the Sim2real Gap in Robotics with Marius Memmel - #695The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 57 min
Listen in VO