In short
NVIDIA AI Podcast: Episode 224 - How Two Stanford Students Are Building Robots for Handling Household Chores
Episode Overview In this episode, hosts Noah Kravitz speaks with Chengshu Eric Li and Josiah David Wong, Ph.D. students at Stanford University, about their project, BEHAVIOR-1K. This initiative aims to develop robots capable of performing 1,000 household chores, utilizing the NVIDIA Omniverse platform and advanced learning techniques. They discuss their experiences, breakthroughs, and challenges in creating robots that can help with everyday tasks like cleaning and cooking.
Key Participants
- Noah Kravitz: Host
- Chengshu Eric Li: Fourth-year Ph.D. student at Stanford, part of the Vision and Learning Lab.
- Josiah David Wong: Third-year Ph.D. student at Stanford, also part of the Vision and Learning Lab.
Subject Matter Background on the Project
- BEHAVIOR-1K: A project aimed at establishing a benchmark for household robotics that includes tasks like:
- Picking up fallen objects
- Cooking
- Folding laundry
- Cleaning up after parties
Development Techniques
- Simulation vs. Real World:
- The team uses the NVIDIA Omniverse platform for simulation due to the limitations of current hardware.
- Benefits of simulation:
- Safety: Avoiding risks of injuries in real-world testing.
- Cost-effectiveness: Reduces the need for physical hardware iterations.
- Reproducibility: A standard environment for different researchers to test and validate their results.
Robotics Learning Approaches
- Reinforcement Learning:
- Robots are trained through trial and error, receiving rewards for correct actions and penalties for mistakes.
- Imitation Learning:
- Robots learn by observing and mimicking human actions.
Implementation Insights
- Modular vs. End-to-End Learning:
- Modular approach: Learning individual tasks separately and then combining them.
- End-to-end approach: Training robots to complete a full task in one go.
Current Progress
- The team is at an early stage of implementing their project, with initial success in simple tasks like throwing trash away.
- They aim to fine-tune the performance and involve the research community by hosting challenges and making tasks available for public testing.
Challenges in Robotics
- Complex Tasks:
- Tasks like folding laundry and cooking are particularly difficult due to:
- Dexterity required for physical tasks.
- The need for intuitive understanding of materials and processes (e.g., cooking requires knowledge of chemical reactions and timing).
- Safety and Reliability:
- Ensuring robots can operate safely in human environments without causing harm or damage.
Future Outlook
- The speakers are cautiously optimistic about the future of household robots, projecting that while advancements will come, widespread adoption may take time due to high expectations from the public.
- They foresee incremental improvements in robotics, starting with more predictable environments such as warehouses before extending to homes.
Call to Action
- Interested individuals can explore the project further at [behavior.stanford.edu](http://behavior.stanford.edu) to try out demos and learn more about the ongoing research.
Conclusion This episode highlights the innovative work being done by the Stanford team to revolutionize household chores through robotics. They aim to bridge the gap between simulation and real-world application, contributing significantly to the field of robotics and AI.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. We're recording from NVIDIA GTC 24, back live and in person at the San Jose Convention Center in San Jose, California. And now we get to talk about robots. With me are Eric Lee and Josiah David Wong, who are here at the conference to help us all answer the question, what should robots do for us? They've been teaching robots to perform a thousand everyday activities. And I, for one, cannot wait for a laundry folding assistant, maybe with a dry sense of humor, to become a thing in my own household. So let's get right into it.
0:44Eric and Josiah, thanks so much for taking the time to join the podcast. How's your GTC been so far? I know that you hosted a session bright and early on Monday morning, had a couple days since. How's the week treating you? Yeah, thanks, Noah. Our GTC has been going really well. Thanks for inviting us to the podcast. We had a really great turnout yesterday. People have been very engaged, and people also ask a bunch of questions towards the end in the Q &A section. I guess people join because they're really tired of household toys. Who's not, right? Yeah, common problem. So before we get a little deeper into your session and what you're doing with training robots, maybe we can start a little bit of background about yourselves, who you are, what you're working on, and where.
1:24And we'll go from there. Yeah, my name is Chung-Shu Lee, and I also go by Eric. I'm a fourth-year PhD student at Stanford Vision and Learning Lab, advised by Professor Fei-Fei Lee and Silva Samarazic. In the past couple of years, I've been working on building simulation platforms and developing robotics algorithms for robots to solve household tasks. Yeah, and I'm Josiah. Similar to Eric, I'm one year behind him, so I'm a third year PhD student, also advised by Fei-Fei Li. And similar to him, I've also been working with this behavior project that we're going to talk about today for the past couple of years.
1:54And I'm really excited to, I don't know, see robots working in real life, and we're hoping that this is, you know, a good milestone towards that goal. Excellent. Before we dive in, for those out there thinking like, oh, these guys are studying right at the heart of it all, and they're in the lab, and they've got these amazing advisors, I'm going to put you on the spot. one thing people might find surprising or interesting or fun about the day-to-day life of a PhD student researcher working in the Stanford Vision Lab. One interesting thing. I think for me, I'll say one interesting thing is that I didn't expect it to be this collaborative.
2:29It might be unique to our project, but I sort of imagine that, you know, you do a PhD, you sort of just grind away on your own, like in a sad corner of the room with no windows. And, you know, you're just not seeing any sunlight. But our room is really beautiful. and I think we get to hang out. So I think I'm really lucky to have people that I can call my friends as well as lab mates and we also get to work closely together. So I think it's something I wouldn't have thought that I would have at a place like Stanford, I guess. That's awesome. Exactly, I want to echo that. I think it's part of because of the nature of our work, which is really a very complex and immense amount of work that we have to assemble a team of a few dozen people, which is very uncommon in an academic lab setup.
3:07So it feels to me that it very much works like a very fast-paced startup where people share the same goal and people have different skill sets complementing each other. So, yeah, I think we had a great run so far. Very cool. Community is always a good thing. So let's talk about your work. Should we start with the session or do you want to start further back with the work you're doing in the lab and what led up to the session? What's the best way to talk about it? Yeah, I guess we can start maybe two, three years back when we first have this preliminary idea of what our project is, which is called behavior.
3:43I think our professor, Fei-Fei Li, had this amazing benchmark called ImageNet from before in a computer vision community, essentially accelerate the progress in that field and essentially set a benchmark where everybody can compete fairly and in a reproduced way and push the whole kind of vision field forward. I think we were seeing that in the robotics field, on the other hand, things, because of the involvement of hardware, each academic paper seems to be a little bit kind of segregated on their own. They will work on a few tasks that are different from each other. And it's really, really hard to compare results and kind of move the field forward.
4:20So we started this project thinking that we should hopefully establish a common ground, a simulation benchmark that is very accessible, very useful, everybody can use. It has to be large scale so that if it works on this benchmark, hopefully it shows some general capability. And it should be human-centered. The robots should work on tasks that actually matters to our day-to-day life. It shouldn't be like some very contrived example that us researchers came up with. And in fact, maybe nobody cares about. So that's very important. So we set up to do this benchmark that we have been working on for the past couple of years.
5:00Why a simulated environment? Why not just start working with robots, training them out in the real world? And as, you know, the hardware and the software and the systems that drive the robots, capacities increase, you can do more. Why work with simulations? Yeah, that's a great question. I think we get this all the time. And there's a couple of answers, I think. One is that I think, to Eric's earlier points, I think the hardware is not quite where the software is currently. So we have all these really powerful brains with, you know, chat, GPT, and stuff that can sort of generate really generalizable, really rich content.
5:30But you don't have the hardware to support that yet. And so I think part of the issue is that it's expensive to sort of iterate on that. Sure. And along those lines, I think, there's, you know, the safety component where because, like Eric was mentioning, a lot of the tasks we want to care about are the ones that are human-centric. It's like your household tasks where you want to fold laundry or do the dishes or stuff where you would probably have humans or multiple humans, you know, in the vicinity of the robot. And you don't want, you know, a researcher to be like trying to hack together an algorithm and then it just, you know, lashes out and it hits you, you know, on the face.
5:58And that's just, you know, you're going to get sued to the ground. So I think simulation provides a really nice way for us to be able to prototype and sort of explore the different options, similar to how these other foundation models were developed sort of in the cloud and then deploy them in your life once you know that they're stable, once you know that they're ready to be used. And like Eric mentioned earlier, like I think there's this aspect of reproducibility where if you all are using the sort of same environments, then you know that the results can transfer and you can be validated by other labs and other people.
6:23Whereas you build a bespoke robot and you say it does something and you can't really validate it unless you buy the robot and, you know, completely reproduce it. So yeah, a few different benefits, we think that are pretty important. Now there are existing simulation engines, I don't know if you'd call them that, but game engines, Unreal, Unity, that are used beyond game development, obviously. and you can simulate things in those environments. Why not go with one of those? Right, yeah, another great question, I think, is the natural fault that a lot of people ask. I think there is a couple limitations with the current set of simulators that we have.
6:57On the one hand, I think you have sort of the, like you mentioned, the very well-known game engines like Unity and stuff. And I think the problem is that you get really hyper-realistic visuals. I think it's, you know, you get these amazing games that are really immersive, and it feels like real life. But I think when it comes down to the actual interactive capabilities, like what you can actually do with your, you know, PS5 controller or whatnot in the game, I think it's definitely curated experiences by the developers. And so there's a clear distinction between what you can and can't do. And that's not how real life works, right?
7:28Where like, you know, there could be a tape that says, do not caution, do not answer. But you can just, you know, walk through that tape in your life and, you know, you don't have to take the consequences, but you can still do that. And I think that's what we want robots to be able to do where we don't, again, to Eric's real point, like we don't want to predefine a set of things. that we wanted to teach it. We wanted to learn a general idea about how the world operates. And so I think that necessitates a need for a simulator where everything in the world is interactive, where it's a cup on a table or a laptop or a door.
7:55And so there's no sort of distinction between like, okay, we curated this one room. And so this is really realistic and it works really well. But if you try to walk outside of it, then it's going to not work. We want it all to work. Yeah. And so did you build your own simulator? It's more accurate to say we built on top of a simulator. And I think this is where we have to give NVIDIA so much credit where, you know, they have this really powerful ecosystem called Omniverse, where it's sort of supposed to be this one-stop shop where you can get hyper realistic rendering. They have a powerful physics backend.
8:24They can simulate cloth. They can simulate fluid. They can do all of these things where, you know, it's stuff that we would want to do in real life, you know, like fold laundry, you know, pour water into a cup, that kind of stuff. And so they provide sort of the core engine, let's say, that we build upon. And then we provide additional functionality that they don't support. And I think together it gives us a very rich set of features where we can simulate a bunch of stuff that robots would have to do if we want to put them in our households. Anything to add, Eric? Yeah, no, I think I do. We really want to extend our gratitude to the Omniverse team.
8:56I think they have hundreds of engineers really putting together these physics engine, rendering engine that works remarkably on GPUs that can also be parallelized, which is actually in our next step, in our roadmap, to make our things run even faster given its powerful capabilities. And it's just impossible to do many of these household tasks without the support of this platform. You mentioned Omniverse, obviously. And so there was a simulation environment called, was it called Gibson, iGibson? And then you extended that to create OmniGibson? Am I getting it right? Right. Yeah, we definitely, so iGibson, just for the audience, is a pretty size sort of OmniGibson that we developed three, four years ago.
9:40And at that time, Omniverse hasn't launched yet. So we used, we wrote our own renderer and then we were building on a previous physics engine called PyBullet, which works very well for rigid body interactions. And then as Omniverse was launched at that time, that was two years ago, that's also when we decided to kind of tackle a much larger scale of household tasks. We decided to work on, for example, 1 ,000 different activities that we do in our daily homes Then we quickly realized that it has gone beyond the capability of what our previous physics engine can do. It doesn't handle fluid. It supports some level of cloth, but it's not very realistic.
10:18The speed is sort of slow. Now we see this brand new toy, I guess, that came right out of the oven. And we thought, let's try this out. So we pretty much kind of started clean from a new, built a new project on top of Omniverse. Many, many things can change. We do inherit some of the design choices that we already made in iGibson that are proven in history to be working quite well in our research world. We inherit a lot of ideas, but we also change a bunch of stuff to make things more usable and more powerful as well in the OmniGibson. So let's talk about robots doing chores. How does one go about training a robot, whether in a simulation or in the physical world, to learn how to do household chores?
11:00Can you walk us through a little bit of what that's like? Oh, that's a great question. Very open-ended question. It's what I do. I ask the open-ended questions and sit back. I think to make it easy for the audience, you can think of it as two maybe broad fields that are generally tackled right now, where one is essentially you throw a robot in and you sort of let it do what it wants. And it's sort of, you can think of it as maybe learning a bit from play where you give a reward. So you can think of like teaching a child, like, you know, they don't really know what to do. And so you have to sort of give them, you know, you punish them when they don't do something good, you give them like a timeout.
11:31And then when you do something good, you know, you give them like a cookie or, you know, some kind of reward. And it's similar for robots where like you throw them in and then naively, the AI model doesn't know what to do. So it just tries random things. It tries touching table, tries like, you know, touching a cup or something. But let's say what you really wanted to do is to, you know, pick up the cup and then pour water to something else. And so you can reward the things where it's closer to what you want to do. Like if it touches the cup, you can like give it a good reward, like a positive reward.
11:55And if it like says, knocks over the cup and spills the water, you give it a negative reward. And so that's one approach I think that researchers are trying. And another approach is where we learn directly from humans, where a human can actually, let's say, teleoperate. So, like, let's say you have a video game controller and can control the robot's arm to actually just directly pick up the cup, pour some water into something else, and then they call it a day. And then the robot can look at the data it just collected and sort of train on that, saying, okay, I saw that the human moved my arm, did this, and, like, sort of poured it.
12:21So I'm going to try to reproduce that action. And so it's these two different approaches where it's sort of scaffolding directly from scratch versus scaffolding based on, like, human intelligence. Right. Yeah. And if you're stringing together a series of actions, like let's say, I mean, even your example of picking up the cup and then pouring the water into a different vessel, is it one sort of fluid sequence or is it, are you teaching sort of modular tasks that you then can string together? Yeah, it's another design decision, right? Like I think there's something called task planning where you can imagine that every individual step is a different training pipeline.
12:56So like, I'm just going to focus on learning to pick the cup and I'm not going to do anything with it, but I'm just going to repeat that action over and over. And then let's say you can plug it in with something else, which says, okay, and I'm going to do like a pouring action over and over. And then if we just string them together, then maybe let's say you can get the combination of those two skills. But others have looked at sort of the end-to-end, what we call process where, you know, you look at the task level, where it's just like pick up the cup and pour it into another vessel. And you just try to do it from the very beginning to the very end.
13:21And I think it's still unclear which way is better. But again, it's a bunch of design decisions and there's a ton of bad luck. I agree. I agree there's no really a consensus. I think researchers have really been poking here and there and trying their luck. And there's pros and cons on both sides. For example, if you do the end-to-end approach, if it works, it works really well. But because the task is longer, it's more data hungry, it's more difficult to convert. On the other hand, if you do a more modular approach, then each skill can work really well. But the transition point is actually very brittle, right?
13:54You might reach some bottleneck where you try to chain a couple of skills together, and then it breaks in the middle, and then it's very hard to recover from there either. So I think we're still figuring this out as a community. What were some of the hardest household tasks for the robots to pick up or easiest ones or even just sort of the ones that kind of you remember because it was interesting, the process was sort of unexpected and interesting? I was going to say the folding laundry example I mentioned is one that maybe it's just, you know, the platforms I hang out on, the algorithms know that I don't like folding laundry.
14:30And I'm terrible at it. I can't fold a shirt the same way twice. But every once in a while, I feel like I'll see a video of a system that's gotten a little bit closer, but it seems to be a difficult problem. Yeah, it's really challenging. I think, to be clear to the audience, we haven't solved all the 1 ,000 tasks. That's our goal also. I think the first step is just providing these 1 ,000 tasks in a really reproducible way in a platform so that you can actually simulate them. But for me personally, I think what immediately comes to mind is one of the top five tasks. So to give a bit of context, like Eric mentioned, we don't want to just predefine tasks.
15:03We want to actually see what people care about. So we actually like pulled a bunch of thousands of people online and we asked them, you know, what would you want a robot to do for you? And so we had a bunch of tasks, more than a thousand, and we whittled them down to a thousand based on the results that people gave. And one of the top five tasks was clean up after a wild party. And so the way we visualized in our simulator was we had this, you know, living room and just tons of glass bottles, beer bottles, like, you know, like just random objects scattered on the floor. And that's just a distinct memory in my mind because I think it really sets the stage for like how much disdain we have for very certain tasks.
15:35And it was clear that people like ranked it very highly because it's, you know, it's very undesirable to do that or clean laundry or, excuse me, fold laundry. I'm getting flashbacks now and I'm wondering if you have taught a robot to patch a hole in the wall. Oh God. But that's a story I'm not going to get into. A hole that the robot maybe made himself when he was trying to do something in the real world. Exactly, exactly. Make a mistake. Yeah, any thoughts, Eric? What do you think? Yeah, I guess some of the cooking tasks seems pretty difficult. Oh, yeah. A portion of our household tasks are cooking-related.
16:03And we did spend quite a lot of effort kind of implementing these complex... I guess we tried to do a bit of simplification, but we want to get the high-level kind of chemical reaction that's happening in a cooking process, for example, baking a pie or making a stew, for example, those kind of things in our platform. And these tasks are pretty challenging too, right? You need to have this kind of common sense knowledge about what does it take to cook a specific dish? What are the ingredients? How much you put in? You don't want to put in too much salt, also not too little salt. You need to understand how much time you put into the oven, how long to wait, and make sure you don't spill anything else.
16:39Yeah, that's some of the longest horizon tasks. Forgive me. I'm sure there's a better way to ask this question. But what's difficult? I can imagine it's incredibly difficult. Well, what's difficult about cooking for the robot to learn? Is it that there's so many steps and objects happening? Is it something about the motions involved? No, you're asking a very brilliant way. I think both, actually, both are kind of the... There are both two challenging aspects. One of them is that it involves many concepts, or like symbolically, you can think of it involves many types of objects. And you kind of chain them together, make sure you use the right tool at the right time.
17:18and also the motions are difficult. Imagine you need to cut an onion into small dices to make some sort of dish. Can't come off a good example. But then the motion itself is very dexterous, right? Imagine that. Sometimes humans cut their fingers when they're cooking. First thing I thought of. Exactly. That's pretty tough. And then I think also it just needs to have some understanding of things that aren't explicit. Like if I put these two chemicals together, it actually creates something third that you didn't see before. and I think a lot of times in current research you sort of assume that the robot already knows everything and so what it's given they can only do stuff like combinatorially but I think cooking is an interesting example where you know you put in I don't know dough into the oven and outcomes like it just transforms into bread and like I think it's it's you know there's the joke about you know you put in a piece of bread and out comes toast and then there's like a comic where you know Calvin from Calvin and Hobbes is like oh like I wonder how this machine works it just somehow transforms it into this new object and like I can't see where it's stored right right And so I think the idea that a robot has to learn that is also quite challenging, too.
18:18I'm speaking with Eric Lee and Josiah David Wong. Eric and Josiah are PhD students at Stanford who are here at GTC24. They presented a session early in the week. We're talking about it now. They're attempting to teach robots how to do a thousand common household tasks that humans just don't want to do if we can help it, which is just one of the many potential future avenues for robotics in our lives. But it's a good one. and I'm looking forward to it. One thing I want to ask you about, LLMs are everywhere right now. And, you know, a lot of the recent podcasts and guests I've been talking to and just people I've been talking to at the show are talking about LLMs as relates to different fields, right?
19:02Scientific discovery and genome sequencing and drug discovery and all kinds of things. There's been some stuff in the media lately about some high-profile stuff about robots that have an LLM integrated, chat GPT integrated, so you can ask the robot in natural language to do something, it can interact with you, that sort of thing. How do you think about something like that? From the outside, I sort of at the same time can easily imagine what that is, but then my brain almost stutters when I try to imagine, like I've used enough chatbots and text-to-image models and that kind of thing to sort of understand, you know, I type in a prompt and it predicts the output that I want.
19:48When we're talking about equipping, you know, a robot with these capabilities, is it a similar process? Is the robot, when we were talking about cooking, I was imagining, you know, can an LLM in some way give a robot the ability to sort of see the larger picture of, you know, now remember when you take the dough out, it's going to look totally different because it's become a pizza. Is that a thing or is that me and my human brain just trying to make sense of just this rapid pace of acceleration and this thing we call AI that's actually touching so many different, you know, disciplines all at once?
20:26Oh yeah. I do think the development of large models, not just large language model, but also like large language vision models will really accelerate the progress of robotics. And people actually, researchers in our field, have adopted these approaches for the last two years. And things are moving very fast. And we're very excited. I think one of the challenges is that what these LM have been good at is still at the symbolic level. So you can think of in the virtual world, it has these concepts. It knows what ingredients to put in, like, margarita pizza, for example. But there's still this low-level, difficult, robust skills, motions, you would call it, how to roll a dough into a flattened thing, how to spring stuff on the pizza so that it's evenly spread out.
21:17All those little details are the crux of a successful pizza. Edible pizza. Even edible, right? I hope the listeners can hear it. I can see in your face as you're talking the, like, you have to get this right on the internet. And I'm with you. Yeah. And so the actual physical implementation of doing those motions is something that, you know, the robotics field, I'm sure, has been working on, but a work in progress. Exactly. I think you hit the nail exactly on the head where it is the execution where you can think of it as, you know, theoretical knowledge. Like, you can, if you're a human, like, the same thing.
21:57Like, okay, you're planning ahead. You're planning the chores that you do. So you, like, list them out. And you know exactly what you're supposed to do. but then you actually have to go and execute them. And so I think the LLM, because it's not plugged in with the physics simulator, it doesn't actually know, okay, I think that if I do this, if I pick up the cup, then it will not spill any water. But if the cup has a hole in the bottom that you don't see, and then you do, and then stuff falls out, then you have to readjust your plan. And I think if you just have an LLM, you don't know exactly what the outcomes are going to be along the way.
22:26And so like Eric was saying, I think it needs to sort of, we say like closing the loop, so to speak, where you plan, and then you try it out, and then you plan again. And I think with that extra execution step, I think is something that's still sort of an open research problem that we're both hoping to tackle. Right. And so where are you now in the quest for a thousand chores, to put it that way? Is it all in a simulation environment? Are you having robots in the physical world go out? And have you gotten to the point where the experiments feel stable enough to try them out in the physical world?
22:59Where are you on that timeline? So when we originally posted our work, which was a couple of years ago, and we've done a bunch of work since then, one of our experiments was actually putting up what we call a digital twin, which is we have like a real room in the real world with a real robot. And we essentially try to replicate it as closely as we can in simulation with the same robot virtualized. And I think we were able to show that with training the robot at the level of telling it, okay, grasp this object, now put it here, and then having it learn within that loop and simulation, we could actually have it work on the real robot.
Read the full transcript
23:31So we tested it in the real world, and we did see non-zero success. So I think the task was like putting away trash or something. I think so, yeah. So we had to throw away like a bottle or like a red Solo cup into like a trash can. And so that requires like navigation, moving around the room, picking up stuff, and then also like moving it back and then also like dropping it in a specific way. And so I think that's a good signal to show that, you know robots can learn potentially but of course this is i think one of the easier tasks where if it's folding laundry like if we can do it you know not well then you know how much harder is it gonna be for a robot to do you know so i think there's still a lot of unknown questions to actually hit you know even a hundred of the tasks much less a thousand all the thousand so but i think we have seen some progress so i hope that we can you know start to scaffold up from there yeah so um what's what's the rest of i don't know the semester of the year uh like for you guys?
24:16Is it all heads down on this project? What's the timeline? Yeah, I think we have just had our first official release two days ago. And I think things are at a stage where we have all our 1 ,000 tasks ready for researchers and scientists to try it out. I think our immediate next step are to try some of these ourselves, you know, like what Google call like dog food, your own products, right? So we're, you know, robot learning researchers ourselves. We want to see how do the current state-of-the-art robot learning or robotics models work in these tasks? What are some of the pitfalls? So I think that's number one, that essentially tell us where are the low-hanging fruits that can really significantly improve our performances.
25:00And second is that we're also thinking about potentially hosting a challenge where researchers can... So everything is even more modular so that people can participate from all over the world. to make progress on this benchmark. And I think that's also in our roadmap to make it happen. Well, if you need a volunteer to create the wild party mess for your robots to clean up. We know who to ask. You know who to ask, yeah. I think along the lines of the challenge, I think a goal is to sort of democratize this sort of research and allow more people to explore. And so we've actually put together like a demo that anyone can try.
25:36So for the audience listening, like, you know, if you're technically inclined, if you're a researcher, but even if you're just like, I don't know, I don't want to say lay person, But, you know, a person that's normally not involved with AI, but want to sort of like just see what we're all about. We do actually have something where we can just try it out immediately and you can see sort of the robot in the simulation and like what it looks like. So I assume there'll be some links hopefully associated with this. But yeah, we hope you can try it out. Yeah, there'll be show notes. If you know the link offhand, you can speak it out now.
26:00But we'll do show notes as well. Yeah, it's on, what would it say? Behavior.stanford.edu. Yeah. Great. Okay. And that reminds me to ask, you mentioned it briefly, but what is behavior? Yeah, that's a good distinction. And so OmniGibson, to be clear, is like the simulation platform where we simulate this stuff. And I think overarching that, this whole project is called Behavior 1K, representing the thousand tasks that we hope robots can solve in the near future. Right. Yeah, that's the distinction, I guess, is that it's the all-encompassing thing, which is not just the simulator, but also the tasks.
26:26And also sort of the whole ethos of the whole project is called Behavior 1K. Okay. Yeah. All right. And before I let you go, we always like to wrap these conversations up with a little bit of a forward-looking, what do you think your work, you know, how do you think your work will affect the industry? What do you think the future of, et cetera, et cetera, is? Robots. So it's 2024. Let's say by 2030. Are robots in the physical world going to be, you know, to some extent, we've got, you know, Roombas. Right, right. You know, vacuum cleaner robots, that kind of thing. And certainly at a show like this, you see robots out there.
27:00But, you know, there is an NVIDIA robot I saw yesterday that's out in children's hospitals interacting with patients. And yeah, it's down on the, I'm pointing, nobody can see it on the radio, but I'm pointing out to the show floor. It's out there, you can see it. Where do you think society is going to be in, you know, five, six years from now as relates to the quantity and sort of level of interactions with robots in our lives? Or maybe it's more than five, six years, maybe it's 2050 or further down the line. I think it's hard to predict because if we look at the autonomous driving industry, let's say as a predecessor, I think even though it's been hyped up as like the next ubiquitous thing for what a decade now, like for quite a while, right?
27:40But we're still not quite, I mean, it's become much more commonplace, but we're still not at like, what is it? Level five autonomy, let's say. And so I imagine something similar will happen with, you know, humanoid robots or something that you see with everyday, you know, interactive household robots where I can imagine it will start seeing them in real life. But I don't think it'll be ubiquitous until, you know, decades. It's my take, but I don't know if you're more optimistic. Yeah, I think to keep it pessimistic is actually a good thing because I think in general, the reason why these things are too hard is because humans have very high standards.
28:15It's like the sub-driving cars. People are okay drivers. These are drivers. So you only have X and X number of miles. You want the robot to be better at it. Much better at it, actually. So I think we're still a bit far away from, you know, a house of robots that can be very versatile, meaning do many things and also do many things reliably. Very consistently. You don't want to break, you know, 20 % of the time. 20 % of your dishes. Exactly. So I think it's hard because we have high standards. But hopefully these robots can kind of come in like incrementally for our life. Maybe first in more structured environment, like warehouses and so on, like doing like reshelving or like restocking shelves or putting Amazon packages here and there.
28:58And then hopefully soon we can have a full-time job of folding laundry robots soon. Excellent. Good enough. Eric, Josiah, guys, thank you so much for dropping by, taking time out of your busy week to join the podcast. This is a lot of fun for me, I'm sure, for the audience. For listeners who want more, want to learn more about the work you're doing, more about what's going on at the lab, read some published research, that kind of thing, are there good starting points online where we can direct them? I think what Eric mentioned earlier, like just go to behavior.stanford.edu. And that's sort of the entry point where you can see, you know, all this stuff about this project, but more, you know, widely you can also then from there get to see what else has gone on at Stanford.
29:39That's exciting. So, yeah, definitely check it out if you're so inclined. Perfect. All right. Well, guys, thanks again. Enjoy the rest of your show and good luck with the research. Thanks, Noah. Thanks for having us.
30:16Thank you.
From the publisher
Imagine having a robot that could help you clean up after a party — or fold heaps of laundry. Chengshu Eric Li and Josiah David Wong, two Stanford University Ph.D. students advised by renowned American computer scientist Professor Fei-Fei Li, are making that a dream come true. In this episode of the AI Podcast, host Noah Kravitz spoke with the two about their project, BEHAVIOR-1K, which aims to enable robots to perform 1,000 household chores, including picking up fallen objects or cooking. To train the robots, they’re using the NVIDIA Omniverse platform, as well as reinforcement and imitation learning techniques. Listen to hear more about the breakthroughs and challenges Li and Wong experienced along the way.




