Jim Fan on Nvidia’s Embodied AI Lab and Jensen Huang’s Prediction that All Robots will be Autonomous

17 Sep 2024 · 49 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Jim Fan on Nvidia’s Embodied AI Lab

Podcast Overview Podcast Title: Training Data Hosts: Sonya Huang, Pat Grady, and other Sequoia Capital partners Episode Title: Jim Fan on Nvidia’s Embodied AI Lab and Jensen Huang’s Prediction that All Robots will be Autonomous Episode Description: Jim Fan, a prominent AI researcher, discusses his journey through AI, the role of Nvidia’s Embodied AI lab, and the implications of robotics becoming autonomous.

---

Key Contributors

  • Jim Fan: Senior Research Scientist at Nvidia, leads the Embodied AI GEAR group.
  • Fei-Fei Li: Jim's PhD advisor, influential in visual recognition and AI advancement.
  • Jensen Huang: Nvidia CEO, predicts all moving objects will become autonomous.

---

Episode Highlights

Introduction to Jim Fan's Journey

  • First intern at OpenAI, involved in foundational AI projects.
  • PhD at Stanford under Fei-Fei Li, focusing on embodied AI.
  • Transitioned to Nvidia, leading innovations in robotics and AI agents.

The GEAR Group at Nvidia

  • Mission: Develop AI agents that can act in both physical and virtual environments.
  • Focus on Project GR00T, aiming to create advanced humanoid robots.

Three-Pronged Data Strategy for Robotics

  1. Internet-Scale Data: Leveraging diverse online data for common sense learning.
  2. Simulation Data: Using Nvidia’s simulation tools to generate vast amounts of synthetic data.
  3. Real Robot Data: Collecting real-world data through teleoperation, though it's more limited and costly.

Future of Robotics

  • Foundation Agent: Concept of creating a single intelligent agent that generalizes across skills and environments.
  • Prediction that in the next decade, robots will integrate into society, performing household tasks and assisting in various sectors.

Distinction Between Virtual and Physical Worlds

  • Similarities in the design of agents for gaming and robotics (input perception, output actions).
  • Challenges in bridging the simulation-to-reality gap in robotics.

Generative Models and Learning

  • Jim discusses potential breakthroughs in autonomous robots that can understand and execute tasks based on abstract instructions, similar to how humans process commands.

Key Technologies Mentioned

  • Project GR00T: Nvidia’s foundation model for humanoid robotics.
  • Jetson Orin chip: Next-gen computing hardware for robotics.
  • Eureka Project: Trained a five-finger robot hand to perform tasks like pen spinning.
  • MineDojo: Platform for general-purpose agents in Minecraft, awarded at NeurIPS 2022.

Implications for AI and Society

  • The increasing affordability and capability of humanoid robots could lead to widespread implementation in daily life.
  • Ethical considerations and safety regulations will be crucial for the integration of robots in society.

Challenges and Considerations

  • The need for robust data strategies to overcome limitations in current robotic systems.
  • Balancing the development of generalist and specialist models in AI.

Closing Thoughts

  • Jim highlights the importance of long-term planning in AI research and emphasizes the potential for transformative advancements in robotics and AI in the coming decade.

---

Key Takeaways

  • The future of robotics is expected to yield increasingly autonomous capabilities.
  • A well-rounded data strategy combining various data sources is essential for advancements in AI and robotics.
  • Humanoid robots will likely play a central role in everyday tasks, bridging the gap between human needs and robotic capabilities.
  • Continuous research and breakthroughs in AI will shape the landscape of robotics, making it a pivotal area for future investment and innovation.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So from the chip level, which is the Justin Soar family, to the foundation model, project group, and also to the simulation and the utilities that we build along the way, it will become a platform, a computing platform for human robots, and then also for intelligent robots in general. So I want to quote Jensen here. One of my favorite quotes from him is that everything that moves will eventually be autonomous. and I believe in that as well. It's not right now but let's say 10 years or more from now if we believe that there will be as many intelligent robots as iPhones then we'd better start building that today.

0:57Hi and welcome to Training Data. We have with us today Jim Fan, Senior Research Scientist at NVIDIA. Jim leads NVIDIA's Embodied AI Agent Research with a dual mandate spanning robotics in the physical world and gameplay agents in the virtual world. Jim's group is responsible for Project Groot NVIDIA's humanoid robots that you may have seen on stage with Jensen at this year's GTC. We're excited to ask Jim about all things robotics, why now, why humanoids, and what's required to unlock a GPT -3 moment for robotics. Welcome to training data. Thank you for having me. We're so excited to dig in today and learn about everything you have to share with us around robotics and embodied AI.

1:39Before we get there, you have a fascinating personal story. Thank you for the first, the first intern at OpenAI. Maybe walk us through some of your personal story and how you got to where you are. Absolutely. I would love to share the stories with the audience. So back in the summer of 2016, some of my friends said there is a new startup in town, and you should check it out. And I'm like, huh, I don't have anything else to do because I got accepted to PhD. And that summer I was idle, so I decided to join the startup and that turned out to be OpenAI. And during my time at OpenAI, we were already talking about HGI, back in 2016.

2:19And back then, my intermentor was on Dr. Kapparthy and Ilya Soskever. And we talked about, and we discussed a project together, it's called World of Bits. So the idea is very simple. We want to build an AI agent that can read computer screens, read the pixels from the screens, and then control the keyboard and mouse. If you think about it, this interface is as general as it can get. Right, like all the things that we do on computer, like you know, Replying to emails or playing games or browsing the web. It can all be done in this interface, mapping pixels to keyboard mouse control. So that was actually my first kind of attempt at AGI, at OpenAI, and also my first journey, the first chapter of my journey in AI agents.

3:06I remember World of Bets actually. I didn't know that you're a part of that. That's really interesting. Yeah, yeah, it was a very fun project and was part of a bigger initiative called Open a Universe Which was like a bigger Platform on like integrating all the applications and games into into this framework What do you think were some of the unlocks then and then also what do you think were some of the challenges that you had with agents back then? Yes, so back then the main method that we use was reinforcement learning There was no LOM no transformer back in 2016 And the thing is, reinforcement learning, it works on specific tasks, but it doesn't generalize.

3:42Like we can't give the Asian arbitrary language and instructed to do arbitrary things that we can do with a keyboard mouse. So back then, it kind of worked on tasks that we designed, but it doesn't really generalize. So that started my next chapter, which is I went to Stanford, and I started my PhD with Professor Faith Ely, and we started working on computer vision and also embodied AI. And during my time at Stanford, which was from 2016 to 2021, I kind of witnessed the transition of the Stanford Vision Lab led by Feifei from, you know, static computer vision, like recognizing images and videos, to more embodied computer vision, when an agent learns perception and takes actions in an interactive environment.

4:31And this environment can be virtual as in simulation, or it can be the physical world. So that was my PhD, like transitioning to embodied AI. And then after I graduated from PhD, I joined Nvidia and have stayed there ever since. So I carried over my work from my PhD thesis to Nvidia and still working on embodied AI till this day. So you oversee the embodied AI initiative at Nvidia. Maybe say a word on what that means and what you all are hoping to accomplish. Yes, so the team that I am co -leading right now is called gear which is a GER and that stands for Journalist Embodied Agent Research. And to summarize what our team works on in three words is that we generate actions because we build in body AI agents and those agents take actions in different worlds.

5:21And if the actions are taken in the virtual world, that will be gaming AI and simulation. And if the actions are taken in the physical world, that will be robotics. Actually, earlier this year, in March, GTC, at Jensen's keynote, he unveiled something called Project Group, which is a video's moonshot effort at building foundation models for human -erobotics. And that's basically what the gear team is focusing on right now. We want to build the AI brain for human -erobots and even beyond. What do you think is Nvidia's competitive advantage in building that? Yeah, that's a great question. So one is for sure like compute resources.

6:02All of these foundation models require a lot of compute to scale up. And we do believe in scaling law. There were scaling laws for like LOMs. But the scaling law for embodied AI and robotics are yet to be studied. So we're working on that. And the second strength of Nvidia is actually simulation. So, Nvidia, before it was an AI company, it was a graphics company. So, Nvidia has many years of expertise on building simulation, like physics simulation and rendering and also real -time acceleration on GPUs. So, we are using simulation heavily in our approach to build robotics. The simulation strategy is super interesting.

6:41Why do you think most of the industry is still very focused on real -world data, the opposite strategy? Yeah, I think we need all kinds of data and simulation and real world data by themselves are not enough. So, at GIR we divide this data strategy into roughly three buckets. One is the Internet scale data, like all the tags and videos online. And the second is simulation data, where we use Nvidia simulation tools to generate lots of synthetic data. And the third is the real robot data, where we collect the data by tele -operating in the robot, and then just collecting and recording those data on the robot platforms.

7:22And I believe a successful robotic strategy will involve the effective use of all three kinds of data and mixing them and also delivering a unified solution. Can you say more about, well, you're talking earlier about, you know, have data as fundamentally the key bottleneck in making a robotics foundation model actually work? Can you say more about your kind of conviction in that idea? And then what exactly does it take to make great data to kind of break through this problem? Yes. So I think like the three different kinds of data that I just mentioned have different strengths and weaknesses. So for internet data, they are the most diverse.

8:00They encode a lot of common sense prior. Like for example, most of the videos online are human centered. because humans, we love to take selfies, we love to record each other doing all kinds of activities, and there are also a lot of instructional videos online. So we can use that to kind of learn how humans interact with objects and how objects behave under different situations. So that kind of provides a common sense prior for the robot foundation model. But the Internet skill data, they don't come with actions. We cannot download the motor control signals of the robots from the internet. And that goes to the second part of the data strategy, which is using simulation.

8:43So in simulation, you can have all the actions and you can also observe the consequences of the actions in that particular environment. And the strength of simulation is that it's basically infinite data. And the data scales with compute. the more GPUs you put into the simulation pipeline, the more data that you will get. And also the data is super real time. So if you collect data only on the real robot, then you are limited by 24 hours per day. But in simulation, like the GPU accelerated simulators, we can actually accelerate real time by 10 ,000x. So we can collect the data at much higher throughput, given the same work of time.

9:24So that's a strength. But the weakness is that for simulation, no matter how good the graphics pipeline is, there will always be this simulation to reality gap. The physics will be different from the real world, and the visuals will still be different. They will not look exactly as realistic as real world, and also there is a diversity issue. The contents in the simulation will not be as diverse as all the scenarios that we encounter in the real world. So these are the weaknesses. And then going to the real robot data. And those data, they don't have the same to real gap, because they're collected on the real robot.

10:05But it's much more expensive to collect, because you need to hire people to operate the robots. And again, they're limited by the speed of the world of atoms. You only have 24 hours per day, and you need humans to collect those data, which is also very expensive. So we see these three types of data as having complementary strengths. And I think a successful strategy is to combine their strengths and then, you know, to remove their weaknesses. So the cute groups robots that were on stage with Jensen, those are such a cool moment. If you had to help us dream in one five, ten years, like what do you think your group will have accomplished?

10:45Yeah, so this is pure speculation, but I hope that we can see research breakthrough in robot foundation model maybe in the next two to three years So that's what we call like a GPT -3 moment for robotics and then After that, it's a bit uncertain because to have the robots Enter like daily lives of people. There are a lot more things than just a technical side The robots need to be affordable and mass -produced and we also need like safety for the hardware and also privacy regulations and those will take longer for the robots to be able to hit a mass market. So that's a bit harder to predict, but I do hope that the research breaks through a coming in the next two to three years.

11:29What do you think will define what a GPT -3 moment in AI robotics looks like? Yeah, that's a great question. So I would like to think about robotics as consisting of two systems, system one and system two. So that comes from the book thinking fast and slow where a system one means this low -level motor control that's unconscious and fast like for example when I'm grasping this cup of water I don't really think about how I move the fingertip at every millisecond So that would be system one and then system two is slow and deliberate and it's more like reasoning and planning That actually uses the the conscious brain power that we have So, I think the GPU 3 moment will be on the system one side.

12:17And my favorite example is the verb open. So just think about the complexity of the word open, right? Opening the door is different from opening window. It's also different from opening a bottle or opening a phone. But for humans, we have no trouble understanding that open means different things when you're interacting. It means different motions when you're interacting with different objects. but so far we have not seen a robotics model that can generalize on a low -level motor control on these verbs. So I hope to see a model that can understand these verbs in their abstract sense and can generalize to all kinds of scenarios that make sense to humans.

12:56And we haven't seen that yet, but I'm hopeful that this moment could come in the next two to three years. What about system two thinking? Like how do you think we get there? Do you think that some of, you know, the reasoning efforts in the LLM world will be relevant as well in the robotics world? Yeah, absolutely. I think for System 2, we have already seen very strong models that can do reasoning and planning and also coding as well. So these are the LOMs and frontier models that we have already seen these days. But to integrate the System 2 models with System 1 is another research challenge in itself.

13:33So the question is, for Robo Foundation model, do we have a single monolithic model or do we have some kind of cascaded approach where the system two and the system one models are separate and can communicate with each other in some ways? I think that's an open question and again they have pros and cons, right? Like for the first idea, the monolithic model is cleaner, there's just one model, one API to maintain, but also it's a bit harder to kind of control because you have different control frequencies. Like the system two models will operate on slower control frequency, let's say one hertz, like one decision per second, while the system one, like the motor control of me grasping this cup of water, that will likely be a thousand hertz, where I need to make these minor, like these tiny muscle decisions at a thousand times per second.

14:25It's really hard to encode them both in a single model, so maybe a cascaded approach will be better. But again, how do we communicate between system one and two? Do they communicate through text or through some latent variables? It's unclear, and I think it's a very exciting new research direction. Is your instinct that will get there in that breakthrough on system one thinking? Like through scale and transformers, like, is it going to work? Or is that, you know, cross your fingers and hope and see? I certainly hope that the data strategy I described will kind of get us there. because I feel that we have not pushed the limit of transformers yet.

15:03On the essential level, like transformers take tokens in and outputs tokens. And ultimately, the quality of the tokens determines the quality of the model, the quality of those large transformers. And for robotics, as I mentioned, like the data strategy is very complex. We have all the internet data and also we need simulation data and the real robot data. And once we're able to scale up on the data pipeline with all those high quality actions, then we can tokenize them and we can send them to a transformer to compress. So I feel we have not pushed transform to limit yet. And once we figure out the data strategy, we may be able to see some emergent property as we scale up the data and scale up the model size.

15:47And for that, I'm calling it the scaling law for embodied AI and it's just getting started. I'm very optimistic that we will get there. I'm curious to hear what are you most excited about personally? When we do get there, what's the industry or application or use case that you're really excited to see this completely transformed the world of robotics today? Yes. So there are actually a few reasons that we chose humanoid robots as kind of the main research thesis to tackle. One reason is the world is built around human embodiment, the human form factor, all restaurants, factories, hospitals, and all equipment and tools, they're designed for the human form and also the human hands.

16:28So, in principle, a sufficiently good humanoid hardware should be able to support any task that a reasonable human can do in principle. And the humanoid hardware is not there yet today, but I feel in the next two to three years, the humanoid hardware ecosystem will mature and will have affordable humanoid hardware to work on. And then it will be a problem about AI brain, about how we kind of drive those humanoid hardware. And once we have that, once we're able to have the group foundation model that can take any instruction in language and then perform any tasks that a reasonable human can do, then we unlock a lot of economic value.

17:08Like we can have robots in our households helping us with daily chores, like laundry, dishwasher and cooking or like elderly care, and we also have them in restaurants, in hospitals, in factories, helping with all the tasks that humans do. And I hope that will come in the next decade, but again, as I mentioned in the beginning, This is not just a technical problem, but also there are many things beyond the technology. So I'm looking forward to that. Any other reasons you've chosen to go after humanoid robots specifically? Yeah, so there are also a bit more practical reasons in terms of the training pipeline.

17:49So there are lots of data online about humans, right? It's all like human centered, all the videos, all like humans doing daily tasks or having fun. And the humanoid robot form factor is closest to the human form factor, which means that the model that we train using all of those data will be able to have an easier time to transfer to the humanoid form factor other than other form factors. So let's say for robot arms, right? Like how many videos do we see online about robot arms and grippers? Very few, but there are many videos of people using their five finger hands to work with objects. So it might be easier to train for human robots.

18:33And then once we have that, we'll be able to specialize them to like the robot arms and more kind of specific robot forms. So that's why we're aiming for the full generality first. I didn't realize it. So are you exclusively trading on humanoids today versus robot arms and robot dogs as well? Yeah. So for project group. In simulation, yeah. Yes, so for project group we are aiming more towards humanoid right now But the pipeline that we're building including the simulation tools Right the real robot tools like those are general purpose enough that we can also adapt to other platforms in the future So yeah, we're building these tools to be generally applicable You've used term general quite a few times now I think there are some folks from especially from the robotics world who think that you know a general approach won't work can you have to be domain environment specific?

19:26Why have you chosen to go after a generalist approach? And the Richard Sutton bit or less and stuff has been a recurring theme on our podcast. I'm curious if you think it holds in robotics as well. Absolutely. So I would like to first talk about the success story in NLP that we have all seen. So before the chat GPT and the GBD3 in the world of NLP, there are a lot of different models and pipelines for different applications, like translation and coding and doing math and doing creative writing. They all use very different models and completely different training pipelines. But then, Chargibit came and unified everything into a single model.

20:13So before Chargibit will call those specialists. And then the GB3 and Chargibit is, we call them the journalists. And once we have the general list, we can prompt them, distill them, and fine tune them back to the specialized tasks. And we call those the specialized general list. And according to the historical trend, it's almost always the case that the specialized general list are just far stronger than the original specialist. And they're also much easier to maintain because you have a single API that takes text in and then spits text out. So I think we can follow the same success story from the world of NLP and it will be the same for robotics.

20:54So right now in 2024 most of the robotics and applications we have seen are still in the specialist stage. They have a specific robot hardware for specific tasks collecting specific data using specific pipelines. But Project Group aims to build this general purpose foundation model that works on humanoid first, but later will generalize to all kinds of different robot forms or embodiments. And that will be the journalist moment that we are pursuing. And then once we have that journalist, we'll be able to prompt it, fine tune it, distill it down to specific robotics tasks. And those are the specialized journalists.

21:36But that will only happen after we have the journalist. So it would be easier in a short run to pursue the specialist. It's just easier to show results because you can just focus on a very narrow set of tasks. But we at Nvidia believe that the future belongs to journalists. Even though it will take longer to develop, it will have more difficult research problems to solve. But that's what we're aiming for first. The usually thing about Nvidia building grew to me is also what you mentioned earlier, which is that Nvidia owns both the chip and the model itself. What do you think are some of the interesting things that Nvidia could do to optimize Groot on its own chip?

22:15Yes, so at the March GTC Jensen also unveiled the next generation of the edge computing chips. It's called the Jensen store chip And it was actually co -announced with project Groot So the idea is we will have kind of the full stack as a unified solution to the customers So from the chip level, which is the Justin Soar family, to the foundation model, Project Groot, and also to the simulation and the utilities that we build along the way, it will become a platform, a computing platform for human robots, and then also for intelligent robots in general. So I want to quote Jensen here. One of my favorite quotes from him is that everything that moves will eventually be autonomous.

23:02And I believe in that as well. It's not right now, but let's say 10 years or more from now. If we believe that there will be as many intelligent robots as iPhones, then we'd better start building that today. That's awesome. Are there any particular results from your research so far that you want to highlight anything that gives you kind of optimism or conviction and the approach that you're taking? Yes, we can talk about some prior works that we have done. So one work that I was really happy about was called URACA. And for this work, we did a demo where we trained a five -finger robot hand to do pen spinning.

23:42So very useful. And it's super humorous with respect to myself because I have given up pen spinning, in launches chart. I'm not able to do it. I'm doing live demo. I will film, miserably, at this live demo. So yeah, I'm not able to do this, but the robot hand is able to. And the idea that we use to train this is that we prompt an ROM to write code in the simulator API that Nvidia has built. So it's called the i6 -SIM API. And the ROM outputs the code for reward function. So a reward function is basically a specification of the desirable behavior that we want a robot to do. So the robot will be rewarded if it's on the right track or penalized if it's doing something wrong.

24:33So that's a reward function. And typically the reward function is engineered by a human expert, a typical robot assist, who really knows about the API. It takes a lot of specialized knowledge. And the reward function engineering is by itself a very tedious and manual task. So what URICA did was we designed this algorithm that uses OOM to automate this reward function design so that the reward function can instruct the robot to do very complex things like pen spinning. So it is a general purpose technique that we developed and we do plan to scale this up to beyond just pen spinning. It should be able to design reward functions for all kinds of tasks, or it can even generate new tasks using the Nvidia simulation API.

25:20That gives us a lot of space to grow. Why do you think? I remember five years ago, there were people that were research labs working on solving Rubik's cubes, the robot hand and things like that. It felt like robotics went through maybe a trough of disillusionment. And in the last year or so, it feels like the space has really heated up again. Do you think there is a why now around robotics this time around? And what's different and where, you know, we're reading that OpenAI is getting back into robotics. Everybody is now spinning up their efforts. Like, what do you think is different now? Yeah, I think there are quite a few key factors that are different now.

25:59One is on the robot hardware. Actually, since the end of last year, we have seen a surge of new robot hardware in the ecosystem. There are companies like Tesla, working on Optimus, Boston Dynamics, and so on, and a lot of startups as well. So we are seeing better and better hardware. So that's number one. And those hardware are becoming more and more capable with, you know, better dexterous hands, better whole body reliability. And the second factor is the pricing. So we also see a significant drop in the price and the cost, the manufacturing cost for the human robots. So back in 2001, NASA had a humanoid developed and it's called a robot knot.

26:42I remember if I recall correctly, it cost north of $1 .5 million per robot. And then most recently, there are companies that are able to put a price tag of about $30 ,000 full -fledged humanoid, and that's roughly comparable to the price of a car. And also, there's always this trend in manufacturing where a mature product, the price of it, will tend towards the price of the raw material cost. And for the humanoid, it typically takes only 4 % of the raw material of a car. So it's possible that we can see the cost trending downwards even more, and there could be an exponential decrease in the price in the next couple of years.

27:24And that makes these still very hard to be more and more affordable. That's the second factor of why I think humanoid is gathering momentum. And the third one is on the foundation model side. We are able to see the system too problem. The reasoning, the planning part being addressed very well by the frontier models, like the GPs and the clouds and the llamas of the world. And these LOMs, they are able to generalize to new scenarios. They're able to write code. and actually the URICA project I just mentioned leverages these coding abilities of the LOMs to help develop new robot solutions. And there are also a search in multi -modal models improving the computer vision, the perception of it.

28:06So I think these successes also encourage us to pursue robot foundation models because we think we can write on the generalizability of these frontier models and then add actions on top of them. So we can generate action tokens that will ultimately drive these humanoid robots. I completely agree with all that. I also think so much of what we've been trying to tackle to date in the field has been how to unlock the scale of data that you need to build this model. And all the research advancements that we've made, many of which you've contributed to yourself around Centauril and other things. And the tools that Nvidia's built with Isaac Sim and others have really accelerated the field alongside tell operation and cheaper tell operation devices and things like that.

Read the full transcript

28:49And so I think it's a really, really exciting time to be building here. Yeah, I agree. Yeah. I'd love to transition to talking about virtual worlds, so that's okay with you. Yeah, absolutely. Yeah. So I think you started your research more in the virtual world arena. Maybe say, word on what got you interested in Minecraft and versus robotics? Like, is it all kind of related in your world? What got you interested in virtual worlds? Yeah, that's a great question. So for me, my personal mission is to solve embodied AI. And for AI Asians embodied in the virtual world, that will be things like gaming and simulation.

29:26And that's why I also have a very soft spot for gaming. I also enjoy gaming myself. What did you play? Yeah, so I played Minecraft. At least I tried to. I'm not a very good gamer. And that's why I also won my AI to adventure my poor skills. Yeah. So I worked on a few gaming projects before. The first one was called Mindogel, where we developed a platform to develop general purpose agents in the game of Minecraft. And for those audience who are not familiar, Minecraft is this 3D voxel world where you can do whatever you want. You can craft all kinds of recipes, different tools, and you can also go on adventures.

30:07It's an open -ended game with no particular score to maximize and no fixed storylines to follow. So we collected a lot of data from the internet. There are videos of people playing Minecraft. There are also wiki pages that explain every concept and every mechanism in the game. Those are like multimodal documents and also like forums like Reddit. The Minecraft subreddit has a lot of people talking about the game in natural language. So we collected these multimodal datasets and we're able to train models to play Minecraft. So that was the first work of mine, Dojo. And later, the second work was called Voyager.

30:47So we had the idea of Voyager after GP4 came along, because at that time, it was the best coding model out there. So we thought about, hey, what if we use coding as action? And building on that insight, we're able to develop the Voyager agent where it writes code to interact with the Minecraft world. So we use an API to first convert the 3D Minecraft world into a text representation and then have the Asian write code using the action APIs. But just like human developers, the Asian is not always able to write code correctly on the first try. So we kind of give it a self -reflection loop where it tries out something and if it runs in an error or if it makes some mistakes in the Minecraft world it gets the feedback and it can correct its program.

31:38And once it's written the correct program, that's what we call scale, we'll have it saved to a scale library so that in the future, if the agent faces a similar situation, it doesn't have to go through that trial and error loop again. It can retrieve the scale from the scale library. So you can think of that scale library as a codebase that the LOM interactively authored all by itself, right? There's no human intervention. The whole codebase is developed by Voyager. So that's a second mechanism, the skill library. And the third one is what we call an automated curriculum. So basically the agent knows what it knows and then knows what it doesn't know.

32:16So it's able to propose the next task that needs to difficult, nor too easy for it to solve. And then it's able to just follow that path and discover all kinds of different skills, different tools and also travel along in a vast world of Minecraft. And because they travel so much, and that's why we call it the Voyager. So yeah, that was kind of our team's one of our earliest attempts building AI agents in the body of the world using foundation models. Talk about the curriculum thing more. I think that's really interesting because it feels like it's one of the more unsolved problems in kind of the reasoning in LLM world generally.

32:55Like, how do you make these models self -aware so that they know kind of how to take that next step to improve? Maybe say a little bit more about what you built on the curriculum and the reasoning side. Absolutely. I think there's a very interesting emerging property from those frontier models is that they can reflect on their own actions and they kind of know what they know and what they don't know and they're able to propose tasks accordingly. So for the ultimate curriculum invoitor, we gave the agent a high -level directive, that is to find as many novel items as possible. And that's just the one kind of sentence of goal that we gave.

33:34And we didn't give any instruction on which objects to discover first, which tools to unlock first. We didn't specify. And agent was able to discover that all by itself using this kind of coding and prompting and skill library. So it's kind of amazing that the whole system just works. I would say it's an emerging property once you have a very strong reasoning engine that can generalize. What do you think there are so many, so much of the kind of virtual world research has been done in the virtual world? And I'm sure it's not entirely because a lot of deep learning researchers like playing video games, although I'm sure it doesn't hurt either.

34:12But what I guess, what are the connections between solving stuff in the virtual world and the physical world and how do the two interplay. Yeah, so as different as gaming and robotics seem to be, I just see a lot of similar principles shared across these two domains. For the embodied agents, they take as input the perception, which can be a video stream along with some sensory input, and then the output actions. And in the case of gaming, it will be like keyboard mouse actions. And for robotics, it would be low level motor controls. So ultimately the API looks like this. And these agents, they need to explore in the world.

34:52They have to collect their own data in some ways. So that's what we call reinforcement learning and also self exploration. And that part, that principle is again shared among the physical agents and the virtual agents. But the difference is robotics is harder because you also have a simulation to reality gap to bridge. because in simulation, the physics and the rendering will never be perfect. So it's really hard to kind of transfer what you're learning simulation to the real world. And that is by itself an open -ended research problem. So for robotics, it's got a sim2 real issue, but for gaming it doesn't.

35:30You are training and testing in the same environment. So I would say that would be the difference between them. And last year, I proposed a concept called Foundation Agent, where I believe ultimately will have one model that can work on both, you know, virtual agents and also physical agents. So for the foundation agent, there are three axes over which it will generalize. Number one is the skills that it can do. Number two is the embodiments or like the body form, the form factor it can control. And number three is the world, the realities it can master. So in the future, I think a single model will be able to do a lot of different skills on a lot of different robot forms or agent forms and then generalize across many different worlds, virtual or real.

36:23And that's the ultimate vision that the Guirateam wants to pursue, the foundation agent. Pulling down the thread of virtual worlds and gaming in particular, and what you've unlocked already with some reasoning, some emergent behavior, especially working in an open -ended environment. What are some of your own personal dreams for what is now possible in the world of games? Where would you like to see AI agents innovate in the world of games today? Yes, so I'm very excited by two aspects. One is intelligent agents inside the games. So the NPCs that we have these days, they have fixed scripts to follow and they're all manually altered.

37:03What if we have NPCs, the non -player characters that are actually alive? And you can interact with them. They can remember what you told them before, and they can also take actions in the gaming world that will change the narrative and change the story for you. So this is something that we haven't seen yet, but I feel there's a huge potential there so that when everyone played a game, everybody will have a different experience. And even for one person, you play the game twice. You don't have the same story. So each game will have infinite replay value. So that's one aspect. And the second aspect is that the game itself can be generated.

37:44And we already see many different tools kind of doing subsets of this grand vision, I just mentioned. There are text to 3D generating assets. There are also text to video models. And of course, there are language agents that can generate storylines. What if we put all of them together so that the game world is generated on the fly as you're playing and interacting with it? That would be just truly amazing and a truly open -ended experience. So we're interesting. For the agent vision in particular, do you think you need GPG -4 level capabilities or do you think you can get there with Lama -8B, for example, alone?

38:22Yeah, I think the agent needs the following capabilities. One is of course, it needs to hold and interesting conversation. It needs to have a consistent personality and it needs to have long -term memory and also take actions in the world. So for these aspects, I think currently like the Lava models are pretty good for that, but also not good enough to produce very diverse behaviors and really engaging behaviors. So I do think there's still a gap to reach. And the other thing is about inference cost. So if we want to deploy these agents to the gamers, then either it's like very low cost, holds it on the cloud, or it runs locally on the device.

39:06Otherwise, it's kind of unscalable in terms of cost. So that's another factor to be optimized. Do you think all this work in the virtual world space? Is it in service of, you know, you know, you're learning things from it that way, you can accomplish things in the physical world. Is it, does the virtual world stuff exist in service of the physical world ambitions? Or I guess, so differently, like, is it enough of a prize in its own rights? And how do you think about prioritizing your work between the physical and virtual worlds? Yes. So I just think the virtual world and the physical world ultimately will just be different realities on a single axis.

39:45So let me give one example. So there is a technique called domain randomization. And how it works is that you train a robot in simulation, but you train it in 10 ,000 different simulations in parallel. And for each simulation, they have slightly different physical parameters. Like the gravity is different, the friction, the weight, everything is a bit different. So it's actually 10 ,000 different worlds. And let's assume if we have an agent that can master all the 10 ,000 different configurations of reality all at once, then our real physical world is just a 10 ,000 first virtual simulation. And in this way, we're able to generalize from sim to real directly.

40:33So that's actually exactly what we did in our follow -up work to your ACA, where we're able to train agents using all kinds of different randomizations in the simulation and then transfer zero shout to the real world without further finding tuning. So I do believe that's Dr. Urika work and I do believe that if we have all kinds of different virtual worlds including from games and if we have a single agent that can master all kinds of skills in all the worlds then the real world just becomes part of this bigger distribution. Do you want to share a little bit about Dr. Rico to ground the audience in that example?

41:12Oh, yeah, absolutely. So for the Dr. Uraka work, we built upon Uraka and still use LOMs as kind of a robot developer. So the LOM is writing code. And the code is to specify the simulation parameters, like the domain randomization parameters. And after a while, after a few iterations, the policy that we train in a simulation will be able to generalize to the real world. So one specific demo that we showed is that we can have a robot dog walk on a yoga ball. And it's able to stay balanced and also even walk forward. So one very funny comment that I saw was someone actually asked his real dog to do this task and his dog isn't able to do it.

41:56So in some sense, our neural network is super dog performance. I'm pretty sure my dog would not be able to do it. I mean, all the ADI. Yeah. Yeah, artificial dog intelligence. Yeah, that's the next benchmark. In the virtual world's sphere, I think there's been a lot of just incredible models that have come out on both a 3D and the video side recently, all of them kind of transformer based. Do you think we're kind of there in terms of like, okay, this is the architecture that's going to take us to the promised land? And let's scale it up. Or do you think there's kind of fundamental breakthroughs that are still required on the model side there?

42:33Yes, I think for robot foundation models, like we haven't pushed the limit of the architecture yet. So the data is a harder problem right now and it's the bottleneck because as I mentioned earlier we can't download those action data from the internet. They don't come with those motor control data We have to collect it either in simulation or in the real on the real robots And once we have that we have a very mature data pipeline that will just push the tokens to the transformers and have it compressed those tokens just like, you know, transformers predicting the next word on Wikipedia. And we're still testing these hypotheses, but I don't think we have pushed the transformers to their limit yet.

43:18There are also a lot of research going on right now on alternative architectures to transformers. I'm super interested in those personally, like there are a Mamba, Recently, there was like test time training. There are a few alternatives. And some of them have very promising ideas. They haven't scaled really to the all the frontier model performance. But I'm looking forward to seeing alternatives to transformers. Have anyone caught your eye in particular and why? Yeah, I think I mentioned the Mamba work and also test time training. Like these models are more efficient, I think, time. So instead of like transformers attending to all the past tokens, these models have inherently more efficient mechanisms.

44:03So I see them holding a lot of promise, but we need to scale them up to the size of the frontier models and really see like how they compare heads on with the transformer. Awesome. Should we close that with some rapid fire questions? Yeah. Oh yeah. Okay, let's see. Number one, what outside the embodied AI world are you most interested in with an AI? Yeah, so I'm super excited about video generation because I see video generation as a kind of world simulator. So we learned physics and rendering from data alone. So we have seen like open air Sora and later there are like a lot of new models catching up to Sora.

44:45So this is like an ongoing research topic. and yeah. What does the world simulator get you? I think it's gonna get us a data driven simulation in which we can train in body AI. That would be amazing. Nice. What are you most excited about an AI on a longer term horizon? 10 years or more? Yeah, so on a few fronts. Like one is for the reasoning side. I'm super excited about models that code. I think coding is such a fundamental reasoning task that also has huge economic value. I think maybe 10 years from now we'll have coding agents that are as good as human level software engineers and then we'll be able to accelerate a lot of development using the LOMs themselves.

45:32And the second aspect is of course robotics. I think 10 years from now we'll have human robots that are at the reliability and agility of humans or even beyond. And I hope at that time, Project Group for the SS and that we're able to have humanoid helping us in our daily lives. I just want robots to do my laundry. Yeah, always in my dream. What year robots can endure a laundry? As soon as possible. I can't wait. Who do you admire most in the field of AI? And you've had the opportunity to work with some great dating back to your internship days. Who do you admire most these days? I have too many heroes in AI to count.

46:13So I admire my PhD advisor, Fei Fei. I think she taught me how to develop good research taste. So sometimes it's not about how to solve a problem, but identify what problems are worth solving. And actually the what problem is much harder than a how problem. And during my PhD years with Fei Fei, I transitioned to embodied AI. And in retrospect, this is the right direction to work on. I believe the future of AI agents will be embodied for robotics or for the virtual world. I was on my Andre Kaparasi. He's the great educator. I think he writes code like poetry. So I look up to him and then I admire Jensen a lot.

46:58I think Jensen, he cares a lot about AI research and he also knows a lot about even the technical details of the models and I'm super impressed. So I look up to him a lot. Pulling on the thread of having great research taste, what advice do you have for founders building an AI in terms of finding the right problems to solve? Yeah, I think we've reached some research papers. I feel that the research papers these days are becoming more and more accessible. And they have some really good ideas and they're more and more practical instead of just like theoretical mission learning. So I would recommend kind of keeping up with the latest literature and also just try out like all the open source tools That people have built so for example at Nvidia we built simulator Tools that everyone can have access to and just download it and try that out and you can try your own robots Yeah, in the simulations just get your hands dirty and maybe pulling on a thread of Jensen as an icon on what do you think is some practical, tactical advice you'd give the founders building an AI, what they could learn from him?

48:03Yeah, I think identifying the right problem to work on. So Nvidia bets on humanoid robotics because we believe like this is the future and also like embodied AI because if we believe that let's say 10 years from now, there will be as many intelligent robots in the world as iPhones, then we better start working on that today. Yeah. So yeah, just like long term future visions. I think that's a great note to end on. Jim, thank you so much for joining us. We love learning about everything your group is doing and we can't wait for the future of laundry folding robots. Awesome. Yeah. Thank you so much for having me.

48:40Yeah. Thank you. Thank you. Thanks.

From the publisher

AI researcher Jim Fan has had a charmed career. He was OpenAI’s first intern before he did his PhD at Stanford with “godmother of AI,” Fei-Fei Li. He graduated into a research scientist position at Nvidia and now leads its Embodied AI “GEAR” group. The lab’s current work spans foundation models for humanoid robots to agents for virtual worlds.
Jim describes a three-pronged data strategy for robotics, combining internet-scale data, simulation data and real world robot data. He believes that in the next few years it will be possible to create a “foundation agent” that can generalize across skills, embodiments and realities—both physical and virtual. He also supports Jensen Huang’s idea that “Everything that moves will eventually be autonomous.”
Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital
Mentioned in this episode:

World of Bits: Early OpenAI project Jim worked on as an intern with Andrej Karpathy. Part of a bigger initiative called Universe

Fei-Fei Li: Jim’s PhD advisor at Stanford who founded the ImageNet project in 2010 that revolutionized the field of visual recognition, led the Stanford Vision Lab and just launched her own AI startup, World Labs

Project GR00T: Nvidia’s “moonshot effort” at a robotic foundation model, premiered at this year’s GTC

Thinking Fast and Slow: Influential book by Daniel Kahneman that popularized some of his teaching from behavioral economics

Jetson Orin chip: The dedicated series of edge computing chips Nvidia is developing to power Project GR00T

Eureka: Project by Jim’s team that trained a five finger robot hand to do pen spinning

MineDojo: A project Jim did when he first got to Nvidia that developed a platform for general purpose agents in the game of Minecraft. Won NeurIPS 2022 Outstanding Paper Award

ADI: artificial dog intelligence

Mamba: Selective State Space Models, an alternative architecture to Transformers that Jim is interested in (original paper here)

00:00 Introduction
01:35 Jim’s journey to embodied intelligence
04:53 The GEAR Group
07:32 Three kinds of data for robotics
10:32 A GPT-3 moment for robotics
16:05 Choosing the humanoid robot form factor
19:37 Specialized generalists
21:59 GR00T gets its own chip
23:35 Eureka and Issac Sim
25:23 Why now for robotics?
28:53 Exploring virtual worlds
36:28 Implications for games
39:13 Is the virtual world in service of the physical world?
42:10 Alternative architectures to Transformers
44:15 Lightning round

More from Training Data

All 110 episodes
Jim Fan on Nvidia’s Embodied AI Lab and Jensen Huang’s Prediction that All Robots will be AutonomousTraining Data · 49 min
Listen in VO