In short
NVIDIA AI Podcast Episode Notes
Episode Overview
- Title: How World Foundation Models Will Advance Physical AI With NVIDIA’s Ming-Yu Liu - Ep. 240
- Description: Discussion with Ming-Yu Liu, Vice President of Research at NVIDIA, focusing on world foundation models and their implications for various industries, particularly in enhancing AI workflows and development.
Key Guests
- Ming-Yu Liu
- Position: Vice President of Research at NVIDIA
- Notable Achievement: IEEE Fellow
Introduction
- Host Noah Kravitz introduces the topic of world foundation models and their significance in AI development.
- Reference to NVIDIA Cosmos, a development platform for these models, announced by CEO Jensen Huang at CES.
Understanding World Foundation Models
- Definition:
- Deep learning-based space-time visual simulators.
- Simulate physical environments, human intentions, and activities.
- Used to generate virtual worlds from various prompts (text, image, video).
- Customization:
- World models can be tailored to specific physical AI setups, accommodating different configurations (e.g., camera placements).
Comparison with Other AI Models
- Differences from LLMs:
- LLMs focus on text generation and understanding.
- World models concentrate on simulation (e.g., generating videos).
- Differences from Video Generation Models:
- World models generate future outcomes based on physical laws and intentions, while video models generate content for creative purposes.
Need for World Foundation Models
- Significance for Physical AI Developers:
- Essential for systems that operate in the real world and can cause physical damage.
- Applications:
- Training Verification:
- Allows for pre-deployment testing of AI models in simulated environments to avoid real-world damage.
- Policy Initialization:
- World models can serve as starting points for training policy models, requiring less data.
- Action Prediction:
- Enables pre-simulation of multiple scenarios to select the best action before deployment.
Measuring Accuracy of World Models
- Challenges:
- World model development is still evolving, and accurate measurement remains a work in progress.
- Key Aspects:
- Must adhere to physical laws (e.g., object permanence and accurate physical predictions).
Introduction to Cosmos
- Announcement at CES:
- The Cosmos platform provides pre-trained world foundation models (both diffusion and autoregressive) for developers.
- Components of Cosmos:
- Pre-trained models, tokenizers for video data, and post-training scripts for fine-tuning.
- Open-weight models available for commercial use, aimed at aiding developers in physical AI.
Diffusion vs. Autoregressive Models
- Autoregressive (AR) Models:
- Generate tokens one at a time based on previous observations (e.g., GPT models).
- Diffusion Models:
- Generate groups of tokens together and iteratively refine them, leading to potentially higher quality outputs.
Applications of World Foundation Models
- Industries Benefiting:
- Self-driving cars and humanoid robotics are highlighted as key sectors that will significantly benefit from these models.
Future Prospects
- Growth and Development:
- The field is still in its infancy, with rapid advancements expected as AI technology continues to evolve.
- Collaborative Efforts:
- Emphasis on partnerships with companies to address real-world problems and iterate on model development.
Conclusion
- Ming-Yu Liu shares insights on the transformative potential of world foundation models and encourages feedback on community resources.
- Mention of a white paper on the Cosmos platform for further reading.
Next Steps for Listeners
- Explore additional resources from NVIDIA regarding world foundation models and their applications.
- Stay engaged with ongoing developments in AI technology as discussed in the podcast.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. NVIDIA CEO Jensen Huang recently keynoted the CES Consumer Electronics Show conference in Las Vegas, Nevada. Amongst the many exciting announcements Jensen talked about was NVIDIA Cosmos. Cosmos is a development platform for world foundation models, which I think we're all going to be talking a lot about in the coming months and years. What is a world foundation model? Well, thankfully, we've got an expert here to tell us all about it. Mingyu Liu is Vice President of Research at NVIDIA. He's also an IEEE fellow, and he's here to tell us all about World Foundation models, how they work, what they mean, and why we should care about them going forward.
0:57So without further ado, Mingyu, thank you so much for joining the NVIDIA AI podcast, and welcome. It's great to be here. So let's start with the basics, if you would. What is a World Foundation model? Sure. So World Foundation models are deep learning-based space-time visual simulator that can help us look into the future. It can simulate physics. It can simulate people's intentions and activities. It's like data strength of AI. Imagine many different environments and can simulate the future. So we can make good decisions based on the simulation. We can leverage wall foundation models, imagination, and simulation capability to help train physical AI agents.
1:43We can also leverage this capability to help the agent make good decisions during the inference time. It can generate a virtual world based on text prompts, image prompts, video prompts, action prompts, and the layer combinations. So we call it a world foundation model because it can generate many different worlds and also because it can be customized to different physical AI setups to become a customized world model, right? So different physical AI have different number of cameras in different locations. So we want the world foundation model to be customizable for different physical AI setups so they can use in their settings.
2:20So I want to ask you kind of how a world model is similar or different to an LLM and other types of models. But I think first I want to back up a step and ask you, how is a world model similar or different to a model that generates video? Because my understanding, and please correct me when I'm wrong, my understanding is that you can prompt a world model to generate a video, but that video is generated based on the things you were talking about, based on understanding of physics and other things in the physical world, and it's a different process. So I don't know what the best way is to kind of unpack it for the listeners, but one place to start might be, how does a world model differentiate from an LLM or a generative AI video model?
3:07So world model is different to LN in the sense that LN is focused on generating text description. It generates understanding. And world model is generating simulation. And the most common form of simulation is videos. So they are generating pixels. And so world models and video foundation models, they are related. And video foundation model is a general model that generates videos. It can be for creative use cases. It can be for other use cases. In world models, we are focusing on this aspect of video generation. Based on your current observation and the intention of the actors in your world, you roll out the future.
3:52Yeah. So they are related, but with a different focus. Gotcha. Thank you. So why do we need the world models? I mean, I think I know part of the answer to the question we're talking about simulating physical AI and all of these amazing things. But what's the, you know, tell us about the need for world foundation models from your perspective. So I think World Foundation models is important to physical AI developers. You know, physical AI are systems with AI deployed in the real world, right? And different to digital AI, it's physical AI systems that interact with the environment can create damage, right?
4:29So this could be real harm, right? Right, right. So a physical AI system might be controlling a robotic arm or some other piece of equipment changing the physical world? Yeah, I think there are three major use cases for physical AI. Okay. It's all around simulation. The first one is, you know, when you train a physical AI system, you train a deep learning model, you have a thousand checkpoints. Do you know which one you want to deploy it, right? Right. And if you deploy individually, it's going to be very time consuming. And so then it's bad, it's going to damage your kitchen, right? So with a world model, you can do verification in the simulation.
5:09Right. So you can quickly test out this policy in many, many different kitchens. And before, you know, you deploy in the real kitchen. And after this verification step, you maybe narrow down to three checkpoints. And then you do the real deployments. So, you know, you can have an easier life to deploy your physical AI. It reminds me of when we've had podcasts about drug discovery. and the guests talking about the ability to simulate experiments and different molecular combinations and all of that work so that they can narrow it down to the ones that are worth trying in the actual, the physical lab, right?
5:49So it sounds like, you know, similar, like just being able to simulate everything and narrow it down must be such a huge advantage to developers. Yeah. And second application is, you know, role model, if you can predict the futures, you have some kind of understanding of basics. You might know the action required to drive the world toward that future. And the policy model, you know, the typical one deployed in physical AI is all about between the action, right action, given the observation. So world model can be used as initialization to the policy model. And then, you know, you can train the policy model with less amount of data because the world model is already pre-trained with many different observations that's on the data assets.
6:34So without a world model, what's the procedure of training a policy like? So one procedure is you collect data and then you start to do the supervised by tuning. Right. And then you may use, yeah. So it's hands-on, it's manual, you have to get all the data. It's a lot, yeah. Yeah. And third one is when world model is good enough highly accurate and fast. You know, before the robot taking any actions, you just simulate different features. And the check which you want to really achieve your goal and take that one. Yeah, it's like I have a data strength next to you before you're making any decision.
7:13Wouldn't it be great? You mentioned accuracy when the models are fast enough and accurate enough. And I don't know if it's a fair question to ask. So ask it, interpret it the best way. But like, how do you determine accuracy on a, or measure accuracy on a world model? And is there a benchmark that, you know, different benchmarks you need to hit to deploy in different situations? How does that work? Yeah, it's a great question. So I think a world model development is still in its infancy. Right. So people are still trying to figure out the right way to measure the world model performance. And I think there are several aspects a world model must have.
7:49One is follow the law of physics. When you're dropping a ball, you should predict it's in the right position based on the fit it looks. And also in the 3D environment, we have to have object permanence. So when you turn back and come back, the object should remain there without any other players. It should remain in the same location. So there are many different aspects I think we need to capture. And I think an important part for the research community is to come out with the right benchmark. so that the community can move forward in the right location to democratize this important area. Right.
8:26So speaking of moving forward, maybe we can talk a little bit, or you can talk a little bit about Cosmos and what was announced at CES. So in CES, Jensen announced the Cosmos World Model Development Platform. It's a developer-first world model platform. So in this platform, there are several components. One is pre-trained wall foundation models. We have two kinds of wall foundation models. One is based on diffusion, the other is based on autoregressive. And we also have tokenizers for the wall foundation models. Tokenizers compress videos into tokens so that transformers can consume for their task.
9:09In addition to these two, we also provide post-training scripts to help physical air builder to fine-tune the pre-trained model to their physical air setup. Some cars have eight cameras, right? And we rely on our World Foundation model to predict eight views. And lastly, we also have this video curation toolkit. So processing videos, a lot of video is already the computing task. There are many bits that need to be processed. and media gather libraries as they're ready to do computation code together, want to help the world model developers leverage the library to read data. Either they want to build their own world models or find one based on our pre-trained world foundation models.
10:00So the models provided as part of Cosmos, those are open to developers to use? They open to other businesses, enterprises? Yes. So this is an open-weight development platform. So meaning that the model is open-weight. The model weights are released for commercial use. We feel this is important to physical ad builders, right? So physical ad builders, they need to solve tons of problems to build really useful robots, self-driving cars for our society. There are so many problems and world model is one of them. And those companies, they may not have the resources or expertise to build a role model.
10:44NVIDIA care about our developers, and we know many of them are trying to make a huge impact in physical AI. So we want to help them. That's why we created this role model development platform for them to leverage, so that they can handle other problems and we can contribute our art to the transformation of our society. Absolutely. I wanted to ask you, can you explain a little bit about the difference between diffusion models and autoregressive models, particularly in this context? Why offer both? What are the use cases and pros and cons? So autoregressive model or AR model, it's a model that we did token once at a time, condition of what has been observed.
11:29So GPT is probably the most popular auto-request model we did token at the time. Deficient, on the other hand, is a model that we did a set of tokens together. Right. And so iteratively removed noises from these initial tokens. Right, right, right. Yeah. And the difference is that for AR model, there's a significant amount of investment in GPT. There are so many optimizations, so they can run very fast. Right, right, right. And deep fusion, because tokens are generated together, so it's easier to have coherent tokens. Right. The generation quality tend to be better. And both of them are useful for physical air builders.
12:14So some of them need speed, some of them need high accuracy. So both are good. Excellent. So far, the most successful autoregressive model is based on discrete token prediction. like in GPT. So you pretty much have a set of integers, tokens, and you put this then during training. And in the case of wall foundation models, it means you have to organize videos into a set of integers. And you can imagine it's a challenging compression task. Right. And because of this compression, the autoregressive model tend to struggle more on the accuracy, but it has other benefits. For example, it's setting it's easier integrated into the physical AI setup.
12:58Got it. I'm speaking with Mingyu Liu. Mingyu is vice president of research at NVIDIA, and he's been telling us about world foundation models, including the announcement of NVIDIA Cosmos, the developer platform for world models that was announced during Jensen's CES keynote. So we've been talking a lot about, you've been explaining what a world model is, how it's similar and different to other types of AI models, just now the difference between autoregression and diffusion. Let's kind of change gears a little bit and talk about the applications. How will Cosmos, how are our World Foundation models going to impact industries?
13:33Yeah, so we believe that, first of all, the World Foundation model can be used as a synthetic data generation engine to generate different synthetic data. And like what I said earlier, the work model can also be used as a policy evaluation tool to determine which checkpoint or which policy is a better candidate for you to test out in the physical world. Right. And also, if you can predict the future, you probably can reconfigure it to predict the action toward that future. So as a policy training initialization. Right, right. And also to have a data strength next to you before any endeavor. So during the next time, spiritual rollout and pick the best decision for each moment.
14:20Are there particular industries? I know work in factories and industrial work, anything involving robotics, but are there specific industries that you see benefiting from world models maybe sooner than others? Yes, I think the self-driving car industry and the human-knowing robot industry will benefits a lot from these role model developments. They can simulate different environments that will be difficult to have in the real world to make sure the agent is behave effectively. So I think these are two very exciting industries the role models can impact. And NVIDIA obviously has a long history, as you were saying, of, you know, it's not just about rolling out the hardware.
15:04There's the software, the stack, the ecosystem, all of the work to support developers. Because if the devs aren't building world-changing things with the products, then there's a problem, right? What are some of the partnerships, the ecosystems relative to World Foundation models? And maybe there's some partners who are already doing some interesting stuff with the tech you can talk about. Yes. We are working with a couple of humanoid companies and sales driving car companies, including OneX, Wabi, Leoto, S10, and many others. Right. So Amidia believes in suffering. We believe that two greatness come from suffering.
15:42So working with our partners, we can look at the challenges they are facing to experience their pain and to help us to build a role model platform that is really beneficial to them. Fantastic. Yeah. So I think this is the important part to make the field move faster. Absolutely. All right. So you talked about being able to predict the future and you talked about just now that things moving faster. What do you see on the horizon? What's next for World Foundation models? Where do you see this going in the next five years or adjust that time frame to whatever makes sense? So I'm trying to be a world model now, try to predict the future.
16:23Exactly. Yep, putting you on the spot. Yes. I believe we are still in the infancy of World Foundation model development. The model can do physics to some extent, but not well or robust enough. That's the critical point to make a huge transformation. It's useful, but we need to make it more useful. So the field of AI advances very fast. So from GPT-3 to ChetGPT, it's just a year or two. Right. Yeah, we forget. It's all going so quickly. Yeah, it's going so fast. I believe physical AI development will be very fast too because the infrastructure for large-scale model has been established. So this large-scale model transformation, right?
17:14And there's a strong need to have physical assistance for dry, so dry clean cars for humanoid. Right. And there are also a lot of investments. So we have the Great Foundation, and many young researchers want to make a difference. And we also have great need and investment. I think this is going to be a very exciting area and things are going to move very fast. I don't want to say that it will be solved in five years or ten years. So I think it's still a long way. And more importantly, we also need to study how to best integrate these role models into the physical AI systems in a way that can really benefit them.
17:56Right. And does that come through just working with partners out in the field, kind of combining research with application and iterating and learning? Yeah, I believe so. I believe in suffering. So I believe that to hang in hand with our partners, understand their problems is the best way to make progress. For folks who would like to learn more about any aspects of what we're talking about, there are obviously resources on the NVIDIA site. And of course, the coverage of Jensen's keynote and the announcements. Are there specific places, maybe a research blog, maybe your own blog or social media channels, where people can go to learn more about NVIDIA's work with world models and anything else you think the listeners might find interesting?
18:42Yes. So we have a white paper written for the Cosmos world model platform. and we welcome you to download and take a read and let me know how, you know, whether it's useful to you and let me know the feedback and we will try to do better for the next one. Excellent. Ming Yu, it was an absolute pleasure talking to you. I definitely learned more about world models and some of the particulars and the applications going forward. So I thank you for that. I'm sure the audience did as well. But, you know, the work that you're doing, as you said, it's early innings and it's all changing so fast. So we will all keep an eye on the research that you're doing in the applications and best of luck with it.
19:22And I look forward to catching up again and seeing how quickly things evolve from here on out. Thank you. Thanks for having me. It's been fun. And I hope next time I can share more, you know, maybe more advanced version of the role model. Absolutely. Well, thank you again for joining the podcast. Thank you.
19:49Thank you.
20:22The End
From the publisher
As AI continues to evolve rapidly, it is becoming more important to create models that can effectively simulate and predict outcomes in real-world environments. World foundation models are powerful neural networks that can simulate physical environments, enabling teams to enhance AI workflows and development. Ming-Yu Liu, vice president of research at NVIDIA and an IEEE Fellow, joined the NVIDIA AI Podcast to talk about world foundation models and how it will impact various industries. https://blogs.nvidia.com/blog/world-foundation-models-advance-physical-ai/ https://www.nvidia.com/cosmos/




