NVIDIA’s Jim Fan Delves Into Large Language Models and Their Industry Impact - Ep. 204

3 Oct 2023 · 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NVIDIA AI Podcast Episode 204 Summary

Episode Title NVIDIA’s Jim Fan Delves Into Large Language Models and Their Industry Impact

Episode Description In this episode, host Noah Kravitz interviews NVIDIA Senior AI Scientist Jim Fan about his research utilizing large language models (LLMs) to create AI agents capable of autonomous play in Minecraft, resulting in the AI bot called Voyager. The discussion centers around the concepts of open-ended AI agents, their training methodologies, and their implications for various industries.

---

Key Concepts & Discussions

Introduction to AI Agents

  • Definition: Jim Fan defines AI agents as models that can proactively take actions, perceive outcomes, and improve their capabilities autonomously.
  • Difference from Traditional Models: Unlike typical language models that respond only when prompted, AI agents possess decision-making abilities independent of direct human instructions.

The Role of Minecraft in AI Research

  • Minecraft as a Testing Ground: Fan describes Minecraft as the "perfect primordial soup" for developing open-ended AI agents due to its lack of strict objectives, allowing for exploration and creativity.
  • MineDojo Project: This project was initiated to experiment with creating AI agents in Minecraft, leading to the development of Voyager.

The Voyager AI Bot

  • Leveraging GPT-4: Voyager employs Chat GPT-4 to autonomously write JavaScript code for game actions, allowing it to adapt and learn from its environment.
  • Self-Reflection and Lifelong Learning: Voyager can debug its code and store successful strategies in a skill library, demonstrating lifelong learning.
  • Autonomous Exploration: The bot can explore Minecraft for hours, adapt to its environment, and develop skills without specific preprogrammed instructions.

Training Methodology

  • Data Sources: The training involved extensive data collection from YouTube videos, Minecraft wikis, and community discussions to create a rich dataset for the AI.
  • Reinforcement Learning: Fan discusses using reinforcement learning from human feedback (RLHF) to align the AI's actions with human-like behaviors based on natural language prompts.

Future Applications of LLMs

  • Industries and Robotics: Fan foresees substantial applications for LLMs in software automation, gaming, and robotics, suggesting that AI agents could revolutionize various sectors by automating tasks and enhancing user experiences.
  • AI Safety Concerns: He emphasizes the importance of addressing AI safety as the technology evolves.

Getting Involved with LLMs

  • Encouragement for Newcomers: Fan encourages listeners to experiment and learn through available online resources and open-source models, suggesting practical engagement as the best way to understand AI technologies.

---

Key Takeaways

  • AI Agents vs. Traditional Models: Understanding the distinction between AI agents that can operate autonomously and traditional models that require prompting.
  • Importance of Environment: Minecraft provides a flexible environment ideal for developing open-ended AI capabilities.
  • Voyager's Capabilities: The bot's autonomous abilities highlight advancements in AI learning and decision-making.
  • Future of AI: The trajectory of AI emphasizes both scaling up powerful models and scaling down for specialized applications, hinting at a future rich with AI-driven innovations.

---

Conclusion This episode highlights significant advancements in AI through the lens of Jim Fan's research on large language models and autonomous AI agents. The conversation underscores the potential of these technologies to transform industries while also noting the challenges that lie ahead in ensuring safe and effective AI deployment. For those interested in further exploration, Fan's insights provide a starting point for engaging with the evolving field of AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. A long time ago, some people created an open-end exploration-based video game called Minecraft. It was a big hit. To this day, a lot of people still play Minecraft. More recently, some other people created a chatbot based on a large language model. They called it ChatGPT, and it was a hit too. In fact, ChatGPT quickly became the fastest-growing consumer application of all time. Well, our guest today has combined those two technology success stories into something pretty awesome. He and his colleagues built an agent that uses GPT-4, the brains behind chat GPT, if you will, to play Minecraft really, really well.

0:53More important than the bot's gaming prowess, however, the bot, which is called Voyager, has implications for how LLMs might be used going forward to automate and advance all kinds of tasks and processes across all kinds of industries and applications in the real world. Dr. Jim Phan is a senior AI scientist at NVIDIA and a leading expert in the field of large language models, among the hottest fields in AI and machine learning right now, to say the least. Dr. Fan is the co-author of the NeurIPS Best Paper Award-winning paper, Mind Dojo, Building Open-Ended Embodied Agents with Internet-Scale Knowledge.

1:28And his work has been featured in publications such as Nature, Wired, Science, and The New York Times. Dr. Fan is here to discuss the latest advances in LLMs, the potential applications of LLMs in robotics, game AI, and other fields, and the challenges that still need to be addressed in order to build truly general-purpose AI agents. So without further ado, let's get right to it. Dr. Jim Phan, thank you for joining us and welcome to the NVIDIA AI Podcast. Thank you for having me. So congratulations on the success with the NeurIPS paper and the Voyager bot and all of the attention it's so well deservedly been receiving.

2:07But we were chatting a little bit before we hit record, and I thought it might be good to start kind of at the basics. You've been working on AI agents for some time now, but the term might be new to some of our listeners. And in the listener's defense, or I should say, being empathetic to the listeners, there's been so much this year in particular about generative AI, LLMs, agents, all of these terms flying around, being misused in some cases. So might we start by having you explain what is an AI agent? What does it do? How does it relate to LLMs? And maybe you can just kind of take it from there to open up the conversation.

2:42Yes. For me, I have been fascinated by AI agents for all my career. And just to put a very simple definition, AI agents are AI models that can proactively take actions and then perceive the world, see the consequences of these actions, and then improve itself. That is a very simple definition. And that is in contrast to many of the logic language models that we see here today. For example, tragedy. You can ask it question and then it gives you an answer, but it never takes actions by itself directly without your kind of explicit instructions. But AI agents are these entities that can make decisions on their own.

3:24So earlier this year, recording in the middle of the summer in 2023, beginning of this year, I remember I experimented a little bit with something, I think it was called AutoGPT, that it was on GitHub and got a little bit of, I think it got some attention on Reddit is how I found it. And so it was a system that would take sort of a larger, broader instruction set and then prompt the chatbot on your behalf and then kind of go from there and reprompt it. And it even, you know, tried to write some code and do some other things, kind of learning as it went. To me, and I guess, correct me if I'm wrong here, but to me, one of the things when I hear the word agent that stands out is the ability to kind of act again and again without re-prompting.

4:10And then also, and this is where my understanding gets fuzzy, they may have the ability to retain what they've learned as they go forward. And I know there are ways around this and the models are always changing and advancing, but with something like ChatGPT or the other chatbots, it seems like they have a very limited memory, so to speak, and they don't carry context forward when you start a new conversation. Are these things relevant to how agents work or am I totally misinformed here? Yes, I think all the GPT is one of an early attempt on turning large language models, which are like passive kind of question answering machines, into active decision making agents.

4:50It's a super early attempt. And after that, the community has built a lot more on top of LOMs to give it decision making abilities. But I also want to mention that actually way before ChatGPT in 2022, AI agents had been around for a long time. And I guess I can give the audience a little bit overview of my own career and also my passion about AI agents. Perfect. So it all started in the summer of 2016 when I was an intern at OpenAI. And there wasn't a transformer. There wasn't large language models at that time. But OpenAI did have a very kind of strong vision for artificial general intelligence or HEI.

5:32And the project that I worked on during my internship was called OpenAI Universe. So the idea is, can we train an AI to control keyboard and mouse and then read pixels off a screen and let this AI use softwares in the digital space, just like a human would do? Right. So after listening to this description, you can see how general this vision is, right? Because if it works, then, you know, replying to email and playing games and browsing the internet all become subset of what this agent can do. So we would call that open AI universe. And for that project, back then we used techniques like reinforcement learning combined with a very sophisticated infrastructure so that the AI can actually read pixels and then control the keyboard and mouse using like Python APIs.

6:23And we integrated like thousands of different environments, like games and desktop software. And the part I was responsible for was the web browser. Can we kind of control the web browser using keyboard and mouse? And there are so many web apps and things that you can do using a browser. But in retrospect, that vision was great, but the technology wasn't there in 2016. And back then, we couldn't really solve it in a general way. Right. There's no way to tell the agent that this is what we want using any arbitrary natural language. And the agent can only work on the tasks that we explicitly train it on.

6:58So it doesn't generalize very well. But that was my first kind of trial in the field of AI agents and I fell in love. And also at that time in 2016, there were other AI agents going on like DeepMind's famous alpha goal where the agent beats the game of goal and then in 2019 open ai did a project called open f5 that plays the game of dota at human champion level right and at roughly the same time deep mind has alpha star that plays starcraft at human champion level so the advancements were really great um but that got me thinking right like all of these agents they can only do one scene out of time, like they play the game of goal or they beat the opponents at StarCraft.

7:44So there's a very clear objective, which is either to maximize the score you get in a game or to beat the opponent as fast as possible. And that got me thinking, right? Like, can we do more than that? Can we have a truly open-ended agent that can be prompted by arbitrary natural language and be instructor to do just open-ended and even creative things. So the time when I joined NVIDIA, I had this idea of turning Minecraft into a playground for open-ended, generally capable AI agent research. And that was how we got started on the project Mind Dojo. The thing that jumped out at me is the difference between a game.

8:27When my kids first started playing Minecraft, I was like, oh, that's cool. Like what's the object? And they said, well, no, it's not an object. Like you go around and you explore and you build things and you gather things and you know, you can build like all kinds of houses and environments and stuff. And I was like, oh, interesting. Yeah. Like I don't get it because I grew up on arcade games where you're trying to get points. Right. And so this difference between a specific objective and a more generalized environment kind of started to make sense to me that way. So anyway, for any listeners out there not familiar with Minecraft, open-ended game, there's not necessarily one clear objective or you're not trying to get a high score.

9:05Exactly. And that's why we find Minecraft to be almost the perfect primordial soup for open-ended agents to emerge because it sets up the environment well. And in the MindDojo work, we had three major components. And the first one is we designed a programmatic Python API around Minecraft so that you can actually control the game using Python code. And that would open up an interface for AI to drive a character in the game of Minecraft. Right. And as opposed to that earlier in your internship, you're not writing, you're not trying to get the AI to command the keyboard and the mouse. It's communicating directly with the Minecraft game?

9:48Yes, it is communicating directly with the game and you can directly issue commands and read the pixels from the game. And the second part of Mind Dojo is we actually collected, I believe, to this point, the largest open-ended agent database. And that is because Minecraft is actually the best-selling video game of all time. And there are about, I believe, 140 million active players on Minecraft. And just to put this number in perspective, this is roughly twice the population of UK and five times the number of people who know how to code. So it's a huge gamer population on Minecraft. And that means this enormous human mass just produces a ton of training data every day on the internet.

10:34And it just so happens that gamers are generally happier than PhDs. So they love to play the game and then stream and share the game experience on YouTube and on like also many other forums. So we collect a lot of data. One part is the YouTube database where we collected about 300 ,000 videos on Minecraft of people just playing it. And also 2 billion words in the transcript because the gamers like to narrate what they're doing as they're streaming. And then we also find that the gamers are so enthusiastic that they compiled an entire Minecraft specific wiki database. And that's about 7 ,000 pages of tables and every crafting recipe you ever need to know in the game.

11:20And we also script all of them. And then we have the Reddit forum. The Minecraft subreddit is one of the most active forums on Reddit. And we see people using it as a kind of stack overflow for Minecraft. They go there and showcase their work and ask questions. And that just becomes a great training data. You have all kinds of training data available. Yes, exactly. And we make all of these data and curations available to the research community. And in the MindDojo work, we also designed an agent that is trained on the YouTube partition of our database, where we have the agent learn to associate the transcript and also the video.

12:01And this association is actually a score between zero and one. So zero means the text is irrelevant to the video. and one means the text perfectly describes the video. Like for example, oh, I am about to chop down this tree and that if the player is actually chopping down a tree, it will be close to one. And you can see that this model that we trained on the YouTube data is a foundation model for reward function because it's encouraging the agent given a natural language prompt to match the behavior described in the video. And this is one of the earliest attempts at reinforcement learning from human feedback, but applied to Minecraft.

12:43And just to give the audience a little bit of background, reinforcement learning from human feedback, also known as ROHF, is a cornerstone algorithm that drives ChatGPT, where it aligns ChatGPT to be more towards what humans prefer. And we apply a similar technique using our foundation model trained on the Minecraft YouTube videos to guide the agent to be aligned with the human instruction. So we can ask that to chop down a tree or find me another portal or dig a hole in the ground, for example, and just give it an arbitrary language. So that was the work in MindToucho. How much more difficult, if at all, is it to train a model on video data as opposed to text?

13:31Yeah, so for us, we just use Transformer Architecture. and apply to video because you don't think of video also as a sentence just that each word is now a frame an image frame right okay and it's still a sequence of frames and anything sequential transfer is uh just uh the best model for it so we we trained uh similarly uh as you would for language modeling and and as somebody who like i said i have kids and they they learn i don't want to say everything but they learn a lot through watching youtube videos And back in the day, I used to make some YouTube videos. So I'm kind of wondering, and I would imagine your data set is so large, it doesn't matter.

14:10But when you were talking about assigning the score between the transcript and the video, I was imagining if the transcript is just the hosts of the show, the gamers, just talking trash with each other and making jokes and doing all of these things that aren't relevant. I was imagining if it might throw your training data off, but I would imagine that's just a drop in the bucket considering how much data you have. Yeah, so we did do some filtering so that we focus on kind of the key terms from Minecraft and then find videos around those key terms. That is one thing. But also, it's really hard to avoid people talking about sponsors and ads.

14:46We also found that when people talk about that, it is reflected in the transcript. So we actually know that when that person is talking about sponsorship, because the word sponsor corresponds to those videos really well. So that is like a natural filtering mechanism that just simply emerged from Charles Horvath, from this magical combination of Charles Horvath and large-scale data. All right. So props to the FCC for putting in those sponsorship rules that actually help AI scientists be more efficient in their work. Fantastic. Yeah. Thanks, Dan. So in reading some of the coverage, it looks like Voyager outperforms the other Minecraft playing bots that are around by quite a wide margin.

15:31And there was something, and we don't want to spend the whole time talking about Minecraft. I mean, I would, but we have limited time. But maybe we can transition in a moment. But there was something that caught my eye about a skill library and a difference between the bots that had the skill library and don't have the skill library. And then that got into things about creating a curriculum and learning and things that seem like they might be more broadly applicable. So maybe you could take that and run in the direction that to you is most important and talk a little bit about it. Absolutely. So when we're developing the first MindDojo paper, we found that the model that we train is only able to handle relatively short horizon tasks, which are tasks that can be completed in maybe a hundred keyboard and mouse click, for example.

16:18But for longer tasks, it's really hard for the model. It doesn't have a very strong kind of long-term finding abilities. And then GPT-4 came, a language model that is just really good at long-term planning, reasoning, and also coding. And then we decided to do a follow-up work called Voyager, which leverages the power of GPT-4 to write code in JavaScript. And there is actually a JavaScript API built by the enthusiastic Minecraft community that can do anything that keyboard and mouse can do in the game. So we leverage that JavaScript API and have GVT4 write code that can execute in the game. But then, you know, the code isn't always correct on the first try.

17:01Just like for human software engineers, we don't get the code typically right on the first try. We need to run it, debug, and iterate. And that's exactly what we mimic in Voyager, where we have GVT4 look at the output. And if there is an error from JavaScript or some feedback from the environment, GVD4 will do a self-reflection and then try to debug that code. And after debugging, this code took effort to write, right? And we don't want to throw that away. So we devised something that you mentioned called a skill library, where GVD4 can store the correctly implemented program into that skill library.

17:36And the next time, if GVD4 faces a similar situation in the game, but elsewhere, it can just retrieve that skill and then directly execute it And it doesn't have to kind of debug, write and debug again. So in this way, we built the skill library to enable something called lifelong learning in Minecraft, which is that as you run GD4, even though GD4 itself is a black box API to us, we cannot update its parameters, but it's still learning in the form of skill library, in the form of explicit code stored in a code base authored entirely by GD4. So as we just let Voyager do the exploration, we actually see it riding kind of dozens of different skills and just improving over time as a Minecraft player, even though GBT4 itself isn't being updated.

18:25So this is what we call the no gradient architecture, where, you know, traditionally, if you do deep learning projects, you will have a model, you will have a data set, and you train it using gradient descent to find some parameters. But here we introduce a new paradigm for Voyager where there's no parameters being updated and the learned model isn't matrices of floating numbers, but it's actually an explicit and interpretable and human verifiable code library. So that is kind of the new thing introduced by Voyager. Were you and your team surprised by how Voyagers performed? And then maybe as a follow-up, if you were, what were some of the things that surprised you?

19:09And as you were talking about the skill library, I was thinking, oh, did it come up with any skills that you'd never seen before or just sort of made you laugh or whatever the reaction was? but maybe even a bit broader, you know, as far as once you got it up and running, and I'm sure that it was tweaked along the way, everything's, you know, always a work in progress. But once it was up and running for a while, what was it doing that maybe you didn't expect, or maybe you thought like, oh, wow, that's really interesting. And that could maybe lead somewhere else. Yeah, that is a great question. I think it was quite a magical experience when the whole thing, when we put together the whole thing and just starts working.

19:47I can only imagine. Because we sit back and then we just see the bot playing Minecraft for hours on it and we didn't intervene. The whole thing is automatic. And the bot actually decides what are the next places to visit and what are the next areas to explore. So we just see it kind of doing traveling. And if it sees a monster, it will develop a skill to combat that monster. And if it runs out of food, it will find ways, creative ways to get food. Like, for example, if it's close by a river, that it will start fishing. If it's in a forest, that it will decide to hunt some animals for food. Yeah.

20:24And we see all of these behaviors just emerge from the Voyager setup, the skill library, and also this coding mechanism. And we did not pre-program any of these behaviors into it. We didn't tell it, oh, first you need to find wood and then stone and then iron. We didn't tell any of that. But somehow Voyager just figured out that to unlock the technology tree in Minecraft, you need all of these materials. And it would just go ahead and collect them by itself. Right. Our guest today is Dr. Jim Fan. Jim is a senior AI scientist at NVIDIA and a leading expert in the field of LLM's large language models, which if you've been following AI at all in the current year, 2023, you've come across the term at least once or twice, I would imagine.

21:10Imagine. We've been speaking kind of specifically about the work that, Jim, you and your team did building Voyager, a bot that leverages LLMs to play Minecraft quite well. Your career, your work previously and now is a segment of it devoted to games and AI and game agents, but it's more than that. So maybe you can talk a little bit about the work you're doing in some other areas. We talked off mic about software agents and robotics agents, and then kind of broadly about LLMs. You know, it's the hottest thing going. You know, myself, I rode the roller coaster this year from being super excited about LLMs to wondering if they're going to take my work away.

21:53I do a lot of writing work in addition to the podcast. And so that's one of the areas that people have been sort of, you know, a little fearful about, like, well, are these LLMs good enough to replace me? this is your thing so let's get into it um where would you like to start yes um so i think the future of lom uh like a major future trend of lom is better and stronger ai agents so i see at least three applications in the industry one is software agent uh where you know starting from the open air universe project that that i talked about at the beginning can we have ai that learns to control the software tools that we use?

22:33And then can we have it automate many of our digital lives? And most of us spend most of our lives every day in front of a computer screen or a mobile screen. So a software agent would really boost all of our productivity. And a second application area is gaming. And I mentioned Minecraft, but I believe AI agents have a lot of applications in many other games or even in the bigger metaverse concept as well. Can we populate metaverse like AR and VR with just human-like AI agents that not only talk like us, but also behave like us. That makes just the whole experience so much more real and immersive.

23:15So I think AI agents will have a big application in entertainment. So is that the idea of an AI agent in a game, is that similar to, and again, I'm leaning on my kids now, thanks guys, to an NPC, a non-player character and what those characters do? Is it kind of an advanced version of that or is it something different? Yes. I think AI can be applied to the NPC characters and those characters will feel alive and they will actually drive stories in a ways that even the game creators did not foresee. So I believe we'll see a big wave of AI native games coming out soon where like AI just redefines the whole storyline and AI customizes the story to each new player.

23:55So everyone can tell their own story. I think that is definitely the future of AI agents and games. So that is a second area. And the third major application area is physical agents, building AI, bringing AI agents and large language models to the physical world, and that is robotics. So I don't think robotics is yet at a general purpose model level, but I believe we are making significant progress towards was that. And actually, in October last year, my team had a work called Vima, V-I-M-A, that basically connects a large language model to a robot R. So you can instruct it using natural language and then have it perform robot manipulation tasks in the physical world.

24:42And this is just the beginning. It is not yet deployed products, but I believe in the next couple of years, we'll have a better robot hardware and we'll also have stronger robot brains powered by large language models and AI agents. And then Avid automate a lot of the platforms that none of us wants to do. You mentioned when you were talking about the development with the Minecraft project, that the release of GPT-4, the move from 3.3.5 to 4 was significant. Can you speak a little bit, and I know you're not working, you're not part of a team that's building the LLMs per se, but can you speak to kind of the state of LLMs right now and from the outside or even my sort of, you know, ringside seat, if you will, because I have the privilege of talking to people like you, but I'm not doing the work you're doing.

25:36It seems like things and the offshoot applications from them were moving very, very quickly. Yes. Or at least the things that the mainstream was exposed to were moving very, very quickly and maybe have slowed down a little bit, although it may just be that GPT-4 has been out for a little while. And so the big mainstream buzz has died down a little bit. And then obviously there are other models that are out there. I think as we're recording this today, Meta just announced that they were releasing a new model into the wild for commercial use. What's the state of play right now? What are you seeing with LLMs now and kind of the latest advances?

26:19And then obviously we're all interested. Where do you think the technologies are headed? And then by offshoot areas like robotics, you were talking about, what's that going to mean? Yeah, I am very excited by the future of LLMs. So in addition to AI agents I mentioned, I'm also excited by the no gradient architecture around LLMs. So we all know that GPTs and LLAMs in the LLM space, they're good at reasoning and also coding. But at the end of the day, they're still taking text in and spitting text out. But there are just a huge amount of tools that will make LLMs so much more useful. For example, can we augment large language models with search?

27:02with a search engine or with some long-term memory, like a vector database, or with kind of professional tools, like financial tools or like chemistry tools, those very specialized things that only human experts use right now. Can we augment LOMs with those software tools? And I believe all of these adding together is the no-grading architecture, where we don't need to fine-tune the underlying LOM, but it's more like the LOM is a core reasoning engine that drives an entire software stack to make it more useful than just question answering. So that's one direction I'm excited about. And the second direction is multimodal AI, where again, current LLMs, text in, text out, but we want more than that.

27:47Can it take images and videos and audio, 3D perception, and can it generate not just text, but also images and audio and all of these modalities that we're used to? So I believe in the future, technologies like speech recognition or stable diffusion like text to image generation will all become a subset of powerful multimodal brain, a single model that understands all of these modalities and the connections between them. And when we see that model, it will unlock a lot more applications and also help embodied agents like robotics as well. So that sort of brain-like multimodal model makes me start thinking about artificial general intelligence, AGI.

Read the full transcript

28:31Is that a step in that direction? Is that your line of thinking? Yes, I think it is definitely a step. For general intelligence to emerge, we will need to give it the full richness of the world. And also we need to give it agency where it can make decisions and take actions proactively and then observe its own feedback from the environment and improve itself. I think these are essential properties to our general intelligence. And so far, the language models are still falling short. But again, these fields, as you said, are moving very quickly. So I'm optimistic that we'll see these breakthroughs in the next couple of years.

29:07So then from a technical perspective, are there specific hurdles or milestones that you see in the future that you or some of your peers and colleagues are working on now that you see as kind of like, okay, we've got to get over this major hurdle to then unlock, I'm thinking in game terms now, right? To then sort of level up to the state where we can start addressing this next set of problems. Two things. One is a very powerful multi-model model can provide a great backbone, and then they can unlock many neural body agent applications. And the second is a better coding model, because these models will be better at reasoning and also long-term planning and self-debugging as well, which will be very useful in algorithms like Forager.

29:59Yeah. Self-debugging will be useful for somebody like me who's, you know, tempted all the time to just open a chatbot and say, write me an app that does blah, blah, blah. And then I'm just going to run it because I don't have the technical skills to validate the code. And, you know, from the sort of layperson consumer perspective, right, that's one of the, but it's something that humanity's dealt with, you know, for millions of years. But I'd sort of fine tune thinking about, oh, I can press a button on a computer, get it to do what I want it to do, and then I'll just run it. And I have no idea, but hopefully it does what I want it to do.

30:33And so, again, asking you as, you know, a user who, you know, using these LLMs or building on top of them and with them, as opposed to somebody who's now working on creating them. I don't mean to put you on the spot to answer for others, but there's been a lot of talk that, again, with GPT-4, to use this example, seemingly such a big step forward from the previous generation, and with the immense expense that goes into financial, human, environmental effects that goes into training these LLMs, that maybe the next breakthrough isn't developing, you know, GPT-6 or whatever, or GPT-5, I should say, or whatever it is, the next like giant model, but that maybe the foundational models we have now are good enough to really unlock a lot of things with, as you were saying, different ways of, you know, building agents and connecting them to other tools and that kind of things.

31:29And that the focus now is more on, I don't want to use the wrong words, but using them in different ways, in different contexts, you know, fine tuning and other things. What's your point of view on that? Is it, are you actually waiting for the next gen of LLMs because there are certain things that, you know, you're hoping or waiting for them to do? Or is it actually more of the case of like, yeah, we've got what we need at least for the next, whatever the chunk of time is to allow for a lot of, you know, really important innovation and breakthrough? Yeah. So how I see kind of this field growing is what I call scaling up and then scaling down.

32:08Okay. So first, we need to scale up even more because GME4 is good, but it's still not good enough. It still makes mistakes in coding. It's context length. It's memory. It's quite limited. Yeah. And it's not multimodal, right? So we still need more powerful models coming out and not necessarily from just OpenAI, but we have also seen many other competitions going on and also models for open source community. So I welcome all of these advances. We need to scale up to get more capability. And then after that, we can scale down because not every application would require the full capability of GPT-5 or 6.

32:48Right. But they will benefit from a very powerful kind of teacher model. And let's say they can distill like a subset of capability from it into a much smaller model or even open source one and then have it deployed efficiently. So I think the cycle of scaling up and scaling down will happen quite a few times as we move forward. So before we direct users to some of the places they can learn more about your work, for somebody who's listening, who, whatever their technical background may or may not be, who's been captivated, and I've talked to a lot of people who kind of refer to LLMs as AI, right?

33:26Because unaware of all the things that have been done. When you said you were an intern at OpenAI 2016, it actually, I did a little bit of double take It was like, oh, they've been around that. Yeah, they've been around that long. AI as a field, obviously, has been around much, much, much longer. But for somebody who is newly kind of has their attention grabbed by everything that's going on, what would you suggest as a way to get involved and learn how to work with LLMs? And I know that can mean a lot of things, so I'll leave it open-ended. Is it still getting the basics of computer science and engineering and developing that foundation?

34:02Are there sort of new skills and ways to learn? I see you nodding over the video link, so I'll let you take it from there. What would your advice be? Yeah, I think the best way to learn is just to do something by yourself. And there are so many resources online right now around all kinds of tutorials and code repos that are open source and also people having a very vibrant discussion on these topics on Reddit and also on GitHub. So I would encourage all of you to get your keyboard ready because we'll do a lot of coding. You know, run some of the open source models like Meta's latest Lama 2, which is a really great open source language model.

34:41It's not yet at the Chaggbt level, but it's getting close and you can get access to it and you can even run it on a very humble CPU machine. Right. So you don't need like super powerful hardware to start playing around with these models. So I would encourage playing around with it and do some problem engineering, try out things like vector database and combining LOMs with search. And please read more research papers from NVIDIA. We've got a very strong publication track record. Yes, to say the least. Feel free to reach out as well. I'm happy to answer any questions about our papers or some other papers that you find interesting in the field.

35:21Which is a great segue for folks who would like to learn more about your work, your team's work, the NVIDIA research papers you mentioned, where's a good starting point online where they can connect with some of your work in particular? Yeah. So my personal website is jimfan.me. I'm also pretty active on Twitter. I'm Dr. Jim Fan on Twitter, and I try to kind of custom the noise and really curate the best quality, latest AI research for you. So feel free to check out. Perfect. Well, Jim, this has been great. I could go anyway for a lot longer, So hopefully we can catch up again in the future. But it goes without saying, congratulations on Voyager, on the NeurIPS Award, on all the other accolades and coverage you mentioned.

36:09There's a great article in Wired that kind of walks through. If you were a listener out there and you're listening to Jim discuss, you know, how Voyager was created and the skill library and the curriculum and lifelong learning, all those different things. I thought the Wired article did a great, great job of laying that out. other coverage as well, plenty out there. Just Google Dr. Jim Pan and Vidya and you'll have lots of reading to do. And then you can hop on your keyboard and start trying this stuff out as you suggested. But thank you so much for taking the time and all the best with what you're working on now and look forward to catching up again, hopefully.

36:43Thank you so much to all for having me.

37:13¶¶

37:30Thank you.

From the publisher

For NVIDIA Senior AI Scientist Jim Fan, the video game Minecraft served as the “perfect primordial soup” for his research on open-ended AI agents.

In the latest AI Podcast episode, host Noah Kravitz spoke with Fan on using large language models to create AI agents — specifically to create Voyager, an AI bot built with Chat GPT-4 that can autonomously play Minecraft.

AI agents are models that “can proactively take actions and then perceive the world, see the consequences of its actions, and then improve itself,” Fan said. Many current AI agents are programmed to achieve specific objectives, such as beating a game as quickly as possible or answering a question. They can work autonomously toward a particular output but lack a broader decision-making agency.

Fan wondered if it was possible to have a “truly open-ended agent that can be prompted by arbitrary natural language to do open-ended, even creative things.”

But he needed a flexible playground in which to test that possibility.

“And that’s why we found Minecraft to be almost a perfect primordial soup for open-ended agents to emerge, because it sets up the environment so well,” he said. Minecraft at its core, after all, doesn’t set a specific key objective for players other than to survive and freely explore the open world.

That became the springboard for Fan’s project, MineDojo, which eventually led to the creation of the AI bot Voyager.

“Voyager leverages the power of Chat GPT-4 to write code in Javascript to execute in the game,” Fan explained. “GPT-4 then looks at the output, and if there’s an error from JavaScript or some feedback from the environment, GPT-4 does a self-reflection and tries to debug the code.”

The bot learns from its mistakes and stores the correctly implemented programs in a skill library for future use, allowing for “lifelong learning.”

In-game, Voyager can autonomously explore for hours, adapting its decisions based on its environment and developing skills to combat monsters and find food when needed.

“We see all these behaviors come from the Voyager setup, the skill library and also the coding mechanism,” Fan explained. “We did not preprogram any of these behaviors.”

He then spoke more generally about the rise and trajectory of LLMs. He foresees strong applications in software, gaming and robotics and increasingly pressing conversations surrounding AI safety.

Fan encourages those looking to get involved and work with LLMs to “just do something,” whether that means using online resources or experimenting with beginner-friendly, CPU-based AI models.

More from NVIDIA AI Podcast

All 115 episodes
NVIDIA’s Jim Fan Delves Into Large Language Models and Their Industry Impact - Ep. 204NVIDIA AI Podcast · 38 min
Listen in VO