NVIDIA’s Annamalai Chockalingam on the Rise of LLMs - Ep. 206

23 Nov 2023 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NVIDIA AI Podcast - Episode 206: NVIDIA’s Annamalai Chockalingam on the Rise of LLMs

Podcast Overview The NVIDIA AI Podcast explores how cutting-edge technologies are transforming various industries, focusing on innovative breakthroughs and sustainable efforts. In Episode 206, host Noah Kravitz interviews NVIDIA Senior Product Manager, Annamalai Chockalingam, about Large Language Models (LLMs) and their current and future influence.

Key Concepts What are LLMs?

  • Definition: LLMs are a subset of generative AI, specifically dealing with language processing tasks such as recognizing, summarizing, translating, predicting, and generating text.
  • Architecture: Built on transformer networks, which allow for advanced processing without the need for extensive data labeling, thereby learning patterns from unstructured data.

Enabling Factors for LLMs

  1. Large-scale Datasets: Growth in internet usage has led to the creation of vast datasets for training models.
  2. Advanced Computational Infrastructure: Enhancements in computing technology enable handling immense datasets efficiently.
  3. Innovative AI Algorithms: Improvements in algorithms allow for better parallel processing of data.

Applications of LLMs

  • Text Generation: Creation of content based on prompts.
  • Summarization: Condensing large texts into key points.
  • Translation: Converting text from one language to another.
  • Instruction: Providing actionable directives to users.
  • Conversational Interaction: Engaging users in dialogue, often seen in chatbots.

The Business Perspective

  • Enterprises utilize LLMs to innovate, enhance customer experiences, and gain a competitive edge. They are also focused on safe and responsible deployment strategies to ensure trustworthiness in model outputs.

Emerging Techniques

  • Retrieval Augmented Generation (RAG): This method enhances LLM responses by providing models with real-time data or integrating third-party APIs, leading to more contextually relevant outputs.

NVIDIA's Role in the LLM Ecosystem

  • Full-Stack Platform: NVIDIA offers a robust computing platform for LLM development, including hardware, software, and services tailored to over 4 million developers and 1,600 generative AI organizations.
  • Nemo Toolkit: A platform for building and deploying generative AI models, featuring Nemo Guardrails for AI safety by allowing developers to implement application-specific safety measures.

Future of LLMs

  • Chockalingam emphasizes that we are in the early stages of LLM development. Future advancements will likely entail:
  • Enhanced customization techniques for applications.
  • Advancements in multimodal LLMs that integrate various forms of data beyond text, such as images and audio.
  • Continued emphasis on operationalizing LLMs for enterprises to ensure scalable and trustworthy applications.

Advice for Developers

  • Get Started: Engage with LLMs by using applications like ChatGPT, and explore pretrained models in NVIDIA’s NGC catalog.
  • Training Opportunities: Attend events like NVIDIA's LLM Developer Day to gain insights from experts and hands-on experience in developing LLM applications.

Conclusion The podcast underscores the rapid evolution of LLM technology and its vast potential across various sectors. It encourages both developers and enterprises to actively participate in shaping the future of AI through experimentation and responsible implementation.

---

For more details and resources, visit [NVIDIA.com](https://www.nvidia.com) and check out their dedicated generative AI web pages.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. Today, we're exploring the realm of LLMs, large language models. Large language models are more than chatbots. They're reshaping the entire world of AI. And our guest today is Anomali Chakalingam, a senior product manager in NVIDIA's developer product marketing group. AC brings a unique blend of tech and business expertise to the role, with experience ranging from working on sustainable energy products at Tesla to strategy consulting with Accenture to writing firmware code for a defense contractor. There's a lot to get into here, and obviously LLMs are the word of the day with good reason.

0:54So let's dive right in. AC, thank you so much for joining the NVIDIA AI podcast, and I'm looking forward to diving into the current and hopefully future a little bit of LLMs. So welcome. Thanks, Noah. excited to be here and excited to chat with you. Looking forward to our conversation and I'll share the little I know about this space. Excellent. So can we start actually back up just a little bit and share a little bit about your own journey into deep learning and then what drew you to working on LLMs at Invidia? Yeah, for sure. So growing up, I've always been super fascinated about technology and computing and how it impacts everyone's day-to-day lives.

1:33We've seen the transformation from having no personal computers to no cell phones to now having cell phones and the proliferation of technology, the web. And now you can even order a car to come at your doorstep at any moment in time. And I've always been super fascinated about that. And I'm an engineer by heart and by practice. So I enjoy delving into the details of technology and understanding how all of these things work, very inquisitive. And my path has kind of led me to deep learning and LLMs and super fascinated about the impact that it can have on society in general. So if I did my homework correctly, you did a joint MBA, master's in computer science and business, and then that kind of set you, I shouldn't say set you, you're probably already on the path.

2:22But then from there, you kind of move through some of these different experiences we're talking about in the intro. So maybe tell us a little bit about that and then about when you joined NVIDIA and what your current role is like. So I did electrical engineering in undergrad, learned a lot about engineering and how things work and lower level compute layers. And that's kind of how I did firmware engineering for a defense contractor. And from there, I was always someone who's interested in learning about the broader business impact and how these technologies can affect people in unique ways, right?

3:02And that's kind of what led me to strategy consulting and then grad school at NYU to learn a little bit deeper about both the technology as well as how it impacts people and businesses. And I joined InMedia about a year and a half ago and then have been spending all my time on the world of LLMs. I like to say that I started on LLMs before ChatGPT was a thing or LMs were a thing. So kind of, uh, with our ever shrinking, uh, you know, timeline of history here, that's, uh, that, that, that's not just, um, you know, what did they say? It's not just window dressing, right? It's a thing. You're an OG.

3:36So are you out, are you out in, uh, on the West coast near NVIDIA HQ? Are you in New York? Where are you located? I am out in HQ right now, uh, close to HQ in the South state. How's California been treating you? Yeah, it's been great, beautiful, sunny California, lots to do, lots to explore out here and enjoying it. Well, as a fellow New York City to California transplant myself, you know, I feel like I can say welcome. So welcome. So let's get into it. Maybe we can start since we've got you here. Can you give the audience kind of a succinct definition of what an LLM is? And then I'm sure you'll have no problem kind of diving a little bit deeper.

4:14Awesome. Yeah. I think that's a great question, right? There's been a lot of noise and a lot of buzz generated by generative AI, LLMs. But the way I think about it is that LLMs are a subset of the larger generative AI movement or the paradigm. So LLMs are kind of like deep learning algorithms that can deal with language, right? They can recognize, summarize, translate, predict, and generate language. And language is kind of the code of thought, right? It's the way humans think. So key differentiator between LLMs and kind of other model architectures or other areas is this model architecture called transformer networks, right?

4:51There's a very close connotation between LLMs and transformer networks. Transformer networks were kind of a seminal piece of work released by Google back in the 2017 attention is all you need paper. And it's kind of changed how we do things. You can trace the authors on that paper. You can trace them out to so many of the companies doing the LLM stuff now. It's wild. Exactly. They've kind of started this entire wave and this ecosystem. So these models are deep neural networks that are built using unsupervised learning techniques. So the key differentiator is that you don't need to spend time labeling all your data and making sure that they're prim and proper.

5:36You don't need to do all that work, right? You can just feed it into this magical model architecture. And then after throwing a bunch of compute at it, it'll learn all the underlying patterns and can understand the world and can talk about a whole bunch of different topics. It's general knowledge and not targeting a specific problem. Right. I'm curious your take on it. you know, the chatbot interface with ChatGPT certainly was a watershed moment and, you know, drew an enormous amount of early interest in people actually going on creating accounts and using it and everything. Generative AI was kind of in the getting mainstream-ish a little before that with some of the image generation models.

6:22The relationship between LLMs and chatbots, and maybe you'll get into this through the conversation, is the chatbot sort of just the beginning for how we interface with LLMs, or is it really a paradigm in and of itself that we're going to look back on? Yeah, I think that's a great question, right? Chatbots are, of course, the first signs of success that we've seen of how to leverage this underlying technology. It provides a great user experience of how you can talk to a bot just through natural language and get bad responses about a whole bunch of different things that are incredibly useful. But I'm sure as this space evolves, we'll see different application type user experiences that pop up.

7:06Something that we're also seeing is a lot of traction in search. How do you search? Not necessarily a bot interface, but just how do you search and a whole bunch of other things. There were jokes. I don't know if they were flying around, but sort of bad jokes. I remember, you know, making and hearing from friends about how, like, the chatbot interface was both the most, like, brilliant breakthrough to bring this technology to the masses and also, like, throwing the whole field of personal computing back 30 years to DOS, you know, kind of at the same time. Exactly. Yeah, chatbots have been around for a while, right?

7:40It's nothing new that's come out. It's really how you augment those chatbots with a brand new technology and experience is completely different now, right? It's undeterministic. The responses that you get each time from these bots are not the same. And that's partly the excitement where you can ask it anything and it'll somewhat give you a different response every time. Yep. So let's get into what's going on at NVIDIA with LLMs and the work you've been doing. Obviously, NVIDIA at the forefront of AI, but talking about LLMs, we've talked about on the podcast before, but talking about NVIDIA's work with LLMs.

8:15Can you lay out for us kind of how the company, how you view the potential, the opportunity and the challenges ahead? Because it's, you know, for I mean, I think we just talk about it all in terms of the big explosion of generative AI and LLMs and chatbots. And it's new stuff for a lot of folks out there and in the business world and all the different areas you plan. So, you know, maybe you can give us kind of an overview. Yeah, I think you hit it. The promise of LLMs are enormous. It's going to impact every industry. It's going to impact everyone's society, from developers to enterprises to startups to consumers and how we live and operate.

8:56The promise is enormous. And ChatGPT, like we talked about, has showed the early signs of success of this space. And we're seeing a lot of applications and enterprises, particularly, who are tapping into this and are trying to drive innovation, trying to develop new customer experiences, and also using it in their internal operations to become more efficient and being a competitive advantage. That's kind of like the early areas that we're seeing. But we're still in the early days. There's lots of work and lots of innovation still to be had. A lot of what we see in terms of products is still very researchy today.

9:35It's not necessarily battle-hardened. What is true today may not be true tomorrow. Exactly. And there's a lot of new research coming out in how do we better do things with LLMs. There's a lot of focus on how do you productionize and operationalize LLMs now that the core underlying work behind how do you get decent responses from LLMs or done. It's like, okay, now how do I scale it? How do I productionize it? How do I operationalize it? How do I serve it for millions of users? How do I make sure that my LLM applications and systems I can build in a safe and trustworthy manner that's somewhat repeatable and not necessarily hallucinating all the time?

10:18And how do you verify what comes out of the model, right? And so just to hone in for a second, When you talk about, use the word safe and talk about safety in this context with LLMs and in particular in AI safety more broadly, do you have kind of a working definition or how would somebody who doesn't quite get what that means, how would you define AI safety? Yeah, I think that's a good question. And I don't mean to put you on the spot, but, you know, kind of because I know it's a little bit tricky, but in thinking about an LLM that's generating some sort of language output, what does safety mean in that context?

10:52Yeah. So I think when you think of safety, you want something that is trustworthy, right? You want something that is reliable. The measure of safety also depends on the kind of users you're serving and the applications that you're building. What is safe in one application may not be safe in another application, right? For instance, you're building a customer service chatbot for your enterprise. So what is safe in that domain is, you know, you don't want to talk about your competitors or you don't want to talk about your internal business workings. but that that very same thing may be safe in another on their domain or their space and like ai safety is you know it's about responsible development and trustworthy and repeatability so that those are kind of the three big areas got it so it's not like a blanket sort of um we don't want it saying the i'm dating myself with the fcc's seven naughty words like it's much more than that and it's also context and sort of situation specific or it can be exactly exactly Exactly.

11:50And this problem, you know, you can solve it every layer of the stack, right? You can solve it at the model layer where you're trying to, you know, train it on verified data sets. Or you can also solve it at the application layer where you're kind of putting guardrails around your model to figure out, okay, this is safe for my application or not, right? So it's a combination of a whole bunch of things and how you can solve these problems. So I've gotten the sense, mostly from I'm hosting the podcast, but some events I've been to in the past few months as we record this from, you know, every once in a while, I get to rub shoulders with engineering folks who are building the things we're all talking about, right?

12:28And the sentiment I kept hearing over the summer in particular was, I've never seen things moving so fast. Like, yes, you know, whatever product you're talking about, whatever research paper we're talking about is a huge step forward or a notable step forward, but it's not slowing down. There's like new ones every day. Is the pace, as we record this now, kind of in early fall of 2023, is it still breakneck pace? Is it slowing down? Is it even speeding up more? What's the rate of development in the LLM world? Yeah, I mean, it's quite astounding, right? It's very hard for everyone to keep up with the pace of innovation, which is also very exciting to see how developers and the community is coming together to really accelerate the pace of innovation.

13:14There's a lot of work done by developers in the open source community. So it's allowing everyone to really understand and keep up with the pace of innovation just because things are open source. But I think there's another trend that is kind of very interesting. Before, there used to be a lag between what was research and what was product and today now a lot of product is research and research is product so that kind of changes the mental model of what people are doing and when they launch research it's almost a launching a product and it comes with an entire muscle of launching a product right they're not just a paper anymore it's interesting to hear you say that because I, I associate it with kind of the, the mobile phone, the smartphone revolution or whatever word, kind of the first wave of, you know, in the two thousands talking about how everything and then software coming out after that and the app stores and talking about how everything was released in beta, right?

14:14So releasing a beta was sort of derogatory for like a piece of hardware that wasn't quite had too many bugs. But then we moved into this era where it was like, well, no, we're moving to software and everything online with the internet and everybody's got an internet computer. So you can ship and then update and ship again and update and so on and so forth. And it's almost like it's even a step, I don't wanna say backwards, cause it's not backwards, but a step closer to the, you know, it's like, oh, forget beta, we're in the age of research, you know, and everybody's shipping these things out and then being able to collaborate on them.

14:46And anyway, sorry, it's just, even from the outside, it's amazing to see. Yeah, yeah, we have a strong community, right? And we're, as NVIDIA, we're kind of seeing it all and seeing many of these developers building on the NVIDIA AI platform. It's been kind of years in the making for us. We've seen the entire AI wave really take off. There's several, over 1 ,600 generative AI companies today building on our platforms in combination of our full-stack solution, hardware, software, services. So I want to get a little later, I want to get into some of the specific NVIDIA tools and frameworks that are available.

15:29But before we do, you talked kind of at the beginning and it's in your background. And from my own perspective, I'm fully the same way that one of the great things about working with technology and technology is seeing it help people in real life, in everyday life and big problems and small problems. What are some of the problems that LLMs and, you know, kind of this wave of gen AI powered by LLMs, what are some of the problems that they can actually solve, you know, kind of in computer science, so to speak, and then also more broadly? Yeah, I think AI is not necessarily new. It's been around for quite some years, maybe the past decade or so, and there's been different people trying to hack away at the problem with different chases in time.

16:14But over the last decade or so, we've kind of seen AI or traditional AIs, as I like to call it, right? It's largely around understanding the world around you and detecting patterns with data for solving specific use cases or specific problems in specific domains, right? These models don't necessarily know about everything, like in this new wave of generative AI, right? They're general in a sense, and they can be used for not only understanding the world around you, but generating new content based on data it's trained on, based on its understanding of the environment and world. And this general capability has really happened now because of kind of like three key innovations that have happened over the past few years.

16:58So the first is the availability of large-scale data sets. We as users have been kind of putting everything on the internet, some good, some bad, all the way from funny videos to very educational content for science, pop culture, or you name it, all these different forms. And all this data is now out there and available, and it's a trove of information for these models to learn from. And we've figured out how to effectively scrape the internet for this mountains of training data. Oftentimes, when you're building these models, trillions of tokens is a common theme these days. The second is advancements to compute infrastructure, right?

17:38You're able to crunch these mountains of data in somewhat of a reasonable timeframe, right? There's been a lot of advancements in hardware. you get access to these hardware, you don't need to go build out your own entire data center. Put it all on a pram, you can easily access through the cloud. There's ways to access computer infrastructure and the power of computer infrastructure is incredibly powerful. And the third is these advancements in AI algorithms like the transformer architecture, etc. So there's been a whole bunch of model architecture innovations, which have all utilized the unsupervised learning techniques.

18:10Particularly, you can figure out just from a whole bunch of data that you feed it, what are all the different patterns? What are all the nuances? What are all the intricacies? And pay attention to some of the most important data and disregard some of the data, right? And you can really learn. And all those algorithms have really come about now. And the second piece is, you know, since you're oftentimes these models, you know, you need thousands of compute nodes and it's a distributed computing problem, right? And we figured out a way to how do you process these data in a non-sequential or parallel fashion.

18:43So you can take this large compute problem, kind of decompose it into smaller bite-sized chunks so that these computers can really hack away at one small problem and then feed it back into the large system, like how you divide data models, et cetera. So those are kind of the advancements that have really allowed for us to come to this generative AI wave. And so what are some of the applications that are starting to emerge? And it's early days, right? Just to remind ourselves, remind everybody that even though it feels like, you know, there's been so much and there have been so many advancements still in kind of the early innings of generative AI.

19:20But what are some of the applications that are out there, whether, again, research or product or things that, you know, just being thought about beyond chatbots? Yeah. So language is super interesting, right? You can encode all knowledge in the world to some form factor of language, right? There's a lot of language that's embedded in all sorts of ways. And it's kind of the way humans think, right? Language is the way humans think, humans communicate. Language is also the way machines think and machines communicate, right? We send machine-understandable language, which is code, to these models.

19:56And it's also the way biological systems think, right? They think through molecular structures or proteins. The language of biology is these proteins and molecular structures. So language is kind of the way of thought for almost every system. So when you're able to understand language and generate new language, it's almost like there's a new computing paradigm that's coming up. And largely, there's like five broad categories of things that you can do with language. You can generate, you can summarize, you can translate, you can instruct, or you can chat. And you can talk to everything through natural language.

20:35And language is beyond just text. Right. And instruct meaning like give a computer an instruction. Exactly. Or take an action. Not just a computer. Yeah. No. Important point. Well said. Yeah. Yeah. And it's across modalities, right? Like, you know, across machines, across humans, across biology. Any domain has a language of its own. You can kind of go solve that. And with a combination of these modalities and actions, you can build applications. Right. So every application will do a combination of generation or summarization, translation, etc. And you can solve any problem. And there's kind of like different user experiences and type of apps that you can build.

21:17Chatbots are one big one. Second is like search, right? How do you search this data, particularly in the enterprise? And then how do you create content or how do you analyze information? The one key thing that we're seeing is this concept of retrieval augmented generation. When you connect these models to other data sources or third-party APIs, these models just get that much more powerful, right? If I can ask a question about, it's like, think you're a human, right? If I give you the access to Google, you're just that much smarter and you can figure out things as you learn by doing a Google search.

21:55Similarly, if you give these modeled access to information that it didn't have before or information that is live or current, it can reason through that information to give you more appropriate responses. Right. So you're talking about two types of retrieval, one from live sources online, being able to browse the web, go to a URL, wherever else. The other, I think we're talking about training an LLM that's already trained on the corpus of the internet and books and everything else up until whatever cutoff date. And then also saying to it, hey, here are all the recordings of the NVIDIA AI podcast ever.

22:34And so you're also going to be trained on that specific body of knowledge. Or here is an enterprise customer's body of HR documents so that you can train it to also specifically grab from what you've pointed it to. Exactly, right? So there's kind of like two ways of customizing models, right? I think we've talked about it when we talk about safety a little bit. You can customize the models themselves to understand the knowledge better. And that's through concepts like RLHF or fine-tuning or these different customization techniques that you see popping up, right? There's entire work of things going on there.

23:13The second is at the top layer, at the application layer. At runtime or at inference time, which is where RAG fits in, right? It's retrieval augmented generation. You can retrieve from data sources at inference time. At inference time, okay. And a lot of information is not necessarily always in the internet, right? Enterprises have a trove of data that are proprietary. They want to keep within themselves. So how does a business leader, for instance, talk about, understand maybe what their sales was for the past quarter? All that information is encoded somewhere in that enterprise's knowledge base.

23:48So you can really talk to your data through the power of sales. Which is, for those who've worked in an enterprise situation, you know, I think it's probably universal across whatever your discipline, you know, your product or your area of the company is, that once you hit whatever the critical mass is, it becomes incredibly difficult to leverage just all of that institutional knowledge. And so, I mean, LLM is applied to that situation. When I sort of stop and think about it, it's kind of mind-blowing, the potential there. Exactly, right. And we're still in the early innings, right? We're just starting to see these enterprises having access to this technology.

24:29So more to come and let's see how this shapes out. So the word ecosystem is a word that's used a lot within technology and certainly within NVIDIA and talking about how hardware and software work together. Let's talk about LLMs in the context of what an ecosystem is for an LLM and kind of how NVIDIA is involved. What's NVIDIA's role in the whole sort of stack or ecosystem that LLMs live and thrive in? So NVIDIA, you know, we build kind of full stack computers and we call it accelerated computing, right? We try to re-engineer all layers of the stack from the chip to the systems optimization layers to the application layers.

25:12We try to build full stack solutions for each domain. And LLMs are a domain that needs full stack innovation and computer to run on. And this is a data center scale distributed computing platform. Data center is becoming the computer and apps run on the data center of computers while being consumed at the edge on personal devices and different kind of form factors. So we have a multi-layered stack of offerings and partner offerings that developers and enterprises can engage with us on. At the very bottom, we have the compute infrastructure layer where we continue to innovate on with specifically things like Hopper architecture.

25:55So this is where GPUs live and these computers run, kind of compute and do math problems. And then the next layer is kind of like the systems optimization layer, right? Since you need kind of thousands of nodes or hundreds of nodes to do many of your compute problems, how do you kind of run those problems across these different compute nodes? So there's a lot of parallelism techniques that we have innovated on and the community has innovated on. And we welcome everyone to do that. And that's that layer. And the next layer is what are the right models and right model architectures. So we talked about transformer models.

26:38And there's a whole bunch of different variations of transformer models that are out there. And there's a lot of work going on there. And we are also continuing to innovate at that layer. And then once you have these foundation models, which are trained on large data sets, then oftentimes you kind of need to figure out how do you customize these, right? How do you customize these models to make it useful for your particular applications? Or how do you make it safe for your application? So we have a whole bunch of offerings at that layer as well. some which are, you know, NVIDIA built like Nemo, Triton, TransferRT, LLM, etc.

Read the full transcript

27:16That help with model training, model customization, model inference, and as well as partner offerings like, you know, Hiding Face or DeepSuite. So, you know, developers and the community has choice of what they want to use at any layer of the stack. Right, right. And then at the very top, you know, there's how do you build application systems? And there's a lot of work going on there. So our kind of pitch is, you know, we offer the full stack excited computing platform developers and enterprises. You know, they can choose from our depth and breadth of offerings at any layer of the stack and we welcome them to engage with us.

27:51And our ecosystem currently spans over 4 million developers, 1 ,600 plus generative AI organization, plenty of apps being built on top of our platform. Can you talk a little bit about what Nemo is, Nemo Guardrails and kind of the associated, anything associated with that? But Nemo is a name I've been hearing recently. Yeah. So, you know, Nemo is kind of a platform for building, customizing, and deploying generative AI models anywhere. So, you know, developers can use as a framework or as a service or in different form factors to develop their generative AI models and deploy them anywhere and customize them for their particular use cases.

28:32And, you know, recently we launched things like Nemo Guardrails, which would help with AI safety. So they're kind of like an application layer offering that allows you to add guardrails to your application specific to your application. So it's like a programmable toolkit. You can go in and program different things to align it to your situation specific safety. Exactly. Exactly. Exactly. And kind of like the inception of all this layers is a layer underneath, right? It's a research contribution that we did with the Megatron LLM open source GitHub repo, right? It's kind of like a seminal piece of work in the community that really helps with how to utilize all these advanced parallelism techniques on NVIDIA's GPUs.

29:19And many models and many frameworks today are built with a derivative of Megatron LLM. And, you know, we even used some of these technologies to build some of the largest models at the time in 2021, which was called Megatron 530B. Right. So it was a 530 billion parameter model, very large. It was state of the art at the time. And now, even now, we're continuing to build models. And still arguably the best named LLM that I'm aware of anyway. It's tough to beat Megatron. Yeah, exactly. We're focused on taking all these research innovations that we're continuing to do across the stack to the enterprise with NVIDIA AI Enterprise licensing that comes with enterprise and production-grade support, stability, security, and allowing enterprises to shorten their time to market.

30:10It gives them an easy-to-use set of software that's ready to go. Excellent. So we've delved deep into kind of the near past and certainly the present moment. And this is the unfair question that everybody wants to hear the answer to. What's next? Where do you see LLMs headed, you know, next three years, five years, even three months out from now? We may have a, you know, something really important that comes up, but where is it all headed? Yeah, I think we're, you know, I think we basically touched upon it before, but like We're still in the very early innings of LLMs. We've seen a lot of people really focus their energy and efforts on building the tools and platforms to make this technology successful.

30:55And now we're starting to see more and more people trying to build scalable applications to go to the end users and enterprises. So that's kind of like the next wave of things that I see happening. It goes back to the question of how do we make these models useful for specific applications? you know there's a whole bunch of customization techniques that are coming up you know rag as a concept right retrieval augmented generation how do you connect these llms and make powerful systems by by what's what's the best way to do that how do you production a lot how do you productionize and operationalize these llms and efficiently and cost effectively and scale these systems you know how do you build trustworthy llms but more importantly i think generally it's going Beyond language, right?

31:40We're seeing a lot of crashing language for machines. Code, code LLMs are a big thing. Multimodal LLMs. So when you augment your LLM with other modalities like visual or visuals or audio, you give just that much more knowledge to the model and it becomes just that much more useful, right? Think you as a human, right? The stuff that you can understand just through text, if I augment that with, you know, some visual or some audio, you just understand it better. Words, right? Exactly. Exactly. That's the same concept with these models. We had Dr. Jim Phan on the podcast not too long ago, and I've been following him online, and he's got these great, you know, pretty pithy little, usually a video with a caption kind of explaining.

32:25And that was the first time I saw, you know, I knew of the concept of multimodal LLMs, and we talked about it a little bit on his appearance on the show. But following him was the first time I saw it in action and saw the LLM analyze a picture. I think it was a picture of a toolkit and an instruction page next to it. It was basically just, do I have the right tools for the job? And there's something about just seeing it with a very simple concept, and then you realize how complicated it is to parse that out. And it feels magical. So cool. Yeah. Model behave the same way. They kind of think in a similar fashion.

33:06You give them more and more information across modalities. They're just able to understand and generate better things. They're able to understand context. So one of the things that I've encountered recently from some folks who are kind of in the business world, but interested in this stuff and following it on their own, but not necessarily a research scientist or software engineer, they're sort of wondering about kind of the application layer, right? And if the models themselves are still changing so rapidly, how do the applications sort of keep pace? Or, you know, how does it all work with developers?

33:45So NVIDIA works with a lot of developers, a lot of generative AI companies and all other kinds of AI companies, as you mentioned. Do you have advice for, you know, folks who want to dive in and play with LLMs, even use some of the NVIDIA tools you talked about? and I was thinking kind of more broadly, you know, business-wise about how fast it's changing, but, you know, sort of open question. What's your advice for developers? Yeah, I think the space is moving fast. So I think the best way to learn is just to get your hands dirty and get started somewhere. You know, there's, like I said, there's, you know, it's a multi-layered problem.

34:19Whichever problem that you're most interested in solving, go try to tackle it. You know, obviously the first thing that anyone should do is to just experience the power of these things, right? Yeah. Go use some of the popular applications out there, like ChatGPT, for instance. Start playing around with that, understand, figure out how those things work. And then if you want to get a little bit deeper, go experiment today with some of the more state-of-the-art models. We can find more state-of-the-art models out in the open source. You can also go on NGC today to find a bunch of pre-trained models across different domains and different types of capabilities.

34:57these, you can also experience it on there today. And then it's that journey, right? It's that first you experience it and then you dig a little bit deeper, dig a little bit deeper, figure out what tools work for you, figure out what tools don't work for you. So just start somewhere. You can use cloud APIs or other things to just start. So I think that's the big piece of that. So start today. And we're offering kind of hands-on training and practical guidance for anyone who's interested to learn and get deeper into this space on the November 17th. So join us on LLM Developer Day to get access to NVIDIA experts and learn from them on how to best develop applications in fall of 2023.

35:40Excellent. We covered a lot of ground and this has been great. It's kind of a nice mix of relatable and easy to follow, but also getting under the hood a little bit for, or I should say of this technology that's really just kind of captured many forms of imagination, let's put it that way, over the past year or so. So to put you on the spot one more time, you've already kind of predicted the future a little bit. Any words kind of in sum, whether it's parting wisdom, insights, advice for folks listening out there from the world of LLMs at NVIDIA? Yeah, there's a lot of innovation happening, a lot of excitement.

36:19So get your hands dirty and get started. I think that's kind of the biggest piece of wisdom I want to give away, just get started somehow and learn about this. It's going to change the world. Excellent. Obviously, NVIDIA.com is the gateway to myriad resources about all of this stuff and then some. But any particular places online you'd direct listeners to start if they want to learn about LLMs at NVIDIA specifically? Yeah, go to Inmedia.com and go to our Genitive AI webpages and learn about all different tools and offerings through there. There's also a whole bunch of blogs that we have that try to educate developers on different technologies.

36:58But even apart from Inmedia, just go to the internet. There's a whole bunch of great, great YouTube videos and medium articles and stuff just out there that's ready to be consumed. Excellent. Well, AC, thank you so much for the time and all of the insight and knowledge. And I want to make some kind of bad joke about, you know, well, let's catch up again in six months when everything's changed 180 degrees. But I would imagine, you know, whether it's six months or a little further down the line, we'll have to do it again and recalibrate on the state of the art and all the great stuff you guys are working on.

37:34Awesome. Yeah, I'll be looking forward to that. Thanks for chatting and enjoying our conversation. Me too. Thank you. Thank you.

38:11Thank you.

From the publisher

Generative AI and large language models (LLMs) are stirring change across industries — but according to NVIDIA Senior Product Manager of Developer Marketing Annamalai Chockalingam, “we’re still in the early innings.”

In the latest episode of NVIDIA’s AI Podcast, host Noah Kravitz spoke with Chockalingam about LLMs: what they are, their current state and their future potential.

LLMs are a “subset of the larger generative AI movement” that deals with language. They’re deep learning algorithms that can recognize, summarize, translate, predict and generate language.

AI has been around for a while, but according to Chockalingam, three key factors enabled LLMs.

One is the availability of large-scale data sets to train models with. As more people used the internet, more data became available for use. The second is the development of computer infrastructure, which has become advanced enough to handle “mountains of data” in a “reasonable timeframe.” And the third is advancements in AI algorithms, allowing for non-sequential or parallel processing of large data pools.

LLMs can do five things with language: generate, summarize, translate, instruct or chat. With a combination of “these modalities and actions, you can build applications” to solve any problem, Chockalingam said.

Enterprises are tapping LLMs to “drive innovation,” “develop new customer experiences,” and gain a “competitive advantage.” They’re also exploring what safe deployment of those models looks like, aiming to achieve responsible development, trustworthiness and repeatability.

New techniques like retrieval augmented generation (RAG) could boost LLM development. RAG involves feeding models with up-to-date “data sources or third-party APIs” to achieve “more appropriate responses” — granting them current context so that they can “generate better” answers.

Chockalingam encourages those interested in LLMs to “get your hands dirty and get started” — whether that means using popular applications like ChatGPT or playing with pretrained models in the NVIDIA NGC catalog.

NVIDIA offers a full-stack computing platform for developers and enterprises experimenting with LLMs, with an ecosystem of over 4 million developers and 1,600 generative AI organizations. To learn more, register for LLM Developer Day on Nov. 17 to hear from NVIDIA experts about how best to develop applications.

More from NVIDIA AI Podcast

All 115 episodes
NVIDIA’s Annamalai Chockalingam on the Rise of LLMs - Ep. 206NVIDIA AI Podcast · 39 min
Listen in VO