In short
NVIDIA’s open foundation model effort (Nemotron) and why NVIDIA is “giving away” AI models, framed as efficiency at the compute/power/economic limit and as ecosystem support for agentic workflows. The episode also compares open vs closed AI, discusses distillation, and explains Nemotron 3 Ultra’s technical design.
Guest backgrounds
Bryan Catanzaro is an NVIDIA leader who returned to build applied research labs after earlier work at NVIDIA (including GPU/AI systems) and a stint at Baidu’s Silicon Valley AI lab with Andrew Ng and Dario Amodei. He helped develop Megatron (large transformer training) and previously worked on GPU deep learning systems.
Key claims
Open AI technologies enable broader innovation like the Internet; the open/closed gap is less important than rapid overall progress. Open models help companies customize while managing sensitive data/guardrails. NVIDIA builds Nemotron to co-design future AI systems (for NVIDIA’s acceleration business) and to support the deployment ecosystem. Distillation isn’t the only driver of progress; China’s openness has helped the global ecosystem.
Notable examples
Nemotron 3 Ultra (4-bit pretraining, hybrid Transformer+Mamba/SSM, MoE, 1M-token context, multi-token prediction, multi-teacher distillation) and the Nemotron Coalition; NVIDIA’s DLSS as an example of AI efficiency.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Need for Efficiency in AI Development
0:00 to 0:32
Exploring how to gain intelligence through efficiency rather than force.
“If you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient.”
Open Source AI Trends and Developments
0:48 to 2:10
Discussing the current state of open source AI and recent advancements.
“We begin this conversation with the state of open source AI and the race between the US and China.”
The Importance of Open Technologies
2:10 to 3:21
Bryan Catanzaro discusses how open technologies are vital for AI innovation.
“Well, it's really exciting to see all of the energy going into open technologies for AI, because we know that open technologies make it possible for people to innovate.”
Drivers of Progress in Open Source AI
3:21 to 4:19
Exploring the various factors that propel open source AI development forward.
“And what's your sense for how far behind open source is compared to closed source?”
Community Contributions to AI Progress
4:19 to 7:12
Examining how the AI community collectively contributes to advancements.
“Is that the communities and big companies like NVIDIA being behind it?”
Perspectives on Chinese AI Development
7:12 to 9:50
Bryan shares insights from his experience in China and the global AI landscape.
“And, you know, the community, of course, cares deeply about this technology.”
The Case for Open Source AI in Business
9:50 to 12:35
Discussing the fundamental advantages of using open source models for companies.
“You know, I was really excited when OpenAI released the GPT OSS models a while back.”
Deep Dive into Bryan's Background
12:35 to 13:20
Bryan Catanzaro recounts his journey in AI and his time at NVIDIA.
“I'd love to go into a bit of a deep dive into Nemotron.”
Journey Through AI and NVIDIA
14:00 to 17:45
Explore the evolution of AI applications at NVIDIA and Bryan's experiences.
“So a GPU is a thing that we make in order to accelerate the world's most important computations, which in 1995 was graphics.”
Real-Time AI in Graphics: The DLSS Project
17:46 to 19:59
Learn about the development and impact of DLSS in gaming.
“So 10 years ago, actually, in 2016, Jensen called me up and said, hey, would you like to come back and build an applied research lab?”
Show all 33 chapters
The Megatron Project: Advancing Language Models
20:00 to 22:08
Discover the Megatron project's role in training large transformer models.
“And this was back in 2017, before Transformers were big and before language modeling started taking over the world.”
Nematron: NVIDIA's Future-Oriented AI
22:09 to 24:42
Understand the purpose and goals of NVIDIA's Nematron initiative.
“very significant efforts into creating its own family of frontier models?”
The Evolution of Nematron Models
24:43 to 28:00
Review the key releases and developments in the Nematron series.
“And these days, that is absolutely not the case.”
Development of Nemotron Models
28:00 to 30:28
Learn about the evolution and iterations of the Nemotron models by NVIDIA.
“And then, yes, and then we continued to develop that.”
The Nemotron Coalition Explained
30:28 to 33:14
Explore the purpose and structure of the Nemotron Coalition and its collaborative approach.
“And NVIDIA is a company that follows through.”
Overview of Nemotron Family Models
33:14 to 35:48
Understand the different models within the Nemotron family and their use cases.
“So that's the idea of the NemoTron Coalition.”
4-Bit Arithmetic in Model Training
35:48 to 39:25
Discover the significance of using 4-bit arithmetic in training AI models.
“so that your model can converge to an excellent result.”
Hybrid Architecture of Nemotron
39:25 to 42:00
Learn about the hybrid architecture combining transformers and state space models in Nemotron.
“So as we get into slightly more technical things, the architecture of NemoTron is hybrid.”
Understanding Mixture of Experts (MOE)
42:00 to 44:35
Learn about the MOE architecture and its implications for AI models.
“which then means that generally you can fit much higher batches on the GPU when you're training and doing inference because the memory requirement is lower.”
Innovations in Latent MOE
44:35 to 46:35
Explore the concept of Latent MOE and its impact on AI model efficiency.
“So with Blackwell, for example, NVIDIA went all in on MOEs.”
Context Length in Language Models
46:35 to 48:41
Discover the importance of context length and its challenges in AI.
“So you could think about it as like, you know, our library of books got four times bigger, and we get to, you know, read four times more books at the same inference cost because of this particular innovation.”
Multi-Token Prediction Explained
48:41 to 51:26
Understand how multi-token prediction can enhance model performance and speed.
“And there's this whole separate discussion around context compaction to make sure that the model doesn't get lost in too many tokens.”
Collaboration in AI Development
51:26 to 55:05
Examine the balance between technology and human organization in AI model training.
“So it doesn't degrade your accuracy at all to turn on multi-token prediction, but it can give you a speed up and it's probabilistic depending on the acceptance rate of your predictor.”
Post-Training and Data Sourcing
55:05 to 56:00
Learn about data sourcing for post-training in reinforcement learning contexts.
“So it's just as much a technology question as a human organization question.”
Data Acquisition Strategies at NVIDIA
56:00 to 58:00
Learn about how NVIDIA sources and generates data for AI models.
“The next big question is, can they become great at law and consulting and then all sorts of different domains?”
The Future of AI Across Domains
58:00 to 1:00:10
Explore the potential for AI to expand into new fields beyond coding.
“Since we're talking about post-training and RL in different domains, just so curious to get your thoughts on where we go from here in terms of generalization.”
Organization of NVIDIA's Research Teams
1:00:10 to 1:03:40
Discover how NVIDIA structures its diverse research teams to foster collaboration.
“And I think that's going to become significantly more complex and diverse over the next few years.”
Challenges in GPU Allocation
1:03:40 to 1:06:52
Understand the complexities of GPU resource allocation at NVIDIA.
“I think organizations that figure out how to collaborate to build AI succeed.”
Bootstrapping Research Ideas
1:06:52 to 1:10:02
Learn how NVIDIA encourages innovation through iterative research processes.
“How do you balance useful research with great exploratory research?”
The Balance of Innovation at NVIDIA
1:10:02 to 1:12:07
Explore how NVIDIA balances top-down strategy with bottom-up innovation.
“big investment and if we can figure this out it will be significant for our company and then we We let the people who are interested in that work on that.”
The Multifaceted Nature of Intelligence
1:12:07 to 1:13:43
Understand why intelligence is not just about raw capability but contextually driven.
“You work at a company, it's one company.”
The Future of AI and Human Adaptation
1:13:43 to 1:17:46
Discuss the implications of AI as an 'external brain' and its societal impact.
“would it find the next CEO by looking for somebody who won the International Math Olympiad?”
Navigating AI Safety and Open Technologies
1:17:46 to 1:22:27
Dive into the significance of open technology for AI safety and societal consequences.
“and I think ultimately this is going to make our lives better.”
Transcript
Automatic transcript. May contain errors.0:00If you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. We build tools. We build external organs that help us solve problems. You know, we have an external stomach. We call it kitchen. Now we're creating an external brain. What is the implications of an external brain? pretty profound. Nobody actually really knows. Hi, I'm Matt Turk. Welcome back to the Matt Podcast.
0:34Open source AI is having yet another moment with powerful new models arriving almost weekly. And my guest today is one of the very best people to unpack it all. Brian Catanzaro leads Nemotron, NVIDIA's family of open foundation models. Now, not everyone realizes NVIDIA has a massive effort to build frontier AI models, but it employs hundreds of AI researchers and Nemotron 3 Ultra immediately became the number one US open weights model when it was released just a couple weeks ago. We begin this conversation with the state of open source AI and the race between the US and China. And then we go deep inside NemoTron.
1:094-bit training, hybrid member transform architecture, mixture of experts, multi-token prediction, and multi-teacher distillation all in plain language. And finally, we get a rare look at how a modern AI research organization actually runs, how you get many brilliant minds to build one model instead of 100 papers. Please enjoy this awesome conversation with Brian Catantaro. All right, Brian, excited to do this. It seems that open source is having a banner year. So you guys at NVIDIA just released Nemotron 3 Ultra, which is an important moment and the best open source, open weight model in the US.
1:49That was just a few days ago. And then even more recently, GLM 5.2 came out and that was another moment. So it seems that things are accelerating in open source AI. It feels like a great place to start. What's your assessment about where we are and how wide the gap between closed source and open source currently is? Well, it's really exciting to see all of the energy going into open technologies for AI, because we know that open technologies make it possible for people to innovate. You know, the Internet is such a great example of that. We actually did have closed Internet. I don't know if you remember things like America Online and Prodigy back in the day.
2:31And they were great. And open Internet has also been amazing, right? Like so many different companies have been able to figure out how to transform their work thanks to an open technology. The application of the Internet to retail is very different from the application of the Internet to health care or manufacturing. But all of them have been totally transformed by the Internet. AI, I believe, is also a very transformational technology and also a technology that needs to be applied in very diverse ways. And because of that, I believe that open technologies for AI are really fundamental. And it's very exciting to see continued investment and development of open technologies for AI from so many different organizations around the world.
3:20And, you know, I hope that that continues. And what's your sense for how far behind open source is compared to closed source? So it's been the big trend of the last few years has been this sort of narrowing gap. Do you think that open source is almost there or the bar keeps getting raised by the closed source models? Well, I feel like this question, it's maybe a tempting question because, you know, it's fun to set up kind of competition. But I actually feel like the whole AI community is moving very fast. And if you look, for example, at the progress in AI, whether it's closed or open, just over the past three months, it's been incredible.
4:04And so if you're in a field that's moving really, really fast, I think that's more important than any particular gaps that might exist between different models. Because the most important thing is, you know, how is AI developing as a field? What do you think the drivers are to continue to progress in open source AI? Is that the communities and big companies like NVIDIA being behind it? Is that the global competition with China? What propels open source AI forward? You know, I think there's a number of things that are pushing open technologies for AI forward. One is just the demand. You know, there's so many organizations that want to customize AI and want to integrate it deeply into their work in a way that really requires open technologies for AI.
4:50And so I think the demand is certainly there. I think also it's just the best way to develop technology. And we've seen this, you know, for many decades that technologies developed in the open move quicker because we can all learn from each other. And in an era where we're undergoing the most exciting thing to happen in technology in our lifetimes with the development and the deployment of AI, what else do computer scientists want to work on other than making AI awesome? And if working together as a community is the best way to do that, then that's also a driver that pushes the community towards openly developing technology.
5:29To ask maybe a slightly cynical question, there is at least a part of the community that's wondering whether open source as an ecosystem, not NVIDIA, but in general, has been progressing in part based on the ability to distill closed source models. and in a world where we're seeing the Anthropics and Fable 5s of the world starting to discourage distillation, do you think there is a chance that open source AI progress may slow down in that context or as a result? You know, in my mind, there's no question that when the technology community decides to make huge investments in the most transformational technology of our time, that there's going to be rapid progress.
6:23And also that that technology is not going to be controlled by a small group of people because that's just not the way that the industry works. You know, we do our best work. We have the most impact with our work when we're able to each think about it in our own way and apply it in our own way. So, you know, I love the closed AI APIs, whether from Anthropic or other people. I think they're amazing. You know, I'm really, really impressed with the work that those labs are doing, but they're not the only labs in the world. There's lots of labs around the world and lots of people have a good idea.
7:01It's not the case that there's only a few labs that have the monopoly on all good ideas. That's just not true. That's not how humanity operates. There's a lot of bright people on this planet. And, you know, the community, of course, cares deeply about this technology. It's obviously so transformational, has such profound impacts on so many things that, of course, many people want to be involved in that. And so I think over time, we're going to see that community-oriented approaches to developing and deploying AI are going to continue to strengthen and be widely adopted, because that's really the history of how we build things as a human species.
7:42Do you think that is globally true as well? So, you know, in particular, with respect to China, this perception that, yes, a lot of people have great ideas around the world. However, a lot of progress from Chinese models were directly inspired or perhaps generated through distillation from the closed source models. Is that just kind of like press rage bait or from the perspective of a leading AI researcher, you're very impressed by the novel ideas that come out of China as well? You know, perhaps unusually, I actually did work at a Chinese company for about two and a half years. I worked at Baidu.
8:30I worked in the Silicon Valley AI lab, along with Andrew Ng, as well as Dario Amadei. And we all worked for a Chinese company and saw how smart, hardworking, creative, inventive our colleagues were at the rest of Baidu. And, you know, that experience has stuck with me. I think it's absolutely false to say that, you know, the achievements of some other country are all being created by sort of, you know, copycat mentality. It's just not it's just not true. Now, do we all learn from each other in the technology community? Of course, you know, of course, of course, we learn from each other. But, you know, I would say, you know, it's been a really good thing for the world that the Chinese AI community has been so open with what they've been building.
9:28I think it's enabled a tremendous number of companies to build things that they couldn't have done without that community. And I think it's also spurred technological progress throughout the AI ecosystem. So, you know, I'm really grateful for the contributions that our colleagues in China have made over the years. And, you know, I would love to encourage a spirit of openness amongst AI labs around the world outside of China as well. You know, I was really excited when OpenAI released the GPT OSS models a while back. And then, of course, Google's been doing great work with Gemma. Absolutely thrilling to see that.
10:10And, you know, we're pushing Nemotron along here at NVIDIA as well. So I think there's a chance for the rest of the world to catch up to China in the sense that, you know, We can understand the benefits of working together as a community to build technologies for AI in a way that I think China has frankly been leading. What is the case for a customer to be using open source models these days? What is your fundamental advantage? Every company is built around a secret. This is a secret that has to do with not just their intellectual property, but also their platform. which has to do with how do they interact with problems and customers?
10:57How do they think about solutions to what their customers need? And it is always the case that the value of AI is greater when it can be more tightly connected with those secrets because AI depends on data critically. So the more valuable the data that goes in, the more valuable the solution becomes. Now, every company, when it's thinking about how to deploy AI, has to think through what are the implications for the core secrets of our company. And there's a lot of circumstances where due to trade secrets or trying to think through the business model or even regulatory requirements that there's data that you really have to treat very carefully by law.
11:43And it is much better to do that when you are able to think that through and implement it yourself, thinking about the integration of AI, the way that AI interacts with customers, the guardrails that are put in place. Every company has a specific understanding of its customers and therefore what the customer needs. And the amazing thing about open technologies for AI is that they allow customization, right? So companies can think this through. They can build things that really matter for them. And, you know, I started out this conversation talking about the Internet and about how the Internet, the deployment of the Internet has been done in very different ways for very different industries.
12:26And there's a lot of desire to do that as we see AI change the way that we work and play throughout the entire economy. This is really spurring a lot of demand for open technologies for AI. Great. I'd love to go into a bit of a deep dive into Nemotron. But before we do that, maybe a few minutes on your story, your background. What was your path to where you are today, including the Baidu detour? So I started work at NVIDIA in 2008. At the time, I was a graduate student trying to figure out parallel computing for artificial intelligence. And I thought NVIDIA had a chance of changing the way computers work with AI.
13:13Which was presumably a lonely quest right in 2008. Oh, it was very chaotic. Back then, people thought I was crazy. And I remember going to ICML in 2008. I published my first paper, Training Models on the GPU, and people asked me why I was there. People said, this is not a good paper for ICML. We just do fancy math here. And I was like, well, but I think computing actually matters a lot for AI. If we could train bigger models that had more capacity to learn, we could probably solve more problems. And they kind of nodded their heads at the door. Like, well, I'm not really sure why you're here. Isn't a GPU a thing for gaming as well?
13:51Oh, right. Yeah, there's also that, right? Which we continue to run into that idea. Actually, a GPU is whatever NVIDIA says it is. You know, we make them. So a GPU is a thing that we make in order to accelerate the world's most important computations, which in 1995 was graphics. And, you know, for a long time now, it's been AI. So anyway, I started at NVIDIA. I was in the research group doing strange things about trying to make compilers, libraries for AI on the GPU. That led to the creation of first Copperhead, which was a Python embedded language that compiled to the GPU, which I think foreshadowed a lot of things in TensorFlow and PyTorch.
14:40And then that led to the creation of QDNN, which was NVIDIA's first product for deep learning on the GPU. And I really enjoyed working on that. But I was always wanting to see more firsthand about the applications of AI. And at NVIDIA, I was mostly working on libraries and compilers for AI. So I thought, well, you know, when Andrew Ng asked me to go build the Silicon Valley AI lab with him at Baidu, I thought, oh, this is a great opportunity because even back then, Baidu was very advanced in its application of AI to its core business. And so that was a fantastic opportunity for me. The Baidu Silicon Valley AI lab was an amazing place full of brilliant people that were working really hard.
15:32What was it like working with a young Dario? Was there any signs that he could become who he has become? Dario was brilliant from the beginning. I remember I interviewed him. I was on the panel. And at the time, he had been working in bioinformatics, so he hadn't been working on deep learning or the things that we call AI these days. But it was very clear that he learned extremely quickly and also that he thought extremely deeply. I think, you know, the thing I admire most about Dario is the strength of his conviction. You know, I've been working in this field for a long time and I've believed also that AI is going to transform the world.
16:21But I don't think that I believed in it as completely as Dario did. And perhaps that was because, you know, my academic training during my PhD was full of a lot of caution. I don't know if you remember, but AI was old and bad in 2005. It will never work. That people did with computers. They started doing it in 1945, right? And so there had been so many grandiose promises that failed to deliver over the years. And so I came to AI with a lot of caution. In fact, back then we used to call it machine learning, which was basically a dodge. Like we just didn't want people to know that we were we were working on AI because then they would be like, oh, we've heard about that.
17:02It never works. Right. So I came to AI with a little bit of this like, you know, academic caution, like, oh, we should, you know, we should hedge a little bit. Like, I don't know if now's the time. And Dario, you know, his strength of conviction and his understanding of the moment of how the technology was developing. This time it was actually going to work. And then the implications of that on, you know, how the technology should be developed, what kind of institutions to build. I think he's done a spectacular job. And so, yeah, working with him, it was always a fun experience. So then you went back to NVIDIA and walk us through the journey.
17:45Yeah. So 10 years ago, actually, in 2016, Jensen called me up and said, hey, would you like to come back and build an applied research lab? And I thought that would be a fantastic opportunity. You know, I've always loved NVIDIA. I've loved the way the company works, the convictions the company holds. You know, NVIDIA is a very unique company. It follows through over long time periods. You know, and I've seen that with CUDA. I've seen it with our deep learning technologies. I've seen it with our ray tracing graphics technologies, our AI for graphics. You know, over and over again, NVIDIA is not afraid to put in five or 10 years worth of research in order to change the world.
18:24You know, and working at a company that has that strength of conviction and the ability to follow through is kind of an ideal thing for me. I just really I just really love the support that the company gives gives its researchers to invent the future. And so I thought I'd come back. The first project that that I worked on actually became DLSS, which some of your audience may know about. But DLSS is our real-time AI for graphics, and it makes a small GPU run like a big GPU. It's about 10 times more efficient because rather than computing the color of every pixel for every frame, we use AI to infer the color.
19:08And, you know, these days, 23 out of every 24 pixels is being generated by our AI model when you're using DLSS to play games. and gamers love it. It's become the standard way of playing games because it's just so much more responsive and it's more beautiful. Our AI, we train it offline on huge data sets and it's able to render graphics in real time more beautifully than traditional methods do. We recently actually announced DLSS 5, which is a fully generative version of DLSS and I am so excited about it. it represents a culmination of 10 years worth of research on how to make real-time graphics much more beautiful.
19:50And so that's part of the journey here for me was real-time AI for graphics. But then at the same time, we also started a language modeling project. And this was back in 2017, before Transformers were big and before language modeling started taking over the world. But, you know, I just had this intuition maybe built on, you know, some of the things that I had seen while working at Baidu. I just had this intuition that, you know, working with text and understanding text was going to lead to better reasoning, which was going to lead to better application of AI in all sorts of domains. And so we started this project called Megatron.
20:35Megatron stands for the biggest, baddest transformer. That's why we named it that. And it was really a systems project to show the world how to train the largest transformer models on NVIDIA's hardware. Back at the time, some of your audience may or may not remember this, but there were being claims made that the only way to train big transformer models was on the TPU. Because after all, the transformer had been invented at Google. And so, you know, we looked at, you know, we loved the transformer paper. We thought, wow, this has amazing potential. We tried it out on our own language modeling tasks, and it worked so much better than the RNNs that we had been using before.
21:13And also, we saw immediately that there was an enormous systems opportunity to co-optimize the GPU, the networking, all of the compilers and software that would enable people to scale transformer-based language models really dramatically. And we thought, you know, this is something that could really have an impact. So we started the Megatron project, which then led to, I think, basically helping the whole industry figure out how to train extremely large LLMs and also led to the foundations of today's Nemotron project, where, you know, NVIDIA trains its own LLMs for its own purposes. So that's kind of the history.
21:54Great journey. Okay, so let's go into all things Nemotron. And before we get into the specifics, that's the obvious question that I'm sure you've been asked many times, which is why does NVIDIA care in the first place to be building model and investing very significant efforts into creating its own family of frontier models? You know, Nemotron has two jobs. The first job is to help us understand how to build the systems of the future. NVIDIA is an accelerated computing company, and that means thinking through the world's most important computational challenges from first principles and designing systems, which includes a lot of software, in order to make it possible for people to invent and deploy things that never could have been done with standard computing.
22:44But in order to do that, NVIDIA has to deeply understand everything about how AI works. That's how we co-design all of the systems and software for our main product line. So the first job of Nemotron is to make sure that NVIDIA continues to exist so that we can continue delivering meaningful acceleration in an era where Moore's law has died. And the acceleration that we get these days comes through specialization. But again, specialization comes through understanding. So that's Nematron's first job is to help NVIDIA understand how to build its core products. Nematron's second job is to support the ecosystem.
23:25One of the most valuable things that NVIDIA has built over the years is all of the people around the world who build and deploy amazing AI using NVIDIA's technologies. And we think that it's necessary for open technology for AI to continue to exist from NVIDIA to help support that. Nemotron's not trying to be the only open technology for AI. We love all technology for AI for the very straightforward reason that whenever AI is further developed and further deployed, it's an opportunity for our business. So we're very explicitly trying to develop our ecosystem because that's good business for us.
24:09But we're not trying to be the only provider of technologies for this ecosystem. We love seeing other companies contribute as well. The most important thing for Nemotron's second job is just making sure that it continues to be possible for companies of all shapes and sizes to build and deploy their own AI. By the way, Moore's law is dead. Is that official? It's been dead for years. It's been dead for years? Why is that? Well, you just look at the progress in semiconductor manufacturing. You know, the original statement of Moore's law was economic, right? It was about we can afford to put twice as many transistors on the same chip in every whatever 24 months, whatever the time period is.
24:51And these days, that is absolutely not the case. it hasn't been for probably five or 10 years, right? Now we are still scaling our systems, right? Through a number of ways. One is just applying a lot more silicon to it, right? We are also getting, transistors are continuing to get smaller and more efficient, although at a slower pace, but they're also getting quite a bit more expensive at the same time. So the, you know, In an era where Moore's law was alive, the best way to make the system of the future was to take the system of the present and then just shrink it and maybe double it at the same time.
25:30But in an era where we've been living for a while now, where you don't get economic benefits from taking your existing design and shrinking it, you really have to be more clever about how you use every part of the system. that that's you know an era where accelerated computing is is much more valuable than ever because the the work of thinking through the problem from first principles and co-designing absolutely everything from transistors to algorithms and applications in order to reduce waste and and deliver meaningful acceleration that's more valuable than ever fantastic to playback what you were saying a minute earlier, it makes good business sense for NVIDIA to be in the model business because one, it helps design better chips and two, whatever is good for AI is ultimately good for NVIDIA, which makes a lot of sense.
26:26That NemoTron effort is reasonably recent, right? It started in 2023, I believe. Maybe walk us quickly through the key releases. I believe in 2023 that was pneumotron 3 8b as a key release or am i missing a step yes yes yes so you know the the you know the the numbering is somewhat lost to time it i almost feel like we're in the lord of the rings and it's like you know there's like some ancient like relics that we're digging up out of an old mine um you know this is a long time ago you know the original what what uh we originally called Pneumotron 1 was actually a project that we did with Microsoft.
27:04We jointly trained a 530 billion parameter model. I believe that was released in 2021. And so this is GPT-3 era. And that's what, at the time, we called it Megatron Turing NLG. Turing was what Microsoft was calling their language model efforts at the time. But that, in retrospect, we called Pneumotron Then along the way, we built a few more. We got up to Nemotron 3. And then Lama came along, and we were really excited about that. We were very happy that Meta was supporting the open AI technology space. And so we started, you know, taking our language model technology and adding it to LAMA models, which then resulted in LAMA Nemotron 1.
27:58And, you know, that was the first reasoning model built on LAMA. We were really proud of that. And that was 2025? Might have been 24, I believe. I can't remember. Somewhere around there. And then, yes, and then we continued to develop that. And, you know, last, so the numbers kind of started over again. We released a Nemotron 2, I believe it was last year. And then we quickly followed that up with Nemotron 3 because we needed to put MOE support in. Nemotron 2 didn't have MOE support, and that made it kind of uncompetitive against other models like GPT-OSS-20B was just so fast because of MOE. And so we were like, okay, we've got to put the MOE in.
28:54So that became Nemotron 3. Now we're in a slightly difficult state because we're working on Nemotron 4, right? But we already released a Nemotron 4, which was in 2024, we released a 340B model called Nemotron 4. And so I'm not exactly sure how we're going to solve this marketing problem. I didn't create this marketing problem. So I'll do my best to make it clear that Nemotron 4 of whatever, whenever we release that is different from the 2024 Nemotron 4. But in any case, we've been working on this for a long time. I think more important to us than any particular generation is just the sustained commitment that NVIDIA has to developing these models.
Read the full transcript
29:40We've been doing it for a while. I think our models have gotten dramatically more useful in the past year, which is a reflection of two things. Primarily, one is that the whole company has come together. So there are many different teams around NVIDIA that now understand how important this is to NVIDIA's future. And so there's dramatically more people and better ideas that are going into Nemotron. And then number two, along with that, we've been able to scale the compute resources that go into it. Obviously, it's very important to have good computing infrastructure to build AI. We've recently increased our investment substantially because we believe that this is really, really key to our company's future.
30:22Fascinating. But just to continue the thought, I think it's really important that everybody knows that we've been doing this for a long time. We are increasing our investments substantially. And NVIDIA is a company that follows through. You know, we followed through over 10 plus years with CUDA, and we're doing that with Nemotron now. That's really helpful because I think the broader world is just starting to catch up to the fact that there is a very substantial open source frontier AI research effort that's been happening. So it's really interesting to hear that, you know, there's been this progression and now there's this family of models that we're going to talk about in a second.
31:00Another important moment seems to be the creation just in March, three months ago, of the Nemotron Coalition. Do you want to explain briefly what that is? So Nemo Tron exists to help support the ecosystem. And we were thinking, well, this is a different kind of AI project than other projects around the industry, right? Because we're not actually trying to dominate in any way. We're just trying to support. We're not trying to control the way that AI is being integrated into all these companies. We're just trying to make sure there's good AI. But we thought, well, maybe if we worked with people while we develop it, then it's going to be more useful for them.
31:46It'll be easier to integrate because we will consider what they need from the beginning. And, you know, Nemotron has always been collaborative. I was telling you that, you know, long, long time ago, our first big model that we trained, we did with Microsoft. Right. It was a joint effort where NVIDIA and Microsoft researchers worked side by side to build that. And that ended up, I think, helping both NVIDIA and Microsoft. I think we both learned a lot from that experience. And so because Nemotron is not trying to compete with other companies, but rather support, because we're going to be putting it out there openly anyway, why not collaborate before the thing is built, rather than Nemotron being a project that NVIDIA does all on its own and then posts on the Internet and says, hey, why don't you try this?
32:33We think it might be good. Why don't we make sure that it's good for the partners that are interested by working with them before Nemotron is even created and incorporating any sort of feedback, evaluations, environments, benchmarks, or any other kinds of technology that other people want to bring. It turns out that the entire ecosystem, there's a lot of companies that really want open models to succeed. And so they They have a self-interest. They have their own vested self-interest in making sure that open technologies are excellent. And so why not work with them and let them contribute however they'd like to making NemoTron better?
33:14So that's the idea of the NemoTron Coalition. It is not an exclusive coalition. We're not trying to be the only model out there. All the companies that we work with are free to continue doing the work however makes sense to them. And yet, you know, these companies want to work with us because they want to make sure that open technologies for AI keep developing quickly and that they have a chance to influence how that happens. Great. What's the current state of the Neumotron family? You got Nano, you got Super, you got Ultra. What do those models do and what are the use cases for them? So Nano is a 30 billion total, 3 billion active parameter model.
33:55Super is 120 and 12 and Ultra is 550 and 55. They're designed really to fit, it's kind of small, medium and large deployment scenarios. Nano can be really capable for things that don't require nearly as much knowledge or reasoning, but obviously for the most capable model, you go for Ultra. Super, in a lot of ways, is our most popular model because it represents kind of a great balance between cost and intelligence. So we kind of like having this small, medium and large approach to building a family just because our customers seem to respond to that pretty well. But, you know, the most important thing from NVIDIA's point of view that people are doing with LLMs is agents, right, is building agentic workflows.
34:54Having an agent working on your behalf, solving problems for you night and day is such an exciting way of approaching the problems that we have to solve. And it's our dream to make Nemotron amazing for that purpose. That's our goal. to double click on this at a hell of all nemotron families focused on urgentic reasoning with a particular focus on making it efficient is that is that the right headline that's right yeah um nemotron has always been um uh speed first approach to building models because nvidia is an accelerated computing company as i was saying we're trying to think through what is the problem here computationally from first principles and um you know nemotron 3 family has a lot of things in it that we're really proud of.
35:41For example, Nemotron Ultra and Super were pre-trained using 4-bit arithmetic. We pre-trained those in MVFP4, which is a not trivial thing to do, to invent the algorithm so that your model can converge to an excellent result. Using such coarse arithmetic required a lot of invention. Really proud of that. Do you want to explain maybe for people what 4-bit is versus 16-bit, for example? You know, actually, there was a fantastic post I saw on Hacker News yesterday where somebody let you upload a picture and then it would basically posterize it, basically reduce the colors to fit different number formats, including NVFP4 and MXFP8 and some of the other formats that are out there.
36:24And so you could kind of swipe around and look what it does to the colors of a picture. And, you know, it's really quite dramatic. Four bits is not a lot of bits, right? That's only 16 values. Now, of course, these are all what are called block scaled formats. So groups of numbers also come with an 8-bit scaling factor. And the specifics of this can get rather complicated. So maybe they're not quite as important. But the reason why we want to do this is because, first of all, we have dramatically higher throughput for these formats in our GPUs, specifically on Blackwell Ultra. And secondly, we know that it's going to save an enormous amount of energy.
37:10One way to think about the computational problem of AI is that we are going to be running at the limit, whatever the limit is. It could be an economic limit, like we only have so many billion to buy servers with. It could be a power limit. We only have so many gigawatts that we can afford to train a model with. Whatever the limit is, we're going to be running at that limit. Every organization is... Why? Because the value of intelligence is so high that people are going to invest because they know that they're going to get return. The value of intelligence is enormous. So if you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient.
37:59We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. And, you know, 4-bit number formats are dramatically cheaper to move around. They take up less space in memory. They take up less picojoules when you move them from the memory or even on the chip, around the chip, much less energy when you compute on them. And so that's really driving the investment in 4-bit formats. And I think these days, 4-bit formats for deployment are very well established. It's pretty straightforward these days to make a good quantized 4-bit checkpoint that you can deploy and that gets you a lot of inference cost and speed advantages.
38:50But using 4-bit formats for pre-training, that's quite a bit more challenging because you have this numeric solver that's optimizing the weights and it can be quite sensitive. So if you don't treat the numbers right, your model can diverge. And instead of actually getting a model done through pre-training, you end up with, you know, basically just that run diverged, which is, you know, always scary. So it took a lot of invention for us to be able to pre-train Nemotron in 4-bit. We're really proud of that. Okay, great. All right. So as we get into slightly more technical things, the architecture of NemoTron is hybrid.
39:36Is that right? So it's a combination of transformer and Mamba state space, which is a slightly more exotic form of architecture. Walk us through that. Yeah, you know, we published a paper in 2024 that showed that you actually get a smarter model by combining state space models with transformers. And we actually did a sweep of, you know, how much of the model should be full attention and how much of it should be a state space model in order to get the lowest perplexity, basically the best language model that you could get. And we found that you actually want it to be mostly a states-based model with a little bit of attention.
40:18And kind of the intuition behind that is that the states-based models seem to be better at kind of this intuitive kind of impressionistic understanding of a sequence. Because they're kind of summarizing the entire sequence into a constant space. That's how they work. So instead of having the ability to look at the entire sequence randomly, they summarize everything at every step into a constant cache or little scratch pad that they're working on. And that constraint seems to actually make them smarter at some tasks that involve global understanding. On the other hand, the advantage of full attention is that it can pick out very specific bits of information and look at those exactly.
41:04It doesn't lose anything. There's no lossy compression going on. and you can actually see the whole thing. And so we found that using both of these together was actually better than using either one on their own. And that is independent of the speed benefit. That is just the model is smarter. And since we published that, I think a lot of other labs have also found this to be true. A lot of models these days are being built with hybrid SSM approaches, For example, QN has done that. Kimi is using what they call Kimi linear attention these days. So it's become, I think, quite widely adopted to use some sort of state space model in conjunction with full attention for the base architecture.
41:55Now, it also has some speed benefits because the amount of memory that you need to hold that state space cache is actually constant with respect to your sequence length, which then means that generally you can fit much higher batches on the GPU when you're training and doing inference because the memory requirement is lower. And it keeps the GPU fuller and busier and therefore, you know, provides some pretty important efficiency benefits as well. So the models are also based on an MOE, mixture of experts, architecture. Walk us through that and maybe remind people what MOE is in the first place.
42:42So mixture of experts is a form of sparsity. The idea is, wow, you want to train a model on the entire Internet. You want it to remember absolutely everything about the history of everything. But when you're answering a particular question, does it seem reasonable that it needs to actually think about the entire universe in order to answer that question? Actually, no. It seems like it's quite sparse, right? It seems like we're using a language model to explore a very tiny space of ideas in order to answer a question or solve a problem. We want the model to be able to draw from the entire universe.
43:16We want to train it so that it understands everything that it possibly can. But when it's actually running, it doesn't really need to see all of that information. There's been a variety of approaches to sparsity that try to take advantage of this property, but mixture of experts has been the most successful. And the way that it works is that the neural network has what's called a router that is learned that is going to decide to send activations to a subset of the experts for every token that's flowing through every layer of the model. It's going to be making choices about which fraction of the model is going to actually get to interact with this token as we try to understand it, build up representations of the problem, and then generate the next token that we're going to output.
43:59So it's a little bit like if I have a company with 550 employees, but 55 of them are in engineering, engineering i want the 55 employees who are specialists to come to my meeting about engineering and not the rest of the company that's right yeah or you can think about it as a library like if you go into a library to do research you don't read all of the books in the library like your first job is to figure out which books do you need to look at in order to find the answer to your question and so um so that's kind of the the idea behind moes now moes have fascinating implications for the systems that we built.
44:37So with Blackwell, for example, NVIDIA went all in on MOEs. That's why we built NBL 72, which allows up to 72 of our GPUs to read and write each other's memory at very high speeds, very low latency. Now, why is that important? It's because as you put a token through the stack of layers, at every layer, you have a router that's routing that token somewhere else. Why don't you partition your experts so that the experts are not sitting, every expert on every GPU, but you have a subset of the experts assigned to each GPU. And then you're routing the tokens between the GPUs very dynamically as you push the token through the network.
45:14Now, this is impossible to predict in advance where the tokens need to go because it's very specific to that particular token for that particular model. And so that's why we built NVL72. And that's why Blackwell is so amazing for inference for today's AI models is because we thought deeply about a mixture of experts when we were building it. And this is speaking to Nematron's first job. If we hadn't been working on understanding AI, we wouldn't have been able to build Blackwell properly. And that has translated directly into increased deployment of Blackwell, which we're very excited about. Is what you just described called latent MOE, or is that a different concept?
45:59Latent MOE is a specific innovation that we have in Nemotron 3 family. And what it does is actually reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So, you know, every token produces a vector. And the idea is like, we're going to take that vector and learn a way to compress it and then send that compressed thing through the network. And then we're going to uncompress it at the other end. And as a result, we save on network bandwidth, and we also get four times the number of experts for the same inference cost. So you could think about it as like, you know, our library of books got four times bigger, and we get to, you know, read four times more books at the same inference cost because of this particular innovation.
46:48Is MOE in general becoming the default architecture for Frontier AI? Yeah, I believe MOEs have been the default in Frontier AI for a long time. They're just a really good combination of inference cost and intelligence. Great, great. But they have drawbacks as well. They take a lot more memory. If you have a very small amount of memory, a dense model is going to be smarter. And they also, they tend to work best either if you're running at batch size one, so you're running basically a single job, or you're running a huge data center with like infinite queries coming in in the middle they can be a little bit tricky another important characteristic of numatron 3 ultra is a 1 million token context the the long context window how important is that in the overall mix and what does it enable the model to do the longer the context length the more challenging problems we can solve with a language model that allows us to do things like append all sorts of information to a query, which could be a code base.
47:54It could be instructions. In the long term, I'm hoping that I have my own personal LLM that's able to read all of my emails and help me answer questions about that. The more information that we can attach to a particular query, the more useful the model can be. Now it can get more and more expensive to reason over large amounts of input data. And so that's one of the reasons why there's usually a limit on how big the context length can be. But with NemoTron 3, we tried to push it as far as we could go. We think a million tokens is a lot of tokens, and you can do a lot of things with that. Prism is particularly helpful in sort of multi-step agentic workflows.
48:41And there's this whole separate discussion around context compaction to make sure that the model doesn't get lost in too many tokens. So like, how do you all think about this? A hundred percent. I mean, compaction, that's a thing if you're using an agentic workflow you deal with all the time. And compaction tends to work pretty well because language models are pretty good at identifying the most relevant things and summarizing, and you're basically trying to summarize your context when you compact it. So compaction is not a bad approach. I think having models that can just natively reason about larger amounts of data is just inherently more useful.
49:23So, of course, we want to push the boundary on that as well. Great. Can you talk about the multi-token prediction, which is also very interesting? If you're running at a low batch size, which is when you are trying to get the most interactivity if you're in a data center. So you want the model to respond as quickly as possible, and it's okay for it to be more expensive. Your cost per token might be higher, but you want the result as quickly as possible. Or if you're running locally, so you might be running a batch size one just because you're the only person using it. It turns out that the GPU has extra execution capabilities that are just lying there unused.
50:03The bulk of the work when you're running in these scenarios is actually fetching the weights from memory. And then you push the token past those weights and then you fetch more weights from memory. But it turns out if you push two tokens or even five tokens through those same weights, it would cost basically the same amount of time because the expensive thing is not doing the math to push the token through the weights. The expensive thing is just reading all of those weights from memory, all those parameters they have to come in. And so the idea with multi-token prediction is to take advantage of this by having the model predict multiple tokens at once.
50:40Let's say that the model predicts five tokens. We know the first token is correct. The next four tokens may or may not be correct. So then what we do is on the next pass, we take those four tokens and we stick them into the model and then run it through. And at the end, we check, you know, the model then predicts another set of tokens, right? Then we check where the extra tokens we predict last time correct. If so, then we just accept them and then we get like a 4x speed up. And if they were incorrect, then we only accept the ones that were correct and then proceed from there. So the benefit of this is it doesn't degrade accuracy at all because you're using the model to double check.
51:25So all this speculation is going to get checked during the next token that you run through the model. So it doesn't degrade your accuracy at all to turn on multi-token prediction, but it can give you a speed up and it's probabilistic depending on the acceptance rate of your predictor. So if your predictor is more accurate, the acceptance rate goes higher, you get a higher speed up. So with our recent Neotron models, we're pretty proud of our acceptance rates, but we're always trying to make them better, always trying to improve that acceptance rate. This is a really good example of accelerated computing.
52:01With multi-token prediction, the speed that you get is a function of the accuracy of your model. The more accurate your model is, the faster the inference is, the cheaper the inference is, the more accurate it is. That's not usually how it works, but in this case, that's how it works. And what that implies is that if we're trying as NVIDIA, as a company, to provide meaningful acceleration to the world's most important computational workloads. This has to be an important part of how we think about it. You know, if there's a 3x cost reduction or speed improvement for inference, which is the most important computational workload of 2026, if that's on the table, and it depends on the accuracy of the multi-token prediction network, then that's something that NVIDIA needs to understand very deeply, because it's going to affect our business directly.
52:47Fascinating. To continue on the tour, multi-teacher distillation. We talked about distillation a little bit up front. What does that mean in the context of Nemotron 3? So with Nemotron 3 Ultra, we did post-training using something called multi-domain on-policy distillation. And what that entails is that, you know, we have many different aspects of the model we want to improve. For example, For example, science understanding is different from math theorem proving, which is different from coding, which is different from agent harness interactions, right? With NemoTron 3, I think we had about 10 or 15 of these teachers.
53:32So the idea is that you take these teacher models and you push them as far as you can go on some specific domain, so you just don't worry about making good at everything, just make it really, really smart at this one domain. Then you have a collection of these models and you want to create one model that learns to be good at everything. And we do that using a specific reinforcement learning technique that a lot of labs these days use called MOPD. And the good thing about this is that because the teachers are supervising, they can give really dense rewards to the student model. Basically, every token is getting supervised.
54:08And so the student can learn really quickly and then become, you know, almost as good as all of the teachers at all of the things. So one benefit of this is that it really helps the team work together better. You know, if you don't have a technique like this and you have, let's say, 500 people working to try to make a model better and one team's like, well, I'm trying to make it better at this thing. And then another team's like, I'm trying to make it better at that thing. There can be a tug of war where it's like, well, who wins? You know, and if you have to make a choice like, oh, I'm going to make them, I'm going to choose to prioritize this one over that one.
54:45Then you make the other team feel like their work doesn't matter. You know, it's just really hard. One of the challenges of building AI in 2026 is that you have to figure out how to get the people to work together, even though you're only building one thing at the end of the day. And so this particular technology has been really instrumental in helping more people work together to make Nemotron stronger. Fascinating. So it's just as much a technology question as a human organization question. Exactly. Okay, fantastic. Let's put a pin in this and get back to this in a second because it's a fascinating topic.
55:16In terms of the post-training that you just alluded to, one of the exciting things that you all did in the context of Nemotron is also to publish the data, the training data. Does that include per industry data for specific reinforcement learning tasks? Yes. That's the beauty of a conversation like this today, where you guys can actually talk about those things. So where does one get the data from for post-training, reinforcement learning focused efforts? Obviously, one of the key questions in the world today is that LLMs or AI systems have become great at coding and great at math. The next big question is, can they become great at law and consulting and then all sorts of different domains?
56:12And part of the black box of closed models is how people go about doing all of this, where do they get the data from? To the extent that you can talk about all of this, I'd be very curious about how you guys have gone about it. It's not an easy question to answer because it is quite complex, but I would say we rely on a number of things. One is that we do purchase data from companies that are building data sets that you can purchase. And to the extent that we have the rights to redistribute or to open up that data, we do as part of our NemoTron data effort. With NemoTron, we are trying to be maximally open with the data that we release because our goal is to support the ecosystem.
57:01Our goal is not to be the only model out there. And we love it when we hear of other models around the industry that are using our data sets to make their AI stronger, because that means we're succeeding in our job to keep the ecosystem thriving and growing. Now, we also are big believers in synthetic data generation. we use an enormous amount of compute, running language models on our own systems to create synthetic data that then helps our models be better at solving problems in specific domains. And we release a lot of that data as well. Now, it's, of course, not very straightforward to do this.
57:44Like, you know, AI is always garbage in, garbage out. So you have to work really hard to make sure that any synthetic data that you create is actually adding value that's actually helping the model generalize and solve problems more intelligently. But those are the primary ways that we go about building our data sets. Since we're talking about post-training and RL in different domains, just so curious to get your thoughts on where we go from here in terms of generalization. So just to build on what I was saying a second ago, like the industry seems to be marching from coding and math, which are domains with verifiable rewards to different industries.
58:24Do you think that this is where things are going and that the AI industry as a whole is going to be able to cover those next few domains as efficiently as coding or math? Coding is really special because it's a very intellectual exercise that created a lot of economic value, which then meant that we had an enormous amount of tokens that we could learn from, as well as tooling that allows us to verify whether our models are actually solving problems. So coding is always gonna have a special place in our heart and something that I think AI is gonna continue to get much better at because we have this special relationship with it.
59:13With regards to other domains, I think what I'm excited about about has to do with significantly more diverse environments for AI to learn in during reinforcement learning. I believe that reinforcement learning is such a general form of teaching an AI how to solve problems. We're just getting started at figuring out how to apply that. And I think as As our environments get more sophisticated, the AI then learns more understanding of the problems that it's trying to solve as well as the implications of the actions that it can take. Then it becomes much better at actually solving those problems.
1:00:02When I look at the environments that we're using today, they're still fairly simple, all things considered. And I think that's going to become significantly more complex and diverse over the next few years. All right. So you mentioned making 500 people work together. And I said that we would get back to it because it's so interesting. So just taking a step back, tell us about the research organization at NVIDIA. How is it structured? How does it all work? Well, NVIDIA is not structured according to an org chart. We have one, but it's not actually the best way of understanding how we work. My team, for example, is not part of the official NVIDIA research team.
1:00:48My team is actually part of the organization that builds the GPU. And my team is not the only team building Nemotron. There's probably 10 teams around the company that have significant involvement in building Nemotron. in different parts of the company, in enterprise software, in our AI software division, the part of NVIDIA that actually designs the GPU also significantly is involved in building Nemotron. So there's so many different teams that have to work together. We always like to say that the mission is the boss rather than the organization. But what that implies is that people have to figure out how to work together, which is challenging in the sense that humans are naturally tribal creatures.
1:01:43And it's not natural for us to be friendly with people we don't know very well or trust co-workers that we don't have success working with in the past. And, you know, actually the name NemoTron reflects that. We had the Nemo team, which was building software for AI, and the Megatron team, which was building primarily focused on systems research for building large language models. And, you know, we decided to work together and then start calling our projects Nemo Tron reflecting, you know, sort of the collaboration between these teams. Since then, Nemo Tron has dramatically expanded. There's so many more teams that are part of the effort.
1:02:28And it's really important that we have structured it in this open way inside of NVIDIA. You know, we are inviting volunteers from around the company to come help build NVIDIA's AI. We think it's very important to the future of the company. And, you know, as that vision continues to develop, more and more people want to join. That's fantastic. We're really excited about that. And it means that we then have to figure out how to organize the work so that everybody has a chance to contribute and feel heard and feel like their ideas are, you know, fairly evaluated on the path towards impact. We have a formal process for doing that.
1:03:09We have an internal website where people share ideas and then those ideas are assigned to one of 25 different leads that are, you know, over various parts of building Nemotron. They interact with those ideas. Some of those ideas get further developed. Some of those ideas get deferred until, you know, the next time we go around building a new model. But we're trying to build Nemotron in an open and inclusive way so that we can really come together as a company to build it. I think organizations that figure out how to collaborate to build AI succeed. Organizations that struggle with control over who owns the AI tend to waste a lot of effort.
1:03:54And so NVIDIA's success and Nematron's success, I think, is directly proportional to our ability to collaborate. It's something that I care deeply about. Fantastic. But you mentioned earlier that despite the fact that you work at the number one undisputed leader in GPUs, you all as a research organization don't have all the GPUs that you would want in the world. So like how does the allocation of GPUs and computes happen? Is that based on how promising an idea is or early success? Do you give GPUs, withdraw GPUs based on success? It's a really complicated question and it's obviously a difficult problem for everyone in the industry to figure out how to allocate their compute.
1:04:44Inside NemoTron, we have a budget for NemoTron and inside NemoTron we allocate compute based on what we think the needs of the project are. We have a hierarchy, so we have a set of programs and inside of each program we have have a set of projects and each of them put forward their requests. And then, you know, we have a two-week cycle where we review requests and we review the budget and then we make decisions in kind of a hierarchical way. And then, you know, compute gets decided that way. Now, having said that, this is something that I think we can still do better at. It's hard when we're making decisions about compute allocation because every researcher is convinced that their idea could change the world if it just got a thousand times more GPUs attached to it, right?
1:05:39And they might be right. It might actually be true. And yet we're running at the limit. We don't have a thousand X more GPUs for every idea that we have. We have to operate within the limits that we have. And so it is a challenging process. we try to incorporate as many people's perspectives into that as possible so that it's as much as possible a shared sense of understanding, maybe not agreement. So there may be times when one project feels like it really deserved more GPUs because the impact of that would have been so high, but it didn't get it. We hope in that circumstance that they have an understanding of why some other project did get more GPUs and why that was considered more of a priority during this particular allocation round for the company so that people can at least understand that there's a reason for the allocations that we have.
1:06:35Having said that, this process is always improving. There's always more work to be done to make this more transparent and more fair. And then, of course, my number one is just to get more GPUs so that we can also fund more things because I would like to do that too. How do you balance useful research with great exploratory research? My belief is that research needs to be bootstrapped. Research is a chicken and egg problem. So it is always the case that every researcher believes if I just had a lot more resources my idea would change the world. Actually, it's important that researchers feel that way, because if you didn't feel that way, you wouldn't have the conviction that's required to go do something crazy and new, right?
1:07:21So you have to believe. And so of course, you start with that belief. But then how do you translate that belief into something that other people can understand, right? That other people are willing to invest in? This is what I call the chicken and egg problem, right? Because once your research idea is obviously good and impactful, it's easy to get resources, but how do you get it to be obviously good and impactful without those resources, right? So the way you solve chicken and egg problems is by bootstrapping. This is an iterative problem solving approach where you do something small, you get some sort of signal about this is a good idea, and you tell people about that.
1:08:00And then you ask for just a little bit more. And if people saw like, oh yeah, that experiment turned out pretty well. That's pretty intriguing. We should probably do a little bit more there. Then you're on track, right? And that over time, iterate a lot, iterate quickly, iterate many times. You can bootstrap to finding significant resources for your idea and also usually attracting more people to come along with it on the way because they have a chance to see that this idea is going to change the world and then they want to be part of it. Is that how the moonshots at NVIDIA got started as well over the years, whether that's in AI or otherwise?
1:08:41So it was bottoms up, somebody coming up with a good idea versus Jensen saying this is what we need to do? Well, you know, Jensen has lots of good ideas, too. And so the company is very responsive to his ideas. And that's that's important as well. But Jensen very explicitly says all the time, this is a company of volunteers. You know, each of us is here because we choose to. We could we could be doing something else in our lives, but we choose to be here. And so, you know, we tend to make decisions, especially for early stage research, it tends to be very bottoms up because, you know, it's sort of an invitation, like bring your best ideas.
1:09:23Let's figure out, you know, what are all of our best ideas and then we'll take a step from there. Now, do we sometimes have top down ideas that are important for the company strategy? Of course, you know, of course. um uh nvfp4 pre-training is one of those you know so we decided as leadership of the company we're going to really invest in in fp4 hardware um now it's time to go invent some optimization algorithms that succeed in using it and um so uh so we told the team we didn't say to the team you have to work on nvfp4 pre-training what we said is there's an opportunity we're making a big investment and if we can figure this out it will be significant for our company and then we We let the people who are interested in that work on that.
1:10:09And as a result, we succeeded. So it is a balance of bottoms up and tops down. But it always has this bootstrapping feeling. Even with something like NVFP4, where there's a significant strategic top-down component, the actual technical solution, which is very intricate and complex and has a lot of moving parts, that came from the researchers themselves. And, you know, that's my belief is that research always comes from the researchers themselves. You can't tell research exactly how to go solve a problem because then it wouldn't be research. It would be engineering. But in a world of AI where the most important problems we have to solve all have this research component, there needs to be freedom for researchers to innovate if we're going to make progress.
1:10:53listening to everything you're saying i'm struck by how entrepreneurial the culture at nvidia still seems to be so like i'm so very large companies i'm i'm sure there's all sorts of politics and you mentioned the tribal instincts like i'm sure all of this is happening but um especially given you know how long the company's been around a phenomenal success the fact that people have been making a lot of money internally it still seems to be very entrepreneurial bottoms up driven and maybe meritocratic. Is that the right takeaway? Yeah, I mean, one thing that's very unusual about NVIDIA is the tenure of its leadership.
1:11:33Jensen Huang has been running the company for 33 years, but he's not alone. There are a lot of other very senior leaders in the company who have been there for three decades or longer, including my boss. And these people remember what it feels like to work at a very small NVIDIA. And they know what it feels like to work at a very large NVIDIA. They have a shared sense of ownership for the company. NVIDIA is a place we often say, no one fails alone. And the point of that, that's just a statement of fact, right? You work at a company, it's one company. You all succeed together, you all fail together.
1:12:13You work in accelerated computing. Accelerated computing is the composition of thousands of technologies. If any of them fail to deliver acceleration, the value is destroyed. It doesn't matter whether the chip is great if the compiler sucks. At the end of the day, the thing that you're selling is time and capability to researchers that are trying to build the future of AI. And if they don't get that, it doesn't matter whether it was the, you know, the transistor or the math unit or the compiler or the library or the networking or anything else along the way that failed to live up to its expectations.
1:12:46the whole thing in composition fails, the whole value is destroyed. And so we have a deep understanding of that culturally at NVIDIA. And it is something that motivates the way that we work together. Maybe to close the conversation, I'd love to zoom out, get your take from the perspective of somebody who's like as deep into all of this as it gets about where things may be going. So like who knows in a few years, but I don't know, in the next year or two, maybe there's some visibility. I read somewhere that you're not necessarily a big singularity kind of person. Is that fair? True. And why is that?
1:13:29Well, I think that intelligence is just so incredibly multifaceted. You know, I always think about this question, like, if a company were to be looking for its next CEO, would it find the next CEO by looking for somebody who won the International Math Olympiad? Probably not, right? Even though it's incredible for people, I could never even compete in any way at the International Math Olympiad. Those people are amazing, right? They have just incredible brilliance. That's not the right kind of brilliance to run a company. If we look, for example, at other aspects of our culture that are really important.
1:14:11For example, musicians. What kind of intelligence does it take to become a hit musician? Don't assume that it's all luck. It's not. These people are working hard and they're very smart in ways that I might not understand with my PhD, right? I might not have that kind of intelligence. And so when I think about intelligence, I think it's just so multifaceted and so contextual. You know, it really depends on the situation. It's not just about raw intelligence. Raw intelligence is kind of like the horsepower of an engine, but an engine running without wheels doesn't go anywhere, right? So intelligence, the impact of intelligence has a lot to do with the context that the intelligence has put in the harness, the platform.
1:14:56And so when I think about that, I think, you know, So the singularity is, although it's an attractive idea, I think that it's really a wrongheaded idea because it doesn't really take into account these other factors. So I believe that artificial intelligence is going to continue to develop at a rapid pace. It's going to unlock significant capabilities for people in every aspect of our world economy, people doing every kind of work. I'm very excited about the opportunities that it's going to bring. I am also a little bit concerned with how we're going to manage the transition. So I do think that transitions are hard for humans in general, like we're conservative generally.
1:15:44And, you know, there is going to be a lot of change. This is a profound change in the way that we think, in the way that we work, the way that we learn. um uh ultimately i have faith in our ability as humans to figure it out um you know we've done it in the past this is how this is who we are um we we build tools uh we build external organs that help us solve problems you know we we have an external stomach we call it kitchen it creates enormous value for us we can eat things that we couldn't eat without a kitchen right now we're creating an external brain. You know, the implications of the external stomach were pretty profound for us as a species.
1:16:23They led to agriculture, which led to organized societies, the way our cities are built. So we think about what is the implications of an external brain? Pretty profound. Nobody actually really knows. But what I do believe in is the power of humanity to solve problems and to learn and to incorporate new technologies in ways that benefit us. I also believe that the problems we face as a planet all require more intelligence, every single one of them, whether that's inequality or climate change or any of the other structural problems that I think are very worrisome that we face. the solutions to those are going to require invention and intelligence.
1:17:07And what that means for me is that the only kinds of tools that we can really create moving forward are going to be AI. Because the problems that we face are all about intelligence. And regardless of the technological approach to solving those problems, the solutions will always be called AI. and so that makes me hopeful for the future but also you know somewhat you know respectful of the challenge that it is going to bring to us as we we try to figure out how to live in a new way with this new external brain but I believe in our our ability to to learn and to change and I think ultimately this is going to make our lives better.
1:17:50Do you guys feel the AI backlash that seems to be forming internally? Is that something that you all perceive, think about? And if so, do you think it's a communication problem that our industry may have, you know, in particular, given what you just said about all the obvious potential of AI? You know, I'm always worried about the way that the public thinks about technology and interacts with it. It matters a lot. And it is definitely the case that societies that want technological advancement have more technological advancement than societies that don't want change. So I think it is actually important to think about it.
1:18:30One thing that's interesting about AI is that I believe it tends to be much more accepted when it is part of everyday life. And at that point, people stop thinking about it as AI. It's just, oh, this is the tool that I use. Like, do you care whether it's AI that's helping you route your car when you ask the map application to help you drive somewhere? Like, I mean, it is. There is actually sophisticated AI that's going into that. But you're not really thinking about that, right? You're just using a tool. And so I feel like people's acceptance of AI, you know, comes with experience, right? The more experience we have working with it, the more we learn how to work with it productively.
1:19:15I think the more comfortable we become with it. Great, Brian. So it's been a fascinating conversation. Maybe as a very last question to make sure we cover it, I want to make sure that we talk about safety. What is the state of safety currently? And where does open source and closed source sort of fit in the safety conversation today? Safety is on everybody's minds right now. You know, watching the Fable release and the way that the government interacted with that, I think, is a consequence of concerns about safety, about these models. You know, they get stronger and stronger and then they could be misused.
1:20:02And, you know, there's different approaches to thinking about safety and trying to define safety. I have maybe a slightly unorthodox opinion about this, which is that I think open technologies are generally safer because there's more sunlight. You know, when more people are thinking about the safety of a technology and evaluating it and then contributing to making it safer, I think that's inherently safer than having a small group of people being in charge of safety for everyone else. I also think with artificial intelligence, because it is really about ideas, it's really about exploring ideas in different ways, that diversity is more safe than monoculture.
1:20:51And what that means is that there's going to be different beliefs. Like diversity isn't just about like the easy stuff. Diversity is about the hard stuff. Like when people have deeply felt disagreements, they really, really totally disagree with each other. Making it possible for people to explore their ideas in a diverse way, I think, is more safe than trying to create a walled garden where, you know, certain ideas are considered safe and certain ideas are considered unsafe. And, you know, this is controversial in today's AI environment, which I think is interesting because we've had hundreds of years of tradition that speak directly to this.
1:21:36You know, in in the United States, for example, we have laws about freedom of conscious conscience and freedom of speech. And, you know, it's not because we didn't consider for thousands of years, would it have been safer if we didn't have those? Right. We tried that. We tried actually having a monoculture about like these ideas are safe to talk about. these ideas are safe to believe. And we found that to be much less safe than a pluralism where we officially don't take a position about what ideas are safe. We actually found that as much safer as a society to support diversity than it is to try to keep everybody safe top down.
1:22:19And so I believe that open technologies for AI are inherently the safest way of building AI. All right. Love it. Controversial take to close the conversation. Brian, it's been fabulous. Thank you so much. We appreciate your spending time with us today. Thanks for inviting me. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests.
1:22:55Thanks and see you on the next episode.
From the publisher
NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models — and then give them away for free? Bryan Catanzaro is VP of Applied Deep Learning Research at NVIDIA and one of the people whose work quietly underpins modern AI: he helped create cuDNN (NVIDIA's first deep learning product), co-invented DLSS, and named and built Megatron, the framework behind how much of the industry trains large models. Today he leads Nemotron, NVIDIA's family of open models — and Nemotron 3 Ultra, released just weeks ago, is one of the strongest open-weights models to come out of the US.
Matt Turck sits down with Bryan for a genuinely deep conversation: the real business logic behind a chip company building its own models, the state of open vs. closed AI, and whether the US is falling behind China in open models. Then they go inside Nemotron itself — four-bit (NVFP4) pretraining, hybrid Mamba-Transformer architecture, mixture-of-experts, multi-token prediction, and multi-teacher distillation — all explained in plain language. Plus a rare look at how a modern AI research org actually runs, what it was like working alongside Andrew Ng and Dario Amodei at Baidu, why Bryan doesn't believe in the singularity, and his contrarian case that open AI is safer than closed.
A reference conversation for anyone trying to understand where AI is really headed.
(00:00) — Cold open & Intro
(01:33) — Is open source AI catching the frontier?
(05:29) — Do closed labs blocking distillation slow open source down?
(07:42) — Is the US falling behind China?
(10:30) — Why companies actually choose open models
(12:39) — A "crazy" 2008 bet: machine learning on GPUs
(15:33) — Working with Andrew Ng and Dario Amodei at Baidu
(17:41) — Coming back to NVIDIA: DLSS and the birth of Megatron
(21:55) — The real reason NVIDIA builds its own models
(24:28) — Is Moore's Law really dead?
(33:37) — The Nemotron family: Nano, Super, Ultra
(35:09) — Built for agents: why NVIDIA bets on speed
(36:02) — How you train a 550B model in 4 bits
(39:25) — Hybrid Mamba-Transformer, explained simply
(42:31) — Mixture of experts — and why NVIDIA built NVL72 around it
(47:26) — Why a 1-million-token context window matters
(49:26) — Multi-token prediction: how the model predicts 5 tokens at once
(52:47) — Multi-teacher distillation: teaching one model from many
(58:01) — Where reinforcement learning goes next
(01:00:16) — Inside NVIDIA's research org: "the mission is the boss"
(01:04:03) — How NVIDIA decides who gets the GPUs
(01:10:53) — Why NVIDIA still feels entrepreneurial after 33 years
(01:12:58) — Why Bryan doesn't believe in the singularity
(01:17:50) — The AI backlash
(01:19:18) — The controversial case: open AI is safer than closed
