Lowering the Cost of Intelligence With NVIDIA's Ian Buck - Ep. 284

29 Dec 2025 · 38 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Mixture of experts (MoE) for frontier AI—how it reduces inference cost (“cost per token”) while boosting intelligence—and how NVIDIA’s hyperscale/HPC infrastructure (NVLink, GB200/NVL72) enables efficient MoE execution via extreme co-design.

Guest backgrounds

Ian Buck, Vice President of Hyperscale and High-Performance Computing at NVIDIA.

Key claims

MoE splits a model into many “experts” and activates only a small subset per token, cutting compute and cost. The main hidden cost for MoE is fast expert-to-expert communication; dense, low-congestion GPU networking is required to avoid idle time. NVIDIA’s NVLink/NVL72 connectivity lets experts run across many GPUs at full speed, enabling large models with sparse activation.

Notable examples

Llama 405B activates all 405B parameters; GPT-OSS ~120B total activates ~5B (~10x+ compression), with benchmark cost reductions (e.g., ~$200 vs ~$75 for scoring). DeepSeek R1 used 256 experts/layer; NVIDIA cites ~15x runtime performance improvement and ~10x cost-per-token reduction when scaling from 8 to 72 GPUs. Mentions KimiK2 (trillion-scale) activating ~32B (~3%) with 340+ experts/layer.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Mixture of Experts

0:45 to 3:32

Ian Buck explains the concept of Mixture of Experts and its significance in AI.

“If you look at the top 10 open models on artificial analysis right now on their leaderboard, they all share the MOE architecture.”

Benefits of Mixture of Experts

3:32 to 5:07

Exploring the advantages of using MOEs in AI models to improve efficiency and intelligence.

“That's a way of making AI cheaper and still, but still being able to encode all the possible information and answer all the questions.”

Architecture and Training of Experts

5:07 to 8:12

Discussion on the internal architecture of MOEs, including routers and expert activation.

“For Llama 405B, I think they currently cost about$200 for them to actually ask what cloud service to get all the answers to create that score.”

The Evolution and Prominence of MOEs

8:12 to 10:39

The historical context and recent advancements leading to the popularity of MOEs in AI.

“and you can train these little routers and combiners, and then you just do that in multiple, multiple layers, and sure enough, at the end of it, you've got a chat model like GPT-OSS or KimiK2.”

Cost and Intelligence in AI

10:39 to 13:08

Examining the balance between AI model size, intelligence, and cost-effectiveness.

“So we know the DeepSeek moment was huge, as you just said, for many reasons.”

Hardware and Model Relationship

13:08 to 14:00

Discussing how advancements in hardware influence AI model development and efficiency.

“So there's always a driving cost and a continuous, like let's increase intelligence and let's lower costs.”

Cost-Effective AI Hardware Development

14:00 to 21:30

Learn how NVIDIA's advancements in AI hardware reduce costs while enhancing performance.

“of inferring here a little bit, but I would imagine it's more expensive to train, to architect to train, perhaps not to run, but total cost.”

The Importance of Communication in MOE Models

21:30 to 28:09

Discover how NVLink improves communication efficiency in mixture of experts models, enhancing performance and cost-effectiveness.

“We're able, instead of talking over each other, just communicate at one as one at the speed of light.”

Extreme Co-Design at NVIDIA

28:09 to 29:50

Learn how NVIDIA's extreme co-design approach enhances AI model efficiency and reduces costs.

“We're super happy with GB200 and what it's been able to do for inference and just keeping the cost and driving the cost of tokens down, down, down while intelligence goes up, up, up.”

Challenges and Confusions in GPU Performance

29:50 to 31:21

Discover the complexities of GPU performance and the importance of co-design in AI.

“And we work really, really hard to continuously work on performance, not just to have the fastest and be the fastest, but also to reduce the cost.”
Show all 12 chapters

The Future Beyond Mixture of Experts

31:21 to 34:24

Explore the potential future developments in AI models beyond the Mixture of Experts (MOE) approach.

“We've been talking about MOE in the context of language models predominantly now.”

The Benefits of Intelligent Models

34:24 to 35:51

Understand how smarter AI models can create opportunities and reduce operational costs.

“We've got lots of different irons in the fire.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10NVIDIA AI Podcast Host:Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. Ian Buck is here with us today. Ian is Vice President of Hyperscale and High-Performance Computing here at NVIDIA, and he's here to discuss mixture of experts, the architecture powering the world's leading frontier models, and how extreme co-design can both be a good idea of the world's world. drive down the cost of generating intelligence today, and future-proof your AI platform for whatever advances come tomorrow. Ian, welcome. Thanks so much for taking the time to join the podcast. Thanks, Tom. Glad to be here. So let's jump right into it.

0:44NVIDIA AI Podcast Host:What is mixture of experts, MOE as we call it? Why does it matter? If you look at the top 10 open models on artificial analysis right now on their leaderboard, they all share the MOE architecture. So can you explain kind of in lay terms what MOE is and why it's suddenly become the standard for frontier AI? Yeah, it's a great question. I think there's a lot of, it's a term that is used in industry and amongst AI researchers, but it's not really understood, like, what does mixture of experts mean? Yeah. We've all heard of neural networks, and that's what these neural networks are. They're neurons, they're parameters, they're components of a AI model.

1:23and when AI got started and really became in the zeitgeist of the world, the neural network was simply, each parameter represented a neuron of the model. And we heard about a 1 billion parameter model and a 10 billion, a 100 billion, now trillion parameter models. Those are basically the neurons of the AI brain that you activate when you ask ChatGPT a question. But something happened along the way. As these models got smarter and smarter and smarter, they naturally got bigger and bigger and bigger. In fact, two years ago, when Lama first came on the scene, there was a 7B Lama, and then there was a 70B Lama, and now we have a 405B, BB billion, parameter model.

2:06And that makes them smarter. They have more information, they understand more things, and they give you better answers. But there was a problem. As they got smarter and smarter and smarter, to get the answer, you actually had to ask and activate every neuron in that brain. So as a result, while the models were getting more and more intelligent, they're also getting slower and slower because you had to ask every neuron and calculate every neuron and perform all the math on every neuron on a GPU. And then it wasn't one GPU, it was lots of GPUs and even more. Along the way, researchers came up with this idea.

2:41And they realized, just like a human brain, we probably don't need all of these neurons to ask every question. Simple questions, probably just a few neurons. or different parts of the brain may encode different information. Let's just activate those. So to make the AI cheaper, the tokens, which is the piece of data that's flying through that eventually becomes a word on the screen, the tokens cheaper, let's only activate the neurons we need to activate. And that's what mixtures of experts is. Instead of having one big model, we actually split the model up into smaller experts. Same number of total parameters, but now we only ask the, we train the model to only ask the experts that probably know that information along the way.

3:21And that's part of the training process to build that model. Once you do that, you can have a model which has maybe 100 billion parameters, 100 billion neurons, but we only ask or activate about 10 billion. That's a compression mechanism. That's a way of making AI cheaper and still, but still being able to encode all the possible information and answer all the questions. So today, most models today are achieving higher and higher intelligence scores by taking advantage of having more than lots of experts and able to ask, have the model, as it comes up to the answer, ask only the right experts in order to get the right answers.

3:57To put some numbers behind it, we have that LAMA 405B, 405 billion parameter. That's one big model. On leaderboards like artificial analysis, you mentioned, it gets an intelligence score of about 28. 28 is just a weighted score of the benchmarks they tested. But all 4 and 5 billion parameters are getting active. Now, fast forward to a modern open model like OpenAI's GPT-OSS model. It has 120 billion parameters, actually a little bit smaller in total parameters. But when you ask a question, it only activates on the order of about 5 billion parameters. So instead of 405 billion parameters and all that math and all that cost, it actually only needs to activate about 5 billion parameters.

4:40That's like a 10 to 1 or beyond compression. Yeah. Making it cheaper. And then it gets an intelligence score of 61. So it is going from 28 to 61. It's going from 405 billion perimeters to 5 billion perimeters. Way cheaper. It's not 10x cheaper. It's still complicated. And we can talk why these MOEs are complicated to run. But artificial analysis does measure the cost to run the benchmark. So like how to run and calculate the intelligence score. For Llama 405B, I think they currently cost about$200 for them to actually ask what cloud service to get all the answers to create that score. They asked GPT-OSS the same thing.

5:21Its tokens are cheaper and only cost about$75. So MMEs are making models, allowing models to get bigger, smarter. It's allowing them to get cheaper. And as a result, advancing AI. Now, of course, across the board, all the leaderboards, they're all these mixture of expert models.

5:40NVIDIA AI Podcast Host:Correct me, bring you back on track if I get off here with the questions. But from kind of a layperson, use that word standpoint, if I'm trying to wrap my head around this idea of mixture of experts, are the experts divided up in ways that I might think about knowledge? You know, this expert handles math and this one handles science and this one handles, I don't know, visual understanding. Yeah, it's a great question. You know, that is the art of training these things. In fact, AI, it's not like hard coded in there. They don't train a separate model for math, doing math questions and a separate model for telling you how to make a pizza.

6:18The AI, the beauty of AI is that the algorithms that these researchers and scientists and companies like Anthropik and OpenAI and everybody else have figured out is that they can just give it the data and they encourage the model to sort of camp, to identify and create these little pockets of knowledge. It's not prescriptive. It's just the data that they're seeing naturally clumps the activity of these different questions to different experts. So, and then in front of those experts, there's this thing called a router. And the router actually is able to just look at the string of questions, like what's the answer is forming, how is it thinking, and then be able to predict, you know what, this one probably goes to that guy or this other guy.

6:58In fact, today's experts, they may have on the order of dozens of experts on every layer of the model. And there's a little router between, and they may actually ask not just one expert, but like at every layer, they may ask two experts or eight experts. And then there's another unit, a model which listens to all the experts. This guy says, I'm pretty sure I got the right answer. Maybe I got the right answer. I don't know. I don't know. I don't know. Combines the answer and then goes to the next one. So that's actually the architecture of it. You know, it's kind of like you could train one person, one brilliant scientist.

7:32You train an Einstein to be able to answer any question. That's really hard. Takes a lot of energy. that's a very expensive person to hire and have on staff. Instead, maybe I can hire a couple of domain experts or teach a couple of different people some stuff. And, you know, I just give them all that question. They can all answer it very quickly in parallel. And the combined knowledge, and that's actually how we work today. We don't work in one person is not a company. Companies exist because we have all this expertise around. And the MOE method is basically applying that to AI. So the models are all trained that way.

8:06There's all sorts of training methods to create the condition where information activations can start grouping and gathering together, and you can train these little routers and combiners, and then you just do that in multiple, multiple layers, and sure enough, at the end of it, you've got a chat model like GPT-OSS or KimiK2.

8:26NVIDIA AI Podcast Host:Yeah. No, MOE isn't new to 2025. The idea of the architecture has been around for a few years. So was it being used? Has it been being used all along, and we just weren't so aware of it? And then why has it kind of come to prominence lately? Yeah. The idea of experts is not new in machine learning. You know, before AI, there was an idea of creating, combining multiple machine learning models together and how to do that with statistically to improve the accuracy. There's all sorts of history and math around that. Applying it to AI, though, is relatively new. You know, the early versions of, we now know, we're chat GPT.

9:04They were a mixture of experts and models, but they were not publicly known. It really wasn't until the DeepSeek moment, which is about a year ago, where it really blew the doors open. Because DeepSeek, those researchers, were the first to really build a world-class MOE-based model. People had written papers about it, but it was one that actually competed and demonstrated the intelligence scores that competed even with the closed-source models. And it was a beast. it was awesome. It had 256 experts in every layer. I mean, it did every single optimization. And as a result, it was extremely cheap to run.

9:44Incredibly complicated, but cheap to run because it was so, it took MOE all the way to the extreme. And maybe many people think it's kind of where OpenAI was, you know, with the original GPT. So now, once we had that moment, you know, the first time DeepSeek was run on even GPU systems, it actually didn't run that well because we didn't have the infrastructure or even the software to run it that well. The DeepSeq engineers had written all this custom code to make it run awesome. But at that point, every model, every researcher realized, hey, this thing's real. We now can see how we do it. They made the whole thing open.

10:17They published the paper. It's a brilliant paper. And it shows the opportunity for MOE. And since that moment, you can see that every model now has shifted to building MOEs. DeepSeq sort of shined a light on how to do it, how to train it, how to do inference and deploy it, and sort of kicked off that revolution of MOEs that we've been enjoying.

10:39NVIDIA AI Podcast Host:Right. So we know the DeepSeek moment was huge, as you just said, for many reasons. Is that kind of, are we going to look back and say like, hey, the lights went on then and, you know, new things will come. But for the moment, is everything MOE? And if not, why? What's kind of the, I don't know, the decision-making process? when would you train a model to be MOE and when would you not? You know, I think all the models that really are focused on providing an intelligent response, it makes a lot of sense why they're MOE. You want to do your best to encode as much knowledge into the neural network.

11:18So it just knows things. You don't need to, like on pencil and paper, write two plus two to work out that it's four. You just know two plus two is four. So the more neurons you can throw into a holistic model, it gives it innate knowledge. It doesn't have to work that out in a reasoning chain or such things. So there's a huge advantage to having models be bigger as long as we don't increase the cost. And that's why MOEs, we want to be able to push the limits of only activating 10%, 5%, 3 % of the neurons, more and more experts. And you can see that in the research and the way the models are evolving.

11:55They're really pushing the limits of some of the modern models. They'll have 300, 400 experts they're trying to combine. Now, getting all those experts and all that communication is complicated. We'll talk about that. But it is innate by having that foundation model with all of those experts. It allows them to then apply all the other techniques of inference, of reasoning. It allows models that are smaller to be distilled and fine-tuned for specific tasks. it creates a foundation for the rest of the AI models around the world. Certainly some of the smallest models, for the more dedicated individual use cases, I've got to put a box around a stop sign, or I've got a ring doorbell that uses AI to detect if it's a squirrel or not a squirrel.

12:42Those small models may not, they need to do one specific thing. Probably I can get it, squeeze it down. I don't need to go to the complexity of an expert system. But anything that wants to be agentic, any kind of agent, And pretty much most of the AIs that we interact with purposefully, they're all MOEs because they can be thrown in the need to know and they need to be able to reason about a wide variety of different stuff. And it makes AI cheaper. Yeah. It lowers the cost per token. So there's always a driving cost and a continuous, like let's increase intelligence and let's lower costs. We can do both with MOEs.

13:18NVIDIA AI Podcast Host:I was going to ask you about that because there's this, it seems like there's this focus happening now. You know, generative has progressed far enough, and certainly it's everywhere, you know, including the news, the business section, if you will. There's a shift kind of from, you know, the biggest models, raw speed, you know, the highest scores, to as you said, how much does this cost? And can we get it to be cheaper while being just as smart, if not more intelligent? So we're calling it tokenomics, right? So not in the sense of blockchain or crypto tokens, but as you mentioned, AI systems generating tokens, reasoning tokens, output tokens, what have you.

13:55NVIDIA AI Podcast Host:So if we're focused on bringing the cost down, how does a more complex system, and I'm kind of inferring here a little bit, but I would imagine it's more expensive to train, to architect to train, perhaps not to run, but total cost. How does a more expensive kind of premium system actually drive the total cost down? Yeah, there's a wonderful symbiotic relationship. that happens in the market between the AI hardware and the models that are being created to serve AI. They inherently, and they kind of have to make sense, you know, if the hardware offers a certain level of connectivity, a certain GPU performance, a certain memory size, obviously building an AI model that's even bigger is going to be hard to take to market or even not possible to efficiently train.

14:39So, you know, since the beginning of the original Kepler GPUs that were used for those first CAT AIs to today's modern GB200, GB300, MVL72X, you can see a pattern where, you know, with every new platform, we advance the state-of-the-art or what the capabilities of what NVIDIA is able to offer. The compute performance, the memory performance, the connectivity, the IO, we'll talk about NVLink, those things enable the next wave of building, to train the next model, but also to do inference. You know, it's the, they add complexity. When we started, we were doing PCIe cards, little graphics cards that you plugged into the server equivalent of your PC and use the floating point calculations and the graphics memory in order to do that computation.

15:27And they were great. When the AI revolution took off, we saw that by adding more floating point calculations and building a bigger GPU, adding things like HP and memory, adding things like increasing the power beyond what a typical PCIe slot will do. we often would increase the performance of what was capable in the AI not by just the percentage of more flops or memory bandwidth, but by X factors. And that's really because the AI models that were able to build were bigger, smarter, and could run more efficiently and could do more things. You know, TCO, people talk about TCO as the cost. You know, TCO actually is just, it's not the goal.

16:07Like in itself, it's just the lowest cost. You want the lowest cost? You buy one GPU. Sure. The goal is actually to improve intelligence and intelligence per dollar, the cost of that intelligence. Or if we're at the same level of intelligence, say this 60 score from artificial intelligence, are we reducing? Are we reducing the cost of that intelligence over time? The tokens that people need to buy are the costs in order to run it. That's really the goal in every generation of NVIDIA architecture. You know, we're looking to figure out what technologies can we incorporate, expand, double down on, invest in, or pull from the community, or pull from our partners in order to deliver X factors of performance improvement, where the model, even the existing models like the current MOEs could get an X factor of performance improvement.

16:57Well, only, you know, we're not afraid to add more cost and more technology on a per GPU basis. You know, the HBM memory is a lot more expensive than the old school graphics memory. But it only increases the cost in percentages. Because you now have HBM and because you have the bandwidth that it offers to and can connect to that much floating point, you can deliver an X factor in total end-to-end performance. Yeah, yeah. And we saw that, actually. When DeepSeaGuard 1 came out, the GPU at the time was the Hopper H200 system. Hopper had eight GPUs in a server. They were all connected with NVLink through an NVLink switch.

17:35So we could effectively build one giant GPU of eight GPUs working as one.

17:40NVIDIA AI Podcast Host:Right. That was really important. The model was so large, it couldn't really fit on a single GPU. It had to use multi-GPU. And the researchers that built DeepSea took great advantage of that. It also had an Envilink capability. So we could actually put every expert on different GPUs. And you could see that. You could paralyze the work. You'd run anything even more efficiently, even faster. And because as those experts all had to talk to each other, they would do that over Envilink. So that was really important. Before we had Envilink, you would have to send things over a PCIe bus and only one could talk at a time.

18:15And it was much slower. Because we have Envilink, all those GPUs can talk to every other GPU at full speed. It's a totally unblocked, literally at gigabytes and terabytes a second of bandwidth without any concern for collision. It was critical for those deep secretures to get good performance. If you fast, so obviously it also happened at a time, which now we can say is when we're in the heart of bringing and building what is now the GB200 and VL72, where we scaled up the number of GPUs we can connect from just eight GPUs in a server to 72 GPUs in an entire rack, a 9x multiple. Yeah. Now, that's a lot more GPUs.

18:56So did the cost go up? And it certainly, obviously, that many GPUs, entire rack with GPUs versus a server is a lot more money. Sure. In fact, we actually even had to add more technology because we needed to take all those NV switches and build a separate NV switch plane. It does cost more. But because we did that, we can actually paralyze and improve the performance of DeepSeek R1 even more. You can take all those experts, and instead of having to try to make it all fit and work within only eight GPUs, we could actually get all 72 GPUs working as one. And that improved performance of just going generation over generation, being able to further paralyze and run all those experts across it could actually increase the performance so much that we actually got a 15x improvement on running DeepSeek R1 versus only adding percents, about 50 % more total cost on a per GPU basis.

19:52NVIDIA AI Podcast Host:Wow, okay. That actually generated a 10x reduction in the cost per token. Right, right, right, right. So we do have to add more technology. We want to keep running more technology. NVIDIA is a technology company, but we turn that technology back into performance, which in the net of it reduces the cost per token because it's that much faster. and as a result, they can actually run, get more out of that rack, more out of the, on a per-GP basis. And we've taken it down from what was Hopper. It cost about a$1 to get her a million tokens, roughly a million words. It's now down about 10 cents. So people look at the rack and they say, it's really expensive.

20:32But the way you do that is actually you've put all that investment in MVLink and all the connectivity and all the next generation software. And you also do all that software work to make it all work really well. And generation over generation, you get that multiple, that 10x multiple in the reduction in cost. That's just one model. That same story is playing out for GPOS-S and everything else. And those are models that were built and trained and designed for Hopper. You know, we're entering into the, you know, starting to see some models come out that are trained on Blackwell. And you're going to see that, you know, now raise the bar and go even further.

21:07So this is the virtuous cycle that we've been working so furiously to help make happen. We might add percents in terms of cost and complexity per GPU basis, but we aim at every generation to deliver X factors of performance. And as a result, we dramatically lower the cost per token by 10X.

21:29NVIDIA AI Podcast Host:As I'm listening to you describe, you know, NVLink and the advances in getting the experts, getting the GPUs to communicate and kind of act as one, I can't help but think like, we need NVLink for like Teams meetings. so we can get everybody. We're able, instead of talking over each other, just communicate at one as one at the speed of light. That's right. Speaking with Ian Buck, Ian is vice president of hyperscale and high-performance computing at NVIDIA, and we're discussing mixture of experts and why it's become the architecture, as it has been for a while, but now getting public prominence, if you will, the architecture behind so many leading frontier models and what goes into not only architecting and training the models, but the infrastructure that really makes them hub.

22:12NVIDIA AI Podcast Host:And Ian, I wanted to ask you, you talked about this a little bit, as I said, with, you know, NVLink and all of the technologies you kind of alluded to as you were describing the MOE architecture. But what is it specifically about these NVIDIA systems that make them such a good and such a unique fit for these complex MOE models and are able to achieve, as you just described, you know, this lowering cost of intelligence measured per token? Yeah, it's an interesting and understandable, it goes back to the original idea about having experts. We're reducing the cost per token by not turning on every neuron, but only turning on the ones we need.

22:51It's a cost savings. And we talked about LAMA, the 405 billion parameter LAMA model, you know, that in order to use it, you got to activate all 405 billion of those neurons, even though they're not all needed. Look at GPT-OSS, it's 120 billion parameters, still a lot. But you only need about 5 billion parameters. So it is smart and is a cost-saving measure. Only it is 5. She also notices, though, that's like a 10x less, actually more than 10x, 1 % of the number of neurons we're actually doing math on. The cost isn't, unfortunately, on GPT-OSS, it's not 1%, actually. It is X-factor slower. It's about 3x less cost.

23:36but it's not, you know, 1 % less cost. Sure, yeah. There's a hidden text to MOE and it's all about how those experts need and need to communicate with each other. In order to get MOEs to run efficiently, those experts are all doing their math very, very, very fast and they all need to communicate with each other very, very, very quickly. And one of the challenges with MOEs is as we go and get sparser and sparser and sparser, which makes the models more and more valuable and we're saving more and more cost, is can we make sure that all that math is happening and all those experts can talk to each other without ever going idle, without ever waiting for a message.

24:20You're buying those GPUs, you're paying for them so they can do the math they need to do. Not to sit around and wait for someone else to send them something. Or worse, the network that connects all these GPUs gets gummed up and now everybody's sitting idle and that's going to go straight to the bottom line of the cost. So that's the key part, and the hidden cost anomaly is communication. We've looked at, you know, can we make it work with just point-to-point? Like maybe I can just connect this GPU with this GPU and this GPU with that GPU. It'll be a much lower cost to actually just directly wire them up.

Read the full transcript

24:53But there's a limit to how much I can do that. If I take one GPU and I connect it to four, well, this GPU now, its IO is split four ways, and I can only do that so far. And even with our hopper systems, we had eight. And there was an Envy switch chip. We built another chip specifically for this. But we can't scale beyond that eight because that's the chip. So if you have point-to-point or a Taurus-like network, you're fundamentally limited by how much MOE, how cheap you can make those tokens. Because the hidden cost of MOE is communication. And if you try to go bigger than the neighboring a point-to-point connection or some kind of loop or message passing thing or use a fabric like Ethernet, they weren't designed for this.

25:40The best answer is no compromises. I want this expert, this GPU, to be able to talk to every other expert at full speed, no limitations, no worry about congestion. I need a network. I want to connect these things so there's nothing blocking. And that's what MVLink is. In fact, that chip that we built is specifically designed to make sure that every GPU and all of its terabytes a second of bandwidth can talk to every other chip at full speed and never compromise on the maximum IO bandwidth we can get out of every GPU. We did that with Hopper with 8-Way and one of the big innovations, and obviously it took a lot of engineering to make that 72 racks, every one of those 72, every one of those GPUs at full speed, no constraints.

26:27And you can see that taking off. You can see the benefit that allows people to go even further and build even bigger models. The Kimi K2 model is even bigger than the GPT one. We now have open source trillion parameter model, Kimi K2, yet it only uses 32 billion parameters when you ask it a question. That's like a 3 % activation of the brain. But it's incredibly complicated. 61 layers, over 340 experts. They all got to talk to each other. And as a result, we now have open models that are truly in parameter scale levels of intelligence. And the cost is all comparable to what, and even lower than what we could ever possibly have with a fully dense model.

27:15It's possible because of that emulating connectivity. NVIDIA is committed to it. Let's keep going down that path, build. We have some of the world's best 30s engineers, signal processing engineer, wire engineers, mechanical engineers, to make all that work without having costs explode and make it all connected. Every one of those GPUs, by the way, is connected with a copper wire to one switch to another switch. There's a reason why it all sits in the rack, because we're running at 200 gigabits per second on every one of those wires. It's PAMP4 signaling, so it's like four bits per wire. It's a 0, 1, 2, 3, and 4, not a 0, 1.

27:49We've gone past the binary at this point. And it's going so fast, it's actually, its wavelength is about a millimeter, I think. So, you know, we're pushing the limits of physics, keeping it all nice and tight, and also doing everything in copper for low cost. We're super happy with GB200 and what it's been able to do for inference and just keeping the cost and driving the cost of tokens down, down, down while intelligence goes up, up, up.

28:21NVIDIA AI Podcast Host:So is this getting into what we call extreme co-design? Yeah. One of the joys of working at NVIDIA is that we're the one company that works with every company in AI. Right, yes. And, you know, we work with them in building their data centers, in getting the latest GPs to them, and explaining the MVL72 architecture, in building and help build a lot of the software that they use. We have teams working on PyTorch, on JAX, on SGLang, on VLM, and all the software that's out there. And as these model makers are building new models or pushing the limits, some inside NVIDIA actually now, but all around the world, we can co-design with them how to take the maximum utility out of those 72 GPUs to manage that hidden cost of communication, to make sure every GPU is running at 110 % on computing on the fewest possible neurons and doing that seamlessly and incredibly fast.

29:17All the while, thinking about the next model. What's that next GPT, that next vision model, the next video model, the next Sora? And making smart decisions about how to add more bandwidth, more communication, more NVLink, and the right kind of floating point. And all doing so without blowing out cost or blowing out power and keeping and leveraging all the work that they've done up to date so that it can be applied moving forward to the future. This is the extreme co-design that we do at Amidia. And some of the folks that I get to work with and probably watching this get to enjoy. And we work really, really hard to continuously work on performance, not just to have the fastest and be the fastest, but also to reduce the cost.

30:04because you talked about tokenomics. If just our software alone could increase performance by 2x, you've now reduced the cost per token by 2x direct to the user and the customer or whoever's going to deploy this AI. I was on a call this morning. We got a model from a customer. They wanted some help. We applied the latest NVFP4 techniques, the latest kernel fusions, the latest MVLink communication IO overlaps. Within two weeks, we hit 2x on their model and gave them the code back and we're not done. There's so many places where we can optimize. I think a lot of people get confused. They see a GPU of a certain number of flops and they say, oh, that's better or faster.

30:49I'll tell you, this stuff's pretty complicated. Manage and run 72 GPUs with 348 experts and all the different kernels and all the different AI and all the different math. We didn't even talk about KV caching and reasoning models and all the tricks and techniques. That's an end-to-end problem. It requires extreme co-design between the hardware, what's the art of the possible, the model builders themselves, and the dense and deep software stack that run on it. NVIDIA actually has more software engineers than hardware engineers specifically for that process.

31:19NVIDIA AI Podcast Host:Right. Yep. So to kind of zoom out for a second, because we've been talking about and kind of get hearkening back to what you just said about, you know, thinking about what's next. We've been talking about MOE in the context of language models predominantly now. And the GB200 NVL72 is really well-suited to that architecture. But is there a risk of focusing too narrowly on this single model trend of MOE? What happens when we get sort of beyond MOE? What happens? Is the architecture still well-suited? Is the cost of tokens still going down? How do you think about that going forward? and how does the design that NVIDIA has today, how is it ready for whatever the next trend might be?

32:03Well, there's one clear trend in AI is that intelligence creates opportunity. As the models get smarter, as they start to learn new things or as they specialize in certain areas, they create opportunities to advance that industry, that science, that application, or just make computers more productive for you and I every day. Yeah. And in order to do that, we need to make the models smarter themselves. We need to use techniques like reasoning, which is only going to generate more tokens. And the only way to advance the state of the art of AI, well, there's lots of ways. One way NVIDIA can help is just reduce the cost of tokens.

32:45And doing that MOE, it's just an optimization technique. If you don't need all the neurons, don't waste time computing on them. That's an idea. That's not unique to LMs and chatbots. That's just a good idea. So we see it made me to realize in different ways and how these networks and experts want to communicate or the shape of the models are actually diversifying in lots of ways. There's lots of different techniques. Mixture of experts is certainly one of them that will stick around for a while. There's lots of other hybrid approaches and other things that people are talking about or trade-offs that you can make in order to reduce cost.

33:23But we see MOEs happening not just in chatbots, but similar sparsity, MOE expert applications being done in vision models and video models. As the models are expanding into science and not just generating tokens, which turn into words that you and I talk about, but work on proteins or working on material properties or understanding or working on things like in robotics or path planning or logic or business applications, all of those will benefit from having a large intelligent model that can be sparsely optimized to only use and leverage the part that is needed for that particular question in that particular use case.

34:01You can always go back down to the squirrel detector in a doorbell, but there's usually a benefit to having a model that is actually able to reason about or has some multimodal aspects. Maybe listen to what's going on, see the things around it. and be able to make intelligent decisions smartly. That is going to continue to grow. And NVIDIA is not just working on MOEs. We've got lots of different irons in the fire. There's lots of different models. The models are diverse. I get to work in HPC as well. The whole supercomputing community has now embraced AI, building all sorts of models for simulating physics and simulating weather and things that look nothing like chatbots.

34:43But they're going to use MOEs. They're going to use every trick in the book because the opportunity is huge. The ability to revolutionize biology, to do drug discovery for cancer research alone is an investment that the whole world's making right now. And they can take these ideas and take our platform and apply them to their domain, their problem, to take an open source model or a general model and fine tune it to be a science model or an application-specific model or a business model, that is possible because they're starting from a really intelligent model that can learn or be used to teach another model to make things possible.

35:27So I'm super excited about MOEs. I'm super excited and will continue to work on reducing the cost per every token. And while that may make our technology bigger, smarter, more complicated at times, and will make it more expensive, it is going to deliver X factors in capability, improvement, intelligent, and as a result, dramatically lower the cost per token.

35:52NVIDIA AI Podcast Host:Ian, for listeners who want to dive in further, we could talk about this all day, but you have things to go build and customers to take care of and all that good stuff. Where can listeners go online? What's the best place to start to dive into MOEs, to the infrastructure you've been talking about, to any and all of it? I check out GTC. You know, one of the things that's, we started this conference a few years ago, well, maybe now over a decade, I guess. I was there for the first one. It's called the GPU Technology Conference. Right. It's not a business conference, although obviously many business people show up.

36:26It's not a demo conference. It's a developer conference.

36:29NVIDIA AI Podcast Host:Yeah. And if you want to learn more, go check out GTC. We put all the presentations online. Jensen's keynote is wonderful. He has a, he'll explain it even better than I can. and you can watch we actually do a few a year now I encourage you to check out GTC go see the old ones and if you're going to be in San Jose in March please come and check it out and attend there's tons of sessions at every level from beginner to deep dive if you want to go down to the hardware all the NVIDIA experts will be there all of the different developers are going to be there it is kind of the go-to place to go learn and also present your work on what you can do with GPUs and the state of the art of AI check it out Perfect.

37:10NVIDIA AI Podcast Host:Ian Buck, again, thank you. And, you know, for what it's worth, Jensen's an amazing presenter. You did a great job explaining all this. So we appreciate you taking the time. And as always, all the best to you and your teams on continued progress. Thank you.

37:55¶¶

38:09Thank you.

From the publisher

Discover how mixture‑of‑experts (MoE) architecture is enabling smarter AI models without a proportional increase in the required compute and cost. Using vivid analogies and real-world examples, NVIDIA’s Ian Buck breaks down MoE models, their hidden complexities, and why extreme co-design across compute, networking, and software is essential to realizing their full potential.  Learn more: https://blogs.nvidia.com/blog/mixture-of-experts-frontier-models/ 

More from NVIDIA AI Podcast

All 115 episodes
Lowering the Cost of Intelligence With NVIDIA's Ian Buck - Ep. 284NVIDIA AI Podcast · 38 min
Listen in VO