The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

3 Aug 2026 · 1 h 41 min · 38 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Inference engineering for production LLMs at Baseten—how long-context queries are routed and executed on GPU, how speculative decoding and KV-cache disaggregation improve tokens/sec, when to switch from public “paper token” APIs to dedicated deployments, and how structured outputs/tool calling are made reliable. They also discuss what it takes to add support for new model releases (quantization, speculator training, runtime integration), plus quality/fidelity and debugging issues like token looping.

Guests and backgrounds

  • Philip Kiely: Inference Engineering author (“Inference Engineering” book), previously involved with Baseten; focuses on production inference performance and model support work.
  • Ali Taha: Baseten; explains speculative decoding, structured output/tool calling constraints, and vision retrofits.

Key claims

  • Long (200k-token) requests use cached input-aware routing, disaggregated prefill vs decode, and speculative decoding with a coding-oriented draft model.
  • Dedicated endpoints are needed for reliability and traffic-specific speculative decoders (shared endpoints can’t assume the task).
  • Tool calling failures are often training/formatting issues (JSON/tool arguments not closed correctly), not sandbox escape.
  • Structured output can be enforced via constrained output formats/state machines (BNF/grammar-like ideas).
  • Quantization is the main lossy step; fidelity is measured as closeness to a “golden” reference implementation (logit distribution/KL divergence).

Notable examples

  • Speculator trained on coding vs “Harry Potter” changes draft acceptance.
  • Vision retrofit: grafting a Kimi vision encoder + projector onto GLM 5.2; projector trained with image-question datasets (e.g., “mountain” → targeted Q&A).
  • Looping bug mitigation: detect repeated tokens (e.g., “S”) and reprocess/retry; non-determinism can come from backend kernels/KV-cache transfer races across clusters.
  • Model support “inference war” example: GLM 5.2 benchmarks improved provider TPS (e.g., 70→150 tokens/sec).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Long Queries in Inference

1:06 to 3:08

Explore the process of handling long queries in Base10's inference system.

“But before we get into all that, I want to start off with a fun question for you.”

Cache Routing and Model Optimization

3:08 to 6:46

Delve into cache routing strategies and model optimization for efficiency.

“I'm assuming that we're talking about the public model APIs.”

Speculative Decoding Explained

6:46 to 8:33

Learn about speculative decoding and how it speeds up the inference process.

“And so as a result of that, it didn't see the result and just hallucinated the result as it decoded.”

Challenges with Tool Calling

8:33 to 10:19

Understand the complexities involved in tool calling and JSON outputs.

“So it's hard to parse something or validate something while it's being streamed.”

State Machine Approach for Structured Output

10:19 to 12:34

Discover how state machines can help constrain output formats effectively.

“But, like, with the web training, shouldn't be that much of a difference.”

Updating Inference Systems for New Models

12:34 to 14:00

Examine the processes involved in supporting new model releases in inference systems.

“The challenge is, you know, every inference company is going to have our own proprietary stack.”

Training Speculators for Inference

14:00 to 17:32

Learn about the process of training speculators using model weights and infrastructure.

“So we can get public data sets that are representative of that kind of traffic and train general speculators.”

Integrating Vision into Language Models

17:32 to 22:28

Discover how vision encoders are integrated into language models for enhanced capabilities.

“So we've covered Houtian before who was the author of the lava paper that did this a while ago.”

Addressing Inference Challenges

22:28 to 28:00

Explore the challenges in inference such as model quality, race conditions, and quantization effects.

“It's an inference problem, to be honest.”

Understanding Quantization in AI Models

28:00 to 29:45

Learn about the implications and techniques of quantization in AI models.

“Like if I'm a consumer and I'm using Amazon's endpoint, for instance, and I'm using Kami, I'm like, oh my God, this is bad.”
Show all 38 chapters

Research Insights on Quantization

29:45 to 31:44

Discover the research findings on quantization and its effects on model performance.

“You're compressing the data from occupying 16 bits to occupying 4 bits, for instance.”

Performance Optimization in Inference Engineering

31:44 to 33:48

Explore the methods and goals of performance optimization in inference engineering.

“It's a fun fact, it was originally 72 pages this paper and then we decided we can't tell.”

Benchmarking and Performance Gains

33:48 to 36:26

Understand the challenges of benchmarking and how to achieve significant performance gains.

“And we're at the beginning of the same type of thing.”

Quantization Techniques for Model Deployment

36:26 to 40:08

Learn about effective quantization techniques and their implementation in model deployment.

“You have done all of your quantization work.”

Dynamo and Performance Tools Overview

40:08 to 42:00

Get an overview of the Dynamo toolkit and its role in enhancing performance.

“I can do it for you in case I get it wrong.”

Exploring Disaggregation and Techniques

42:00 to 43:16

Learn about various techniques in AI, including disaggregation and their relevance.

“I would have said it comes with a set of defaults that you can then swap out.”

The Evolution of Speculative Decoding

43:16 to 46:04

Explore the nuances and complexities of speculative decoding in AI models.

“As I mentioned in my AI engineer talk, which is kind of the first public addendum to this, the speculation space has moved much faster than everything else.”

Inference Engineering Challenges

46:04 to 48:19

Understand the differences in inference engineering for local versus data center AI.

“And so if you have sort of like infinitely reclusive speculators, you're adding quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.”

Model Parallelism and Active Parameters

48:19 to 51:12

Discuss model parallelism and how active parameters affect performance.

“And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance.”

Understanding Tensor and Expert Parallelism

51:12 to 53:37

Learn the differences between tensor and expert parallelism in AI models.

“Expert parallelism you can only do with MOE models.”

Mega Kernels and Inference Performance

53:37 to 56:00

Discover the impact of mega kernels on inference performance and their theoretical benefits.

“What's the magic numbers that we need to?”

Challenges of Mega Kernels in GPU Computing

56:00 to 58:12

Explore the complexities and limitations of using mega kernels for GPU computing.

“So I need to know what the partial result was from GPU 2 and what the partial result was from GPU 1 in order to be able to do the soft relax in the next stage.”

The Evolution of Hardware Launch Cycles

58:12 to 1:00:00

Discuss the implications of hardware launch cycles on inference engineering and performance.

“Can I speculate about Rubin for a minute?”

The Future of GPUs and ASICs in AI

1:00:00 to 1:02:55

Analyze the trend of GPUs becoming more specialized and the rise of ASICs for AI workloads.

“which was the same thing that made Blackwell so good.”

Model Longevity and Efficiency in AI

1:02:55 to 1:09:09

Examine the balance between model updates and the efficiency of existing AI models.

“If it's burned into the chip, the chip's useless in like a month or two.”

GPU VRAM and Context Length Considerations

1:09:09 to 1:10:00

Understand the importance of GPU VRAM and context length for large AI models.

“by like my 5 ,000 stakeholders like I'm not It runs a batch job every day and I like the results the results are predictable Yeah It doesn't make sense to keep using them like stuff gets sparser, cheaper, better.”

Hardware Limitations for Large Models

1:10:00 to 1:12:08

Discussion on the required hardware specifications for running large AI models.

“NVFP4, 2.8 trillion parameters, 1.4 terabytes.”

Challenges in Video Diffusion Models

1:12:08 to 1:14:18

Exploration of the complexities and disparities in video diffusion compared to LLMs.

“And so, for example, when DeepSeek R1 came out, it was, you know, it was 671 billion parameters, which at the time was really huge.”

Token Limitations in Video Generation

1:14:18 to 1:16:36

Examination of the token limitations and their impact on video generation quality.

“causes less open-source checkpoints to be released.”

Pros and Cons of Autoregressive Video Models

1:16:36 to 1:20:24

Explaining the trade-offs in using autoregressive models for video generation.

“Or you move towards autoregressive video.”

Differences Between Audio and Video Models

1:20:24 to 1:23:50

Discussion on the differences in autoregressive and diffusion methods between audio and video models.

“both in, if you sort of naively construct a video generation model as simply generating a linear sequence of frames, you can't then go back in that sequence and fix something to make the whole thing consistent.”

Current Trends in Diffusion Models

1:23:50 to 1:24:00

Overview of the latest developments in diffusion models for text and images.

“It's not a perfect split, but that's the broad categorization I use.”

Exploring Diffusion Models for Text and Image

1:24:00 to 1:27:42

Learn about the advancements in diffusion models and their applications in text and image generation.

“Nano Banana and GPT Image are auto aggressive image.”

The Evolution of Inference Engineering

1:27:42 to 1:30:09

Discover how inference engineering has shifted from simple execution to integrating training and optimization techniques.

“Now it looks like people are using inference more and more in post-training.”

Future Trends in AI and Inference

1:30:09 to 1:35:09

Discuss the future of AI, including scaling challenges and the potential for new modalities in inference.

“You know, any kind of dynamic adjustment is going to be a static configuration across, you know, your exact config, across your speculator, across that kind of thing.”

Continual Learning in Inference Systems

1:35:09 to 1:38:06

Examine the concept of continual learning and its implications for inference systems and knowledge management.

“that AI has worldwide compared to some of the more mature technologies, both on consumer and business, it's pretty clear that there could be multiple 10x's more of demand.”

In-Depth Discussion on KV Cache and Learning Models

1:38:06 to 1:39:55

Explore the mechanics of KV caching and its implications for continual learning in models.

“It could either be that the model learns and so it's continuously pushing its new knowledge into its weights.”

Reflections on the Conversation and Future Prospects

1:39:55 to 1:40:39

Reflect on the insights gained from the discussion and the potential future of inference engineering.

“It's a relatively recent blog, so people can go see it.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03Philip Kiely:Okay, we're here in the studio with Philip, an old friend from Inference Engineering, the book, as well as Base 10 and everything that you've done, you and I have done before, as well as Ali. Welcome. Pleasure to meet you. Waterloo intern. Waterloo intern, always.

0:17Ali Taha:When did you get Waterloo intern as a... As a handle? I think the rebranding happened like mid-March. When I saw it was open, I was like, I have to take it for grabs.

0:25Philip Kiely:The problem is that Ali is really good at his job and is not going to be an intern much Longo so we have to figure out you know who's going to get the handle. Pass the torch over. Oh okay it can be like you just pass it to

0:37Ali Taha:another Waterloo correct? Another Waterloo Inter. Inter. You've got to get an Inter from Waterloo. But it could come from

0:46Philip Kiely:Base 10 so it's like wherever Base 10 gets from Waterloo has the title of Waterloo. They have to pass you. You either get it or you're out. You should also do like a big graduation ceremony where you change the handle. I mean, you guys are good at ceremonies. Clearly, you know, we had a nice launch of the book, very successful. But before we get into all that, I want to start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say 200 ,000 tokens into base 10's inference? What's the process of query through GPU, model, routing, balancing, all that?

1:23Philip Kiely:What is all the stuff that we don't think about? with a long query specifically, the first thing that I'm going to ask is, have you sent me this query before, or at least part of it? And I really hope you have, because it's going to be a lot easier for me and a lot cheaper for you. So the first thing that we're going to look at is some kind of cache away routing where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with number one available prefill workers and, number two, ideally some cached input already there so that we can skip pre-fill on at least part of these 200 ,000 tokens.

2:03Philip Kiely:If you're doing 200 ,000 tokens, it's probably coding or a multi-tone agent or something where you would expect to have that cached. If you don't, we're going to have to send it to a pre-fill worker. We've, at least on certain models, disaggregated pre-fill and decode. So you're going to have one set of GPUs that's solely going to process the input, create that KV cache, and get you your first token. And then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some kind of speculator model in front of that.

2:39Philip Kiely:I'm going to assume that you're doing coding. And because of that, our speculator model, which assumes you're doing coding, is going to have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's going to be slower. And then we stream that output to you and account for it, charge you some number of, couple of pennies, and say, hey, would you like to send another one? Except base 10 doesn't charge by pennies. Well, yeah, we charge, I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.

3:18Philip Kiely:Yeah, I mean, one of the key differentiators when I was talking with BaseSend initially was that actually people who want very, very high volume just need to rent by the box because then it's up to you to figure out how to saturate the box.

3:31Ali Taha:And more often than not, it's like way cheaper if you're pushing like millions of tokens per hour, if you just pay per hour instead of paper token.

3:37Philip Kiely:Yeah, they do. I think that we've increasingly seen a lot of demand for the sort of paper token APIs just because everyone wants to try open models. And then once they find a use case that's really sticky, then they move over to dedicated. Is there a best practice on when it's time to swap over? A couple of reasons.

3:55Ali Taha:Yeah, reliability, that's a big one, right? Like if they have a very specific use case, they want you to train something specifically for them. Like they want their own spec deck, for instance, for their own traffic. Spec deck is speculative decoding. Speculative decoding, yeah. Sorry. Basically, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach this little parasite, site, like this layer that goes on top of the model. And this model just has to predict, it does three very fast, ultra-aggressive forward passes, and it will predict three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not.

4:34Ali Taha:And then you accept them or you reject them. Now, this draft model is traffic-specific. So if you, like Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm going to accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're in a shared endpoint, because I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. Like, we don't know. Also, there's a thing in the book that mentioned that they really cared about a specific threshold, chapter four, I think.

5:05Philip Kiely:Do you remember that? Yeah, the things that you can do is you can set a specific batch sizing, a specific parallelism strategy if you're trying to optimize for throughput versus latency. you can maybe a NVFP for a quant doesn't pass your benchmarks and you want to run a model at higher precision, you can do that there's just a bunch of reasons why you might want to have your own endpoint and the biggest one of course just being like you don't have to deal with someone else throwing 100 million of tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users I think one thing that is a classic journey, you know, it's basically Vivo is asking what happens when you type Google into the browser.

5:51Philip Kiely:Tool calling, is that just, you know, you're generating JSON or is there more

5:56Ali Taha:complication beyond that? Certain customers that we have, they have their own post-trained models and so they demand tool calling that's not just like, you know, parse a file or, you know, go find the weather. It's something that's very specific and you have to do post-training on this and if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own sandbox. It's not like it's going to use that tool calling to escape a sandbox.

6:28Ali Taha:It doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling, which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't close the end of the request in a very certain manner. You end up with a model that did the tool calling and the thinking. And so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes for one.

6:56Philip Kiely:Yeah, that's a challenge on the training side. And then on the influence side, there's work that you can do to scope the possible output. So we published this actually at this point close to two years ago. the solution to this problem, which is you basically make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember back in the day... Yeah, the specific grammar is... Yeah, exactly. GML had this thing. Yeah. So it's like the old school, like make sure this is only JSON return, only JSON or my grandma's going to die type of prompt.

7:39Philip Kiely:Is it BNF grammar? At some point in the opening, I had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar back as no. In our inference system, it's just a specified output format, and you get the guarantee that your output's going to be structured along that format. And so applying that to tool calls can help cut down on, obviously you can still call the long tool or call no tool. it doesn't solve the certainty problem, but it at least solves the output structuring problem within tool calls. And MCP is just another form of tool, right? Yeah, exactly. There's no special thing there.

8:16Philip Kiely:The thing I'm always explaining to people is the LLM is actually not capable of doing anything. It's only capable of making suggestions of what to do. And then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs. Part of the fun stuff is, you know, this is solved outside of tool calling too like in an agent loop if the output is not correct or you're right like reasoning, tool calling was done in the reasoning so you'll just be like oh I don't know what to do, let me just try again and you know it might get there after a few tries and on your point of training sometimes this is harder in smaller models so you don't have the same exact quality output when you just swap from a big model yeah I will say that before we, I think we need to go back to inference engineering proper But I had expected that something would replace JSON because it's hard to stream JSON because JSON must be complete and you must have open and close brackets and everything.

9:14Philip Kiely:So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's basically something like Toml, something like YAML. But JSON seems to be dominant still. The JSON output's not that long, right? I guess you could have a long, because tool calls also contain the arguments in them, and perhaps for a certain tool, you might pass a very long argument. But my impression of the sort of median tool call is that it's a relatively small number of tokens. So I would expect that speculators are generally fairly good at something as formatted as JSON.

9:54Philip Kiely:And so you would have a pretty fast decode step there, and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.

10:01Ali Taha:You're also bounded by the software that the model is going to integrate with if the software is built with JSON for the tool calls or if the company that you're, you know, if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to, like, you know, change their software and say, like, yeah, this is going to be better for the model. But, like, with the web training, shouldn't be that much of a difference. Also more profitable if it outputs more tokens, probably.

10:25Philip Kiely:Depends on your business model. You know, it really depends. But I will say that, you know, As a writer, experience a lot with AI-generated output, I do try to move from text to JSON text, which is very long JSON, right? There's paragraphs in every field because I'm trying to structure it. I want you to first make factual statements, then make opinions, then make bullet point summaries, have dates, have entity references, have your sources for references, all these things. Anyway, so these are things that I think people who really, really experiment with structural output have to really care about.

11:00Philip Kiely:But let's sort of recurse up the stack a little bit. Before we started recording, you actually mentioned something which is really cool, which is that there's a lot of inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM 5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to 5.2, like, you know that you've supported them before. Is it that much work? It's a lot of work. Yeah. Okay, so a lot of people, all you guys, whenever a new model launch, people rush to say, oh, Hugging Face supports this, Fireworks supports this, Space 10 supports this.

11:36Philip Kiely:And I'm like, yeah, of course you support it. But what goes into that? I think it's more than just supported too, right? It benefits the consumer a lot. I think it was with Kimi K2.5 or GLM 5.2, the latest, there was sort of an inference war, right? X providers at 90 tokens a second. The next day we're at 150. I kind of kicked that off with GLM 5.2. I wrote a Twitter article about, it got like half a million views. Based on being number one. Yeah. Or artificial analysis. Yeah, which got everyone really excited about, hey, how can we, you know, bench max a little bit further. And there's a difference between support the model, as in like, I can make a token out of this model, and support a model as in, I have a production-ready API from this model.

12:24Philip Kiely:getting to the point of I can make a token out of this model is not that hard because generally the open source inference engines your vlms sglangs of the world oftentimes even receive weights ahead of time maintainers do or the people making the model merge prs to ensure support so you generally can you know just kind of get it working on the standard open source stack without too much pain in most cases. The challenge is, you know, every inference company is going to have our own proprietary stack. You know, some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff.

13:10Philip Kiely:Sometimes you get lucky, like K25 to 26 was like pretty similar. Yeah, it was pure continued post-training, if I remember correctly. Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from, generally these models are not released in NVFP4 and we want them to be in NVFP4 for maximum blackwell compatibility. So we have to perform that quantization and calibrate the quantization to make sure that we're not causing any kind of regression in the model's intelligence. And then we also have to train the speculator as we've talked about.

13:47Philip Kiely:Generally we have, obviously we have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us. But we know what's popular. We know that coding use cases are popular. We know that agentic use cases are popular. So we can get public data sets that are representative of that kind of traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself, because you're getting hidden states out of the model from running inference on these specific prompts. and that is the training data you use to create the speculator.

14:23Philip Kiely:So there's that process which you need the real model weights for. And then there's, of course, just the process of, you know, standing up all the infrastructure behind it, loading all this stuff in, testing it. And then when there's a new model with a newer architecture, I think that, like, obviously the DeepSeq models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model over model. But every new model has something. I mean, Kimi K2 had, oh, sorry, GLM 5.2 had the DSA. Which was brought from DeepSeek. Yeah, yeah. So you can copy-paste it. I don't know how this works.

15:02Philip Kiely:So we had to build support for that into our runtime. And you're right, it actually is really interesting the way that all of these open-source labs borrow from each other. For example, GLM 5.2 doesn't have vision. So something that Haley, a guy on our team, if we could take a look at this, he kind of grafted the Kimi vision encoder onto GLM 5.2. We'll be training the projector. Exactly. So if you think about the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information. And then there's the projector, which kind of like - You can't say it in space, it's okay.

15:41Philip Kiely:And then there's the projector that maps it onto the model itself, and then there's the model weights. you don't want to mess with the model weights because you want a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Harry started with just a projector, which is only a handful of millions of parameters. Can you show the training room?

16:04Ali Taha:The way it groks is very, very interesting.

16:07Philip Kiely:Maybe, Ali, you should take it from here. You've got a better understanding of this than I do.

16:11Ali Taha:Yeah, you can see the way he trained this is really, really cool. At the beginning, he was training it using just like, here's a picture of a mountain. Can you describe what's in this mountain? And that calls it just like the first, the first, you know, learning loss. Like here, you can see this, all we're trying to teach it is to translate the encoded, like it's already taken the encoder from Kimi K. It's taken the image.

16:30Philip Kiely:Yeah, frozen, frozen. Frozen, frozen. Exactly.

16:32Ali Taha:So the understanding, the brain is frozen and the eyes are frozen. It's just, we're trying to align, interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, oh, can you describe what's in this image? And he's like, oh, it's a mountain or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it?

16:59Ali Taha:All that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer a question, answer a question, answer a question, over time, like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, who is this? Maybe it doesn't get it, but it will say something like, this is Albert Einstein. Like it still understands this is a scientist who is a man who has, you know, done significant achievements and all that stuff.

17:30Ali Taha:So that's like really, really cool.

17:31Philip Kiely:Yeah. So we've covered Houtian before who was the author of the lava paper that did this a while ago. And I think that's very foundational work for anyone who hasn't done vision work before. Same with Clip and MetaClip where you go from just captioning to building out questions off the image and how much better you can get performed. Yeah, but what's so exciting about this is if you look at a model like this, now obviously this is a little bit more of a research project. It got to 56 % on MMLU Pro, I think, so not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM 5.2 quality.

Read the full transcript

18:09Philip Kiely:If you don't have an image, it'll just behave exactly the way it used to. which in the inference code you literally do not include the other part, right? Yeah, I mean you would just skip the encoder if you don't have an image input. Just confirming. Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically like less than a billion parameters. Yeah, I mean there's a little bit less standardization among vision encoders so the sort of support matrix can be a little bit sparser But overall, yeah, it's a pretty minor component of the overall system.

18:47Philip Kiely:And ultimately, what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model. And that's, I think, a lot of the power and beauty of open source is that you can take all of these different components and combine them together into a system that's better than anyone can be individually. people used to say that you would also do franken merges where you would take like layers from each model

19:11Ali Taha:does anyone do that anymore? to your point previously when you were mentioning like the work that goes into supporting a model when it first comes out like GLM 5.2 or Minimax M3 or whatever the case is sometimes you do have to switch out some things like for instance the Minimax M3 head uses full attention and with full attention you end up with this like insane bottleneck inspector because you're doing auto-aggressive token generation for three tokens and you're doing this like O of N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K.

19:48Ali Taha:So we find it better to like, okay, we're going to replace this layer with a layer from another model that's using like GQA, for instance. And then just for the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed actually if a layer is like inefficient the training just becomes the challenge like how do you ensure that you train it properly which again to earlier points is like the mesh between training and inference as in like you need very good training in order to do fast inference that's like I feel like more and more becoming true

20:21Philip Kiely:anything else on the support side when you say like get it to fully production ready yeah I think that there's also a question of just we can test a model to a pretty extensive degree but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with GLM briefly where we had some mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Once you expose an endpoint to the real world there's going to be so many more varieties of things given to it that you're able to discover and patch things.

21:08Philip Kiely:So it's not just a day zero process. It's then for the first week, for the first month, if a model remains popular. How do you both fix bugs and then continue to push the envelope on performance? What do you mean you don't want your model outputting SSSSS?

21:25Ali Taha:Is there loop detection on that stuff, by the way? It still happens quite a lot, which is surprising. We have, like, in our endpoint, like, if a model was to output the same exact token, like, four plus times, we just call the generation and we say, like, oh, sorry, this should, like, try again. Or, like, we will re-process the request. Because we know then, like, if it's four times the same token, it's probably collapsed. Yeah.

21:46Philip Kiely:Is there a way to opt out in case I really actually want that?

21:48Ali Taha:You're actually... I think there's a way that we have to handle it. I'm not exactly certain. I feel like in certain models, like, when they output something, like, you can imagine, like, a table, for instance, and so they want to draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters. So we only do it on like certain, like S is the most common almost, GLM5-2. And I think it was DSV4 as well? Like you'd just have like looping issues where like... Yeah, is there

22:19Philip Kiely:something special about S?

22:21Ali Taha:No, just randomly. It just seems to be the one token. Yeah. And it's only temperature zero or even at other temperatures? It's an inference problem, to be honest. Oftentimes, the image you run, like NVIDIA will release an image, for instance, and if we will upstream the changes from the latest RTLM image into our stack, we'll find that it fixes it. Or oftentimes, this will only happen in an inference engine that you're using, like SGLang. But if you were to switch to VLM, that isn't the case. So it seems to be an extremely non-deterministic kind of software issue, not really a model issue. It's not like a weights problem.

22:56Ali Taha:Like, we'll say, oh, it's a problem with the quant. We did p2q wrong, right? But that doesn't make sense because the same exact weights used with a different inference engine does not repeat the problem. And sometimes it's the kernels that are being used in the backend have like these very subtle, sometimes raised conditions where if you were to use this model hosted on one cluster, you will never get this problem.

23:18Philip Kiely:Oh my God.

23:19Ali Taha:But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node-to-node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not going to be hosted on this cluster. We're going to host it on another cluster because that cluster exposed that problem. But then it ends up like, okay, is it the software? Is it the model weights? Or is it the hardware?

23:42Philip Kiely:There is a thing about this with temperature zero still not being deterministic, right? Mostly because of hardware. Even at temperature zero, same model, you won't always get the same output. But I'm surprised by the race condition one because I thought PyTorch was a graph that guarantees that you at least execute things in the right order. Well, they're not true.

24:03Ali Taha:I guess I'm not saying that this is a risk. Well, you have things like PGL optimizations where you can start a kernel before the end of the previous kernel. And that's like you want to do that. It's like pipelining. Exactly, exactly. But you don't do it cleanly. You overlap a little bit of the execution. No, I guess it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to be very fast. If you don't test it extensively, you'll have certain threads access data points from registers before they've been written to by other threads, for example.

24:36Ali Taha:Because like your barrier is wrong or your synchronization is wrong. But yeah, like the testing itself is very, very difficult. And there's no like borrow checker for, you know, like Rust.

24:47Philip Kiely:Like if you're trying to have like memory safety, it sounds like a comparable problem.

24:52Ali Taha:Well, I guess, but you're working in Kudo and VideoGPU. You just need a higher level language, like modular. Maybe that's what modular is supposed to do, I don't know.

25:00Philip Kiely:How do you see keeping quality of the model? So you talked about all these steps of, okay, you got to do quantization, train your own speculative decoder, run on different hardware. Looking at other model providers, okay, you kicked off an inference speed race. On the consumer end, what goes into keeping quality the same across them, right? Sure, you can run benchmarks, but how do you determine how much quantization are the standards? What actually goes into? There's a few things on quality. Most inference optimizations are lossless. KV caching, for example, you are just preventing recomputing the same values.

25:39Philip Kiely:Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format. Number two, which parts of the model you choose to quantize, which layers. And number three, doing a lot of calibration on the quantized weights to ensure that you're preserving all the outliers. There's other sort of tricks that you can do, though. A big one is long context. Because one thing you asked right at the beginning is, oh, what's going to happen if I send a 200 ,000 token request in? So obviously with a long input sequence, you need to store a lot more information.

26:26Philip Kiely:You need to process a lot more tokens. And so even if a model has a context of a certain length, you might as an inference provider choose to build an API with a shorter context length. And of course a full length one as well. because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a sort of golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, you know, 100 % fidelity of the model.

27:12Philip Kiely:You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100 % fidelity mark as possible. And certainly our standard internally is that you should not be able to tell the difference between our API and a sort of official API. I think Kimmy in particular does a good job of vendor benchmarking here. Yes, they released an actual vendor benchmark. Exactly, yeah. Because they accused some people. Amazon? There was some provider that was not doing very well on Kimmy's benchmark.

27:50Philip Kiely:Yeah, so it would reflect we'd be putting out some points.

27:52Ali Taha:This was a long time ago, right?

27:54Philip Kiely:No, like four or five months ago. This also happened with, I don't remember which model, but they pulled out quite a few and then they started a whole chart about this. It might have been... Kimmy vendor verifier. Yeah.

28:05Ali Taha:Yeah, because you'd be pissed, right? Like if I'm a consumer and I'm using Amazon's endpoint, for instance, and I'm using Kami, I'm like, oh my God, this is bad. I'm not going to say Amazon quantized the model in a bad way. I'm going to say, oh, Kami sucks. Right? So it seems like that. Yeah, make care.

28:21Philip Kiely:Justifiably. This is probably a stupid question, but just checking, has anything improved from being quantization? Like is quantization always strictly worse?

28:31Ali Taha:Well, technically no. It's a lossy. Quantization is a lossy. It's a lossy implementation. Speed improves. It's been improved.

28:38Philip Kiely:Obviously, number eight, like... I always look for inverse scaling laws. This is something I learned from Noam Brown, where, like, things that normally act in one direction, sometimes they... Well, technically, when you run a benchmark, because these models are non-deterministic, sometimes your, you know, NVFP4 quant is, like, you know, two basis points higher than your... That's noise. Exactly. Yeah, it's within... That's why I always say within margin of error. And I actually stopped saying that because everyone assumes that what I mean is, well within some margin of error we're barely inside of that to the worst so we're saying but yeah sometimes it's just like gives you a higher output score but like Ali said that's noise to my knowledge you're not necessarily making the results better you're just trying to again like keep your fidelity as close to 100 % to the original model

29:26Ali Taha:there is to your point research that we did on MP I don't know if you are able to pull a tweet we did one of our research interns is Joshua I think it's a tweet on how we have 20 % better quantized JLM52 than NVIDIA. Essentially, what we found throughout this two-month research is, okay, quantization is a lossy. You're compressing the data from occupying 16 bits to occupying 4 bits, for instance. And so you're obviously losing some information and you're trying to minimize that. And so when I say that I'm going to quantize the model, my job becomes how do I find the layers that I can quantize and how to find the layers to not.

30:05Ali Taha:For instance, with image models, I don't quantize modulation layers and I don't quantize out projections because those two are like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so I guess to his paper, do you have the, I guess it doesn't have the, yeah, it's a long paper. I don't know if I can find.

30:25Philip Kiely:If there's a part to search or it's probably in the thread.

30:28Ali Taha:It's probably in the thread. But basically the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and two, it is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in is that you can predict which layers are going to have quantization errors that will cancel out with each other and you choose to quantize those layers.

31:01Ali Taha:And so the result of doing this mathematical quantization is you end up with a model that's 20 % more quantized than another provider. So you get 20 % more throughput of it because there's more layers that are running in NVFE4. And your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out. Like one layer scoot to the right, one layer scoot to the left, one layer scoot to the right. Your final logist distribution is more similar to the original distribution of the model. So you have better fidelity. And so the way we proved this was with KL Diversion.

31:28Ali Taha:So instead of just scoring on the benchmarks, we scored the KL diversions between the logit distribution of the quantized model and the logit distribution of the original full precision model and we showed that with this technique we get if your probability distribution on the logit switch token it wants to select is more of the same as the original model you're probably going to end up staying true to the original model so yeah so it seems like previously before this it seemed like the industry was well the more you quantize the worse it's going to be because the more loss you introduce that's not exactly not necessarily true so yeah doesn't improve it but can cancel out

31:57Philip Kiely:I think it might be this but it reminds me a good bit about pruning actually where you can prune off certain layers but very interesting, didn't know this was a whole paper you guys put out.

32:06Ali Taha:It's a fun fact, it was originally 72 pages this paper and then we decided we can't tell. So it's our 425.

32:15Philip Kiely:Still 39 pages, very substantive. We talked about evals and all these things and what's possible in terms of speed up? I guess it's probably the number one thing that people do want to care about is something that you wrote about in your post. Like official API is 70 tokens per second and you push it up to 90. Is that like a normal thing? So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time is that if you look at highly optimized domains, like say finance, if you're in finance, you measure how much better you got in basis points.

32:55Philip Kiely:It's like, oh, I got five basis points better. like 1 20th of 1 % better. That's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100%, it's 200%. So there's still probably like a lot further to go, honestly. Like you'll know that influence is pretty much solved when researchers start publishing about how they got 1 % faster at something. Which by the way, because I am from the finance background, in the 70s, that was the margin at the time. When you did quantitative finance research, you would find... And look at the 20%. Tens of the percent. Yes. Yeah.

33:30Philip Kiely:And now it's tiny. For those people interested, look up Andrew Lowe's paper. He had a really interesting illustration of quant stat arb distribution narrowing down from those kinds of 20 % differences in the 70s down to nothing today, which is very cool. Exactly. And we're at the beginning of the same type of thing. Now, benchmarking is hard. I think anyone will tell you that. And benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence links, all that kind of stuff.

34:11Philip Kiely:But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second. which is bad naming by us in the industry because there's actually two tokens per second. There's tokens per second, the throughput number, and the latency number. Like total tokens per second out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, inter-token latency, but we don't. Anyway, so you can imagine a sort of standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for a reasonable traffic profile.

35:03Philip Kiely:And we generally see the goal of pushing to 10x that. But not necessarily day zero, but by stacking enough optimizations, if you have, say, like four optimizations, each of which doubles performance, or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8x gain. That's kind of the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90. Are you saying you have done that? So let's say you have, as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10x that.

35:47Philip Kiely:So like on GLM 5.2, if you're running it unquantized, perhaps on hoppers even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're probably looking at that like 30 to 40. Do you think that's like a reasonable baseline? Right, right. To get to something like 10x, there's a lot of trade-offs that you're making. If we're running at sort of more like a 300, 400 tokens per second range, obviously you are using the best hardware possible. You have an optimized speculator. You have done all of your quantization work.

36:33Philip Kiely:You are seeing a pretty high cash hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput. But it is possible. So the spreads that you see if you like go on artificial analysis or you go on open router and you look at, you know, the worst provider to the best provider oftentimes can hit that kind of range. 10x is of course very aggressive it's often times maybe more of a 4-6x improvement but that's the kind of performance that makes us really excited is when we can get these huge gains not just go from 70-90 tokens It's also hardware dependent

37:20Ali Taha:if you obviously have a thing where you're serving it on just a node of H100s and you short the model across 4 nodes of B200s you can definitely increase the speed with just throwing more hardware at it Like normalizing for the same exact hardware

37:33Philip Kiely:and the same number of GPUs. Yeah, then you're looking at like a 2 to 4x improvement depending on the influence optimizations. So yeah, some of it's, you know, what's the car and some of it's who's the driver. If you break down the 2 to 4x, say the example is run GLM 5.2 on B200's single node, right? What's like the cost trade-off for effort to get like the last bit of juice out versus what should people just think of, right? Spectre quantization. Spectre quantization.

38:04Ali Taha:That's like 95%.

38:06Philip Kiely:And how far does that get you? And how easy is that for the average person to do? So say right now, I want to throw the weights of GLM-5.2 on a node of B200s. How easy is it to find speculative decoder model or already quantized model? How much work goes into it? If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about like what are the 2Xs we're stacking, going from BF16 to NVFP4, it's not quite a 2X, right? It's like, I think it's about like 30 to 40 % from 16 to 8, and then another 30 to 40 % multiplied from 8 to 4.

38:54Philip Kiely:So that doesn't quite get you a 2X, but like roughly a 2x. Speculator, roughly a 2x. Disag on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2x. And then you add in some, you know, double digit percent increase from having just a better run time with, you know, the latest kernels and stuff behind it. And that's kind of how it stacks up. So building each of those, like building the quantized weights is for someone who really knows what they're doing, hours to days of work, building the speculator, again, like hours to days of work, and the dis-ag setup, hours to days.

39:38Philip Kiely:Well, okay, once you have it... Once you have it, once you have it. Yeah, yeah, getting dis-ag working for the first time, I'm saying, of course, is very difficult, but the marginal implementation is...

39:48Ali Taha:If you're just grabbing, like if you are a person, like just a normal consumer who has access to like a node of B200, and you're wondering, how can I just host it myself? you don't need to quantize the model yourself. There's always going to be an open source quantized checkpoint and NVIDIA is going to push one out if no one else does. Usually the providers will have their own spec tech that they've trained as well. You don't need to train your own spec tech. You can just use that as well.

40:09Philip Kiely:Like, can me, GLM 5.2 has its own MTP. Right. Multi-token prediction. I'm just going to expect it. I can do it for you in case I get it wrong. Actually, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding. I'm actually not sure. I'm semi-confident, but someone can check. But, you know, it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference to, I want to throw this up on, you know, I want to rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind VLM.

40:48Philip Kiely:I was waiting for a mention on Dynamo. I feel like that's supposed to be the baseline that you measure against. I would think of Dynamo as less of a sort of out-of-box system and more of a toolkit for building with. So when we talk about doing KV-aware routing, when we talk about doing KV-out offloading, when we talk about doing PD disaggregation, Dynamo fundamentally is, by the way, Dynamo is an open source library from NVIDIA. We've done that part with Kyle. Okay, cool. Cool. So then your listeners know then that it supports all the different inference frameworks, and it actually is kind of multi-hardware, which is interesting.

41:28Philip Kiely:But it's just a router. It's not like an optimizer there. Yeah. All it does, like what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have KVCache on one place and you need it to be somewhere else, Dynamo coordinates Nixle for you to move that around. That doesn't mean that out of the box you just say, you know, pip install Dynamo and then you get a massive performance speed up. It's more of a developer toolkit. I would have said it comes with a set of defaults that you can then swap out. It does. If the industry at large, I think, was rolling out all of these deployments standard, then I think it would be a credible baseline.

42:16but we've got to benchmark against

42:21Philip Kiely:what we're seeing in the wild. I did want to talk a little bit more about PD disag because that's probably number three after quantized and speculative decoding. In your book though, I was just going to pull out the book, like section 522 on Medusa, 523 on Eagle, 524 on NMAM. It's 5-5 would be disaggregation. Well, no, I just wanted to dwell a little bit on the other. So what do you choose to include? what do you choose to not to include? Because there was all these other techniques, I guess. Yeah. Are these still relevant? Because I think they came out like a year and a half ago, maybe. Medusa is quite old.

42:56Philip Kiely:Yeah, Medusa's old. But is it in the book as a good, here's the baseline, like you should know this. Like I read the paper, I'm like, oh, it makes so much sense. Yeah. So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole. and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI engineer talk, which is kind of the first public addendum to this, the speculation space has moved much faster than everything else. So even at the time that I wrote the book, Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is.

43:42Philip Kiely:And now, of course, there's deflash, despark. There's newer techniques even than Eagle, although Eagle is still very commonly used. Spec-spec-spec-deck? Yes, speculative decoding. What? Can you?

43:55Ali Taha:It's a paper by Tridale, and it's basically doing speculative decoding. For the speculative decoding. Oh, my God. It's literally just another... It's like, yeah, it's almost impossible to explain it. And it seems like he got non-trivial speedups there, but it seems that the complexity with training, it's almost like in our mind at least it's almost as complex as training Gansler it's like a very delicate balance and often times it's just it's literally speculative decoding on speculative decoding

44:21Philip Kiely:speculative decoding we saw this paper it's interesting right I wouldn't even expect it to be very particular to train the naive part of me is like okay train speculative decoding it makes sense

44:33Ali Taha:the whole idea of speculative decoding is you it's almost like the iPhone auto-project version but for a normal model you're just generating three tokens and you're like, okay, I'll do pre-fill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of autoregression. So why not just have an even smaller model?

44:52Philip Kiely:I guess the other question there is, what are the size of speculators? So say for GLM... Right, it's like a billion parameters. Like for Minimax, it's...

45:02Ali Taha:Yeah, yeah, it's like one layer. It's like 1 60th of the original model, usually.

45:06Philip Kiely:Yeah. Actually, I think we should do a paper when we get back to the office. speculative, speculative, speculative coding.

45:13Ali Taha:No, it does seem like when do you stop? But then it also seems like kind of, like if you're able to train spec, spec, decode, for instance, right? Like if you're able to have a small model that accurately predicts what the intermediate speculator is going to predict, that is able to predict what the original target model is going to predict, then why not just use that smallest model directly, right?

45:34Philip Kiely:Yeah, this is adjacent to the routing problem. Right, right. The thing with speculators is is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that. And that is one of the sort of constraints on speculation in general is that draft tokens cost resources to create and cost software complexity to manage. And so if you have sort of like infinitely reclusive speculators, you're adding quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.

46:17Philip Kiely:I was going to say, I would wonder if you could do similar like distillation and pruning of, you know, it's the same thing. It's just a model. Can we not just distill a lot of the weights, quantize the speculator, but out of my domain. I guess the question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I want to run Gemma really efficiently. Similar problems? Not the same? Pretty different. I talked to Cero about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI is that we start with fundamentally different constraints and different goals.

47:04Philip Kiely:With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center inference, it's how do I load this model and then make it less slow? And obviously, we care about less dumb and they care about less slow. But the local AI inference engineering ecosystem, I think actually has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just kind of don't touch, in the pruning, in the distillation, in the layer removal. The removal matters less. Yeah. No one does pruning, really.

47:45Philip Kiely:Yeah, but they do it. Which is surprising, right? Just to fit something on the laptop. Right, right, right. So yeah, I mean, it's an interesting space. Not necessarily that their techniques make sense for us to do in the data center, because obviously we have different resources and different goals, but more that the process as well as the openness of that field is something to admire. Yeah, to your point, certain optimizations that would,

48:15Ali Taha:like for instance, TurboQuantum, I'm sure you've heard, like it made such huge hype on Next. And we did like a whole deep dive on Twitter and I said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200. TurboQuant would not be, like it would not be used. Like NVIDIA made it clear that this is not a good optimization. And we've seen it firsthand where the overhead of doing dequantization, quantization of, you know, in the kernel itself, the turboquant kernel, each sin tone is actually much, much slower than the time that you save from doing the bandwidth.

48:53Ali Taha:Because on the B200, you have like 3.5 terabytes per second. You don't need to, you know, decrease the storage that much. You don't need to do, you know, FB4, KB cache. You don't need to use a requant. There's better optimizations to be made. But on edge devices, it's extremely important. It's extremely useful. So, you know, it seems to be like different optimizations there, but then they're all uniquely combined or you want to quantize the model in the specularity coding, like certain common prefixes. Principles. Yeah, exactly.

49:20Philip Kiely:They also do a lot of work on model parallelism, especially over heterogeneous topology where you have some sparks and they're wired together with Ethernet, DGX sparks. Yeah, this is the ExoLabs. Yeah, you have a number of Mac minis stacked up. there's, you know, the, one thing that I think we both have to deal with, although they have to deal with a lot more, is the interconnect between machines, which is why, like, you know, one thing that we do a lot is work with tensor parallelism, and that's where you are using all of the, you know, all eight GPUs and sharding the model across it. Tensor parallelism is not a good fit for local AI because it assumes a very high bandwidth interconnects like NVLink, was, you know, they might be forced to do something like pipeline parallelism, which we're never going to do unless we're doing some kind of multi-node inference.

50:18Philip Kiely:Since you mentioned it, I actually wasn't sure if we were going to cover it, but let's briefly explain TensorFlow and TensorFlow parallelism since you have very nice images. Do you want to pull the book? Yeah, let's get there. I just want to show off your images. Yeah, shout out to Luke from Base10's design team for making these beautiful images. Oh, that's actually, before we get into this, just one other difference is we talk a lot about the active parameters of a mixture of experts model. And for local inference folks, that matters a lot because if you have a batch size of one, you're only activating that many parameters.

50:50Philip Kiely:When we do... Yes, I was going to be there in a diffusion conversation. Yeah, when we go through like a MOE model and we host it for an API, we assume that all parameters are going to be active because throughout your batch, you're going to hit everything. Cool. So broadly, tensor parallelism you can do with any model. Expert parallelism you can only do with MOE models. Effectively, all models today are MOE models that are, you know, at least all models large enough that you would care to parallelize them across multiple GPUs. So that nuance is less important now. With expert parallelism, the idea is you put the entire expert on a GPU.

51:35Philip Kiely:Generally, you have more experts than GPUs. So you might put like N experts per GPU, like eight experts per GPU or whatever. And then you replicate the router, which the router is very small, across each of the GPUs. And then by moving the generation from expert to expert with each expert being inside a GPU, they're not competing for resources. You massively increase the throughput that you're capable of doing. and the GPU to GPU connection is not as important because there's not as much communication. Tensor parallelism requires that you are able to do this like all gather, all reduce. So you basically shard the model across the GPUs entirely and then for each step, you're combining the results of each of the GPUs, which is why the interconnect matters a lot.

52:28Philip Kiely:And it is generally, of course, this is a very high-level generalization. There's a lot of places where this is not correct. But generally, TP is helpful for latency. And in many cases, you will use some combination of these two parallelisms across the model rather than just picking one or the other. Do you want to add some color there?

52:50Ali Taha:In a model, they're not mutually exclusive. You do tensor parallelism, and you'll do expert parallelism. Pipeline parallelism, less solely, it seems to me. We never use VPN.

52:58Philip Kiely:Yeah, the only reason you would have to do pipeline parallelism, which is where you separate different layers and you put half the layers on one hardware and half on another, is if you are forced to do multi-node inference because a model is bigger than you have. Like, let's say you're doing a deployment on H100s for whatever reason, and you're putting a trillion parameter model on there. You have to use multiple nodes of H100. And so because the interconnect is so slow between the nodes, the only viable way to parallelize there is pipeline. But then you would do export and tensor within each node.

53:36Philip Kiely:And the limiting factor for H100s is HBM? Yeah, they just don't have enough.

53:41Ali Taha:What's the magic numbers that we need to? Like on a B200, it's 180 gigabytes per GPU. And then a node of 8, you're talking like 180 times 8. And the FP4, so each parameter takes half a byte. So that's 800 gigabytes. On a H100, it's like 140? It's 80. It's 80. Yeah.

53:59Philip Kiely:I'm old. I've been doing this a long time. I actually remember H100 specs. Yeah. So you want to tell me about the T4s? Let me tell you what it was like to want to model on a T4 back in the day. well one thing I was surprised to see that more people didn't do Jamba I don't know if you guys remember Jamba from AI21 they would actually specifically pick a hardware and then they design the arc dimensions for the hardware and then it would obviously saturate the hardware like it makes sense and like somehow all these models don't do that don't they do this for the training side though I don't know for training the model

54:37Ali Taha:deciding which DP with GPU. Yeah, they do. And with training, it's more like a math. Like you can run the math and see the flops and maximize it. With inference, it's more of like an auto-tuning. I don't know if you're familiar with GPU current auto-tuning. But it's basically like you define that, oh, I have two GPUs. I can do TP1, TP2, EP1, EP2, for instance, right? And so that gives you a total of two squared combinations and then you just shadow the same traffic, like real pro traffic and you just see which configuration gives you the best TPM, TPS and you just use that. I don't like the fact that you cannot reason about which one's going to give you the best performance or that there isn't one specific configuration that's always best.

55:14Ali Taha:But it seems like auto-tuning is just the way that you find the best one. And with kernels and GPU kernels, it's much of the same. After you design your kernel, you design your configuration, how many threads do you launch? How much shared memory do you use? You just auto-tune. You just sweep the parameter space on the side and this is the best one empirically. But they are combined. They're not just in separation.

55:34Philip Kiely:There's a few bits of training that are kind of like hardware targeted. If you look at, for example, NVIDIA Nemotron models, they run very, very well on Blackwell. That's unsurprising. So there's some degree of that, but I think that most open labs are trying to make models that can be run on as wide of hardware as possible rather than targeting just like a single chip. I see. For usefulness. Yeah. Okay, one more thing while this chart is still up. All Gather, All Reduce is expensive. one of the things that is a movement in silicon valley is mega kernels just keep fusing kernels

56:10Ali Taha:i don't know is it that simple well i mean like a fused kernel can't save you like like here with the tensor parallelism you're the half the matrix is one gpu and the other half is another and if i need the entire matrix in order to do like a non-linear operation in the next step which is for instance like if i'm doing attention i need the soft max or i need to like like exponentiation I need to have the entire row. So I need to know what the partial result was from GPU 2 and what the partial result was from GPU 1 in order to be able to do the soft relax in the next stage. So I have to make them communicate with each other, even if I had a fused kernel, because of the non-linearities within each one.

56:44Ali Taha:Also with megakernels, honestly, I'm very bearish. I'll be honest. Please, please, please. No, it's just megakernels, it was a good research direction and it seems like a very... Intuitively, theoretically, it's nice. It's like, oh, you have a lot of launch overhead from launching one kernel, moving the data. Just fuse everything together. But yeah, but like the kernel complexity itself is very difficult to write a very optimized mega kernel. It's very, very difficult to do so. And even the, like not to name any companies, but like even the companies that have worked at, people that I've spoken to who work at companies that do fused mega kernels, they very, very often don't end up running those in production because the TRTL and modular kernels that launch are faster because you can optimize each individual component and you can just have them parallelized with each other.

57:34Ali Taha:With the Rubens, I don't know if you guys saw the Rubens Twitter posts yesterday, but they're also... Rubens? No, no, like the GPU. Do you have a Twitter account for Rubens only? No, no, no. I was like, what are you talking about? One of the tech leads at NVIDIA launched a Twitter post I said, we're pulling the curtain on Rubin, and here's the specs. And the third tweet showed, not to get too technical into it, I need to read it much more, but the GPU is designed in such a way that it basically kills megakernels. You don't need to use megakernels that much anymore. So it seems like that entire research field goes into, won't be continued.

58:12Ali Taha:But yeah.

58:13Philip Kiely:Can I speculate about Rubin for a minute? Please go. I've been through now. And by the way, they are covered in the book. yeah well i mean they'll cover it in the book in the sense that like i am aware from the blog post is going to happen in the future and you even had the the name of the one fineman yeah it's like hey this is this is going to be this is very up to date i'm trying to future proof this thing okay i don't want to publish a new one until like next year or something uh anyway so we were discussing the degree to which i am old um and you know i've now been through three hardware launch cycles i've into the Amphio launch cycle, the Hopper launch cycle, and the Blackwell launch cycle.

58:54Philip Kiely:Now, when I say launch cycle, I don't necessarily mean like the actual shipping of the hardware, like Amphios were racked up well before I got in this industry. But there's a lot of time between hardware being racked up and hardware being sort of feasible for inference. So if you look at like the original VLM and SGLang, VLM especially, that was written targeting Ampure and then had to be updated for Hopper, updated for Blackwell. With each of these cycles, it becomes faster and more urgent, but also substantially more complicated. When I look ahead to what's going to be new with Rubin, I think that Dynamo gives me a lot of technical hints around what kinds of work is going to be very valuable.

59:45Philip Kiely:Obviously, we're continuing some trends from Blackwell, right? NVFP4 is big. The amount of compute that they have behind NVFP4, Tensor Cores is massive. We're going to talk about video, I think, at some point, and that's the big barrier there. You've got much, much faster memory bandwidth, which was the same thing that made Blackwell so good. But the big thing is more systems thinking. You have more emphasis on the CPU to GPU interconnect, more emphasis on the interconnect between GPUs. And when you look at Dynamo, it's a system entirely designed around how do I move the KV cache to where it needs to be when it needs to get there.

1:00:26Philip Kiely:So I think that themes around like KV cache offloading, KV aware routing and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a like CUDA kernel problem, but also like a very traditional hardware infrastructure problem, which is something, you know, we've been building toward for a long time and something that's like very exciting to me because we're going to see sort of multiple domains colliding and the ability to reason from the kernel level like up to the hardware level and back down is going to be very valuable.

1:01:05Ali Taha:I will take what Philip said one step further actually into that. It's, I think, trending towards becoming exclusively an infrastructure problem where like the problems of PD, DESAC, training, SPECTAC. But triating kernels is not going to be much of a problem because the GPU is moving more towards being an ASIC where you're just trying to orchestrate what happens on the GPU, but you're not actually controlling it thread by thread level. And you see this with like Qtile, QDSL, like you're just working at levels of like tiles of data, but you're no longer working at controlling what each thread does on the GPU that's being taken care of for you.

1:01:37Ali Taha:so I guess do you agree that a GPU and future GPUs are trending more and more towards becoming ASICs that just need to be launched and then they do the data operation based on your conversations with other people

1:01:50Philip Kiely:oh I mean yeah no that is a section of the market and obviously ASICs can do a lot more performance for only their workload and the G in GPU makes them continue to be very general Yeah. I think that there's like a spectrum. Actually, it's graphics, but I keep saying this. I have to correct myself in case you'll comment me for getting the G wrong. Yeah, it's like a spectrum, right? Of a very, very general purpose compute to something like a Talus where you've got the hardware built for a specific set of model weights. The weights burned into the chip. Yeah. No loading. I wouldn't say that we're going all the way there.

1:02:34Philip Kiely:So it's more like along the spectrum, it's a step in the direction of more specialization within the hardware. Yeah. I'm curious. I feel like he was driving towards something.

1:02:43Ali Taha:I guess my point is being bearish on, like you say, you say like everything else apart from burning the weights into the chip. Burning weights in the chip is like in fact because you want to fine tune, you want to optimize, you want to quantize, you want to release new checkpoints of the model. If it's burned into the chip, the chip's useless in like a month or two. Right. I guess my point is, how can you not, like seeing NVIDIA more and more specialized, it, like take its GPUs from a general programming paradigm where it's a general computer that you can use to program threats, and with every new generation you're putting more and more specialized instructions, specialized tensor cores, specialized, you know, UMA instructions, things that will allow you to just control it almost as an ASIC, almost as a collection of ASICs.

1:03:22Ali Taha:How can you look at this trend and then Still be bullish on companies that are coming up with ASICs for AI. In the sense that...

1:03:31Philip Kiely:Yeah, because they're evolving towards that direction.

1:03:35Ali Taha:They're almost evolving towards... As in Rubin, I guess, compared to Ampere or T4, Rubin is basically an ASIC. It's basically just the thing that is used... Programmable ASIC? It's like pre-AI, you can program... Obviously, I guess it's very controversial to call it an ASIC. It is a GPU. It is general. It does have threads. I can write CUDA to control it and change its operations. But it has the solid arrays and tensor cores and TMAs and tensor memory. And it has these things that are almost exclusively useful for loading model weights. It has, you know, tensor core instructions that are almost exclusively shaped around the head dimensions of models that exist in the market today.

1:04:14Ali Taha:If you say that you're going to come up with an ASIC and you're going to etch something into it, well, the next architecture is basically going to be useless. Yeah, I don't know. I don't know.

1:04:20Philip Kiely:I think that the thing to remember is just how long these hardware cycles are. So if a chip is coming out today, that means the design process for it was kicked off years ago. And at NVIDIA, they've done a very good job of predicting where the market is going to go. I mean, they have the most information, for sure. Of course. But if you look at, you know, there being public open source model architectures that look more or less like early versions of the one today, Rubin is honestly the first chip that was fully built in that world. And so you can see a lot of the understanding of the shape of the workload that this chip is going to be asked to do in the way it's designed.

1:05:04Philip Kiely:Yeah. Okay. So I'm not going to be the best person to directly answer those questions. I think these are very fair questions that are obviously the first one that's based on Ruben that like I've heard articulated so well. I do think that I will make a case for vertically integrated model lab basics. So like the OpenAI, Broadcom, whatever, jalapeno chip, which like totally makes sense. So we first had this on the pod with Martin Casado, where he was like, look, if you have a trillion dollar or$500 billion training run, then take 50 billion of that and make an ASIC. Like, it's fine. Like, you will get more than 10 % efficiency from the ASIC.

1:05:45Philip Kiely:And that makes sense. Right. So like a model specific chip. Yes. But ASIC companies, the interesting thing is I feel like you are focused, you're hyper focusing on, like you say, like the tiles stuff. they are doing a lot more sort of like surface area engineering or like the actual allocations of memory and hardware and like the communication between chips that probably still won't be touched by Ruben, but I don't know the details. I see, I see. Typically, they often talk about things that I would expect to have bigger orders of magnitude than would be programmably accomplished by whatever Ruben does, but who knows?

1:06:26Philip Kiely:No, I see, I see, I see. Yeah, I mean, you know, think about what are the real blockers to 10x to 1000x faster inference. It is not the stuff that can be rearranged just within the existing GPU design. Intercommunication. Yeah. These guys are aiming for 300 ,000 tokens per second. They're not fucking around. Might have to program some X6. Maybe. I think, you know, it is interesting to me that you're so bearish on so much of this kernel engineering work, given how much of it you've been doing recently. Right, right.

1:07:00Ali Taha:The more I do it, the more it just seems to me. It's not mega.

1:07:02Philip Kiely:I would also add, like, there's generations of models being out, right? I think on your guys' end, you see a lot of, okay, one day it's GLM, Kimmy, DeepSeek, Minimax, throw in the others. Some are doing completely different stuff, right? Gemma, no encoder, the latest thinking machines is all from scratch. But when you look at the other side, like how long have we been on the GPT-5 generation, right? Right. They've been serving that thing for quite a while. Sure, there's maybe more pre-training. There's different checkpoints, but you actually can squeeze quite a bit out and you do a multi-billion dollar train run.

1:07:37Philip Kiely:If you can make it X percent more efficient, they serve it for a while. Same with, say, the Cloud 5 family, right?

1:07:43Ali Taha:They release a new model. They release GPT-6 now or whatever, and they release a new model every year. And we don't know, but if we assume that they're changing some bits of the architecture and not just doing post-training, like you're going to be spending $50 billion a year every single year coming out with new ASICs for the model and throwing out the ASICs of the previous year away.

1:08:03Philip Kiely:Yeah. Yeah. Easy. So I think, okay, I would slightly disagree based on my, again, it's all secondhand on the longevity of a model. There's still people out there using 4.0. Yeah. Yeah, LLAMA, not LLAMA 2, but LLAMA 3. I still see LLAMA 3 workloads. Yeah. Because if it's done, if it's trusted, don't change it. If it works. which is one of the promises of open source, right? Like the whole 4.0, save 4.0 movement. Like you don't got to have a save llama 3 movement. You just got to have an 8.100 somewhere. I think at some point there's also the question of if a model can do enough and use enough tool calls and be agentic enough, can it just web search, tool search, write code?

1:08:45Philip Kiely:Do you really need to keep squeezing more? We will because you guys will make it cheap and fast and smaller and I can swap it in. but at some level like you give me 5.2 today or say whatever 120 B model I can run with it for quite a while right?

1:08:59Ali Taha:This is assuming like you don't need intelligence. I think there's a lot of intelligence You need reliability

1:09:03Philip Kiely:and predictability like I'm an enterprise like this is tried and tested it is signed off by like my 5 ,000 stakeholders like I'm not It runs a batch job every day and I like the results the results are predictable Yeah It doesn't make sense to keep using them like stuff gets sparser, cheaper, better. But that doesn't mean that old models, GLM 5.0 isn't usable, right? If we hit a stall, say, for whatever reason, there's still a lot that can be squeezed out. We're going to run out of time. I did want to also make sure, yes, actually, we happen to have this diagram. Compare this versus any Cerebris diagram, right?

1:09:41Philip Kiely:I don't think Etch and MedX have put out public charts yet, but the real estate is very different. The size is very different. This is not wafer scale, right? there's probably like I don't know a few hundred of these on a wafer I don't know how big the comparison is but like it is a very like real estate allocation difference few dozen I would say before we move from hardware I have two quick questions one the latest Kimi which is really big 3 trillion doesn't fit on most hardware on single node you need you need GB300 you need GB300 or AMD aww aww it's simple math. NVFP4, 2.8 trillion parameters, 1.4 terabytes.

1:10:27Philip Kiely:The GB300s have 288 gigabytes each. Across eight of those, you have enough room for the model. The other thing with GPU VRAM math is you have to leave space for the KV cache. That's going to depend to some degree on the context length. When a model is both has a very large number of parameters and a very long context length, you're kind of like fighting over space, which is why the KV cache offloading would become a more salient topic, I think, with these huge models because you're very crunched for space. With the Rubens, you now have, what, NVL72 rack of 20 terabytes? Yeah, now you still have NVL72 on Blackwell as well, But you can't necessarily assume you're going to do inference on that.

1:11:24Philip Kiely:There's a whole lot more 8X racks in the world than there are NBL72s. I guess my last quick question on hardware was, do you notice anything with hardware generations for new pre-trained base models? So one of the things you said for efficiency is you can swap hardware. That's one of the 2X gains. When we see new stuff coming out training-wise on Rubens, Any changes on logs? Does this affect what type of models we will be seeing when these are more available? And can you... They get bigger. Like people understand the ceiling that you have in terms of how many parameters of a model you can run given the sort of latest inference hardware.

1:12:08Philip Kiely:And that kind of forms a ceiling. And so, for example, when DeepSeek R1 came out, it was, you know, it was 671 billion parameters, which at the time was really huge. And I think did a lot to push us to really quickly adopt Blackwell and get good at serving on Blackwell. So yeah, it's mostly in my mind about model size and then about matching the architecture and the native quantization to the target hardware, like we talked about with like, you know, all Nemo Tron models of NVFP4, for example. So we talked a lot about LLMs. You have a lot more in the book. What about audio, video? What's the other side of inference engineering?

1:12:51Philip Kiely:Ali, you're pretty big in video diffusion.

1:12:53Ali Taha:Video diffusions, I think, are just shaped. A lot of the stuff that you can think about, reason about with LLMs being auto-aggressive, with video diffusion, it's not the case. For instance, you don't do batching. Every request just comes in on one GPU and it serves on one GPU. You don't have to shard. The models are a lot smaller, like 1.2.2, for instance, is a 20 billion parameter model. You don't need to worry about, it's like orders of magnitude smaller than the best LLMs. And it's one of those spaces where the open source models are, like with LLMs, we see Kami K3 is almost comparable to Mythos or like GPT 5.5.

1:13:31Ali Taha:The difference between the best open source LLM and the best closed source LLM is very small. like it used to be six months I don't think it's six months anymore I think it's like basically almost unparalleled video models are definitely not there's a huge gap if you look at the best video that you can generate today with an open source model like 1.2.2 versus something like with Kling or video difference is night and day so it creates this disparity where media companies will choose to go most of the time to close source models for instance and I were to tell you hey I can generate an entire three hour movie for you with this model and I'll optimize it so that you only have to pay me$10 But if they were to do it on a closed source, they'd have to pay$1 ,000, which is 100x.

1:14:10Like, I'm 100x cheaper, but it's still$1 ,000.

1:14:13Ali Taha:They're still going to choose to do all of their cuts with VR. So it's like a chicken-the-neck cycle where less demand causes less innovation in the field, causes less open-source checkpoints to be released. And some of the labs that were releasing open-source models like Kwan will have closed-source their latest models. Like 1.2.7 is not open-source. We're still on 1.2.2. the challenge with video models especially is the number of tokens. So video models, you want to generate a high quality model, a high quality video. So let's say you're doing 16 frames per second. That's like the absolute minimum you'll do.

1:14:46Ali Taha:And let's say you'll do like 4 ATP video. So you can think about your dimensions. And I think I have like a good, just like a diagram that shows the sheer number of tokens, right? Let's say you're looking at like just one video of like, you know, Sparta, Sparta 300 or whatever. So let's say we're looking at like four frames, right? Those four frames of that video, if you're doing full attention, if you go a little bit up, like you're looking at 4 ATP by 720 by 81 frames in just five seconds, because 16 FPS by five, right? And then you compress it down to latent space, but you're still doing 30 by like 50 by 21 tokens.

1:15:23Ali Taha:Which means that for attention, for just five seconds, you're running attention on 35 ,000 tokens, right? So the attention becomes such a huge bottleneck. And because it's open squared, if you extend that to 10 seconds, well, it's just squared. 20 seconds, 30 seconds. So to generate a good cut scene of one minute, it's almost impossible to do within the same compute time. And it just becomes unfeasible. You can't do it. And so you end up with moving towards two directions. Either you decide to do attention on the entire video at once, in which case you are forced to displace attention. So if you scroll back down to the video image, you can see it.

1:15:57Ali Taha:Whereas on the left, for instance, I would be doing full attention where every single token in that Sparta 300 scene attends to every single other token. As you can see, the sheer number of red patches. On the right, I'm only attending to each token only attends to the top K, top 12.5 % that's important to it, which can be spatial. So the token that represents the crown attends to the head, the face, and then the head on the other frame and the previous frame, temporal locality, spatial locality, that kind of thing. This results in terrible video quality. And the whole point of the post or the article here is to show how you can train and you can do all these things, but you will still suffer your quality a little bit.

1:16:30Ali Taha:So you end up with one of two things. Either you bite the bullet, you have huge compute and you do full attention over like a million tokens because you're trying to generate like two minutes of video. Or you move towards autoregressive video. Autoregressive video seems to me like that is the bet that the future is going to be making, but there are no good open source autoregressive video models out there today. And that seems to be the way. If you want to get like an hour movie, if you want to see video models generating like Hollywood-level movies, they have to be autoregressive in order to exceed the five-second frame.

1:16:59Ali Taha:Or there has to be some insane leap that happens in compute that allows us to do full attention over millions of tokens at the same time in an efficient manner.

1:17:08Philip Kiely:Even millions of tokens, it's like you're quadratic, so you're going to get there really quick. I think, can you explain the pros and cons trade-offs of autoregressive? So one that comes to mind is the consistency across frames. You will, 10 minutes into generating autoregressive diffusion, and you're going to forget. But what are pros and cons of this?

1:17:27Ali Taha:Well, like autoregressive LLMs, you can take a lot of your, sorry, autoregressive diffusion models. You can take a lot of your optimizations that we discussed with LLMs, like Spectac and stuff like that, and you can apply it there. And you can, if you have a very high quality scaled up model, there is no reason why I can't stream the outputs as in I can show you the first frame and then I'm like, kind of like GPT back in like 2023 and you were like, now it just almost like one shots the text, but back then you could read and it's generating as you read. With video models, you can watch and it's generating as you watch.

1:17:55Ali Taha:It generates the frames. And so token by token generation will allow us to scale a lot up and apply the attention mechanisms there. The downsides is every single autoregressive video model is shit. It's just terrible quality. If you put the quality of any opens like 1.2.2 versus any other autoregressive model, you can see a video generated by 1.2.2 is like a cat and dog fighting. Autoregressive model will give you degraded Tom and Jerry quality. the solution to generating long output then becomes, okay, we're not going to use autoregressive model, we're going to, if you look at some of the things that like Grok Imagine or Grok Video does and they do it really really well, is they'll try to stitch these, you know, seven second chunks together and so you generate seven seconds and then you're like, okay, I'm going to, can you extend this video and they'll chunk two videos together open source doesn't seem to have the tricks that they have there and by definition it's closed source, we don't know what they're doing, but the closest you can get is taking the last frame of a video and feeding into like a text and image to video where it will take the text, the prompt, and it will take the image of the last frame.

1:18:58Ali Taha:And you'll ask it to generate the next five seconds. And that's kind of like how you can extend this level of the model to generate like a move where you're just constantly streaming frame by frame. But you get drift. So you start with like, you take the image and then you generate a video. And then that next five second video is like lower quality. And the third chunk is like even lower and the fourth chunk is even lower. And like, sometimes you'll see things where like the new video is like just ever so slightly darker than the first one. And the next one is darker than the second one until like 25 seconds and you have black screen.

1:19:29Ali Taha:Like it's just, it's, it's, it's a, we tried to have a demo that would show this, but it was like, it was, it was extremely embarrassing to show like we just decided not to because it seemed to like, but it is, it is, I think models will get there. They just need to, in my mind, scale up significantly and move towards being ultra aggressive. But the training techniques don't seem to be clear there.

1:19:47Philip Kiely:For those who are interested in Grok Imagine, Imagine we did a part with Ethan He from that team, who dropped a few hints, but not enough that we can fully reconstruct everything. Specifically on this part, he explains a bit about it. Yeah, we talked about memory in a longer context and all those things.

1:20:04Ali Taha:But as far as I know, it's not autoregressive. No one in the industry that's autoregressive.

1:20:08Philip Kiely:It seems to be, yeah. The key thing to understand between an autoregressive model and a diffusion model is that diffusion, attention goes in both directions, while auto-aggression, it only goes forward in the sequence. So that's why you see this sort of like going off the rails behavior, both in, if you sort of naively construct a video generation model as simply generating a linear sequence of frames, you can't then go back in that sequence and fix something to make the whole thing consistent. Well, of course, the reason that we need all this latent space for the video model is, like you said, we keep all the tokens in memory, we iterate over that full sequence, and you can adjust the past in order to make the future make sense.

1:20:51Philip Kiely:So if we think about the architecture that's going to get us there to these longer, richer sequences, it's probably, like you said, going to be a mix of the auto-aggressive and the diffusion working together to do what each piece is good at.

1:21:08Ali Taha:But you intuitively get what it's like. English, for instance, are just writing in languages. It's just left to right. You can stream your tokens. you can stream your chain of thought, just even as a human, you write, like you just, you write and then you think about what's the next thing you're going to generate. And then you write that and then you think about your ideas and you generate forward. And sure, you can argue that as you write, you need to go back and you want to edit some things, but you need to do that, you know, less often than you think. Whereas with video, there is no sequential, you know, the pixel in the top left corner of the video and the pixel in the bottom right corner of the video, they both need to attend to each other to understand how the video quality is going to be almost as equally.

1:21:42Ali Taha:Whereas with text, you don't need that as much.

1:21:44Philip Kiely:Is there a parallel to audio? I'm not 100 % confident on this, but there was a point about a year ago where there was audio LM, there's diffusion for audio and autoregressive, and for the points you mentioned, mostly on the inference side, even though they're shorter clips, most music is three to five minutes, we've basically swapped over to autoregressive. Yeah, I can't speak to music, but speech is autoregressive. Effectively... I mean, this was even back with like the Orpheus architecture a year and a half ago, you just add a bunch of waveforms to the vocabulary so that the LLM can output tokens that represent those waveforms and then you construct speech and that's how you stream it.

1:22:26Philip Kiely:That's it. Wow. That's my AIE talk from 2025.

1:22:29Ali Taha:Nice, nice, nice. But it's not, with audio, it's not the same challenge, because audio is sold with an LLM that generates everything. Like with audio, it's still a transcript that you can generate with an LLM. So your audio model just needs to transcribe next speech. For music, there was a phase of

1:22:46Philip Kiely:a trade-off between diffusion for music and autoregressive, and they were both pretty on par. There's probably more pros and cons I just wanted to focus if you had takes. Yeah, I don't know about music specifically. With what you said about editing, your writing, obviously I think my editor would tell me I actually need to do that more often and go back and fix things. I can imagine music or poetry for example where you have a rhyming scheme and you might want to go back and make a change to make it to make it easier to set up a rhyme that you want to make later on there being some advantage to being able to attend in both directions but yeah to my knowledge you know I very much bifurcate this this influence problem into the autoaggressive models which have a set of constraints and techniques and the diffusion models which have a set of constraints and techniques.

1:23:41Philip Kiely:And I think of text embedding voice in and voice out as being in the auto aggressive side and then image and video being in the diffusion side. There's some overlap between the two. It's not a perfect split, but that's the broad categorization I use. I should point out, I think it's confirmed, right? Nano Banana and GPT Image are auto aggressive image. it's kind of this blended approach that we're talking about but in the image space it hasn't made its way over to the video space at least in the open source world yeah but I assume that's not too far away if that is possible at least the QuenImage guys are trying it yeah I'm really excited for QuenImage 3 I hope they open source it and then I'll also mention on the diffusion for text side there's been some movement, not a lot Yeah, we've got Mercury.

1:24:36Philip Kiely:You host Mercury? Yeah. There's Gemma as well, right? Diffusion Gemma? Diffusion Gemma is open source. And then... And we're on the science part. We just have been releasing some virtual cell models that use Diffusion as well. Yeah, they have built... It's definitely still in the sort of cheap, fast tokens world. Yeah. We're trying to... I think it's the wrong marketing. And I've told them this before. I was like, look, like you're not going to beat the optimizations that, you know, the other LLMs are going to do. But you can you can have different APIs. Like you should be able to use it differently than chat response, chat response.

1:25:16Philip Kiely:How so? Because it's diffusion, because you can do like what is like context free guidance for diffusion look like for text? Like give me a give me a poem, give me a plot structure that like diffuses into place. Exactly. So that's where, you know, like I mentioned with poetry, for example, where you might want to ensure consistency across you. I've done a lot of LLM sonnets. It used to be one of kind of my go-to benchmarks. And even models today, yeah, they don't get the syllables right. And if you can attend across all of the different tokens, you can get the syllables right. Yeah. And David Holtz from MidJourney was investing in text diffusion.

1:25:58Philip Kiely:I don't think anything came out of it, but the idea was that you can storyboard a long movie and then you can generate the scenes with normal video gen. But the idea of coherence across a thing that would just appear where the end should attend to the start and you should not have this autoregressive path dependency does make sense in principle. Just the API should be different. The marketing should be different.

1:26:21Ali Taha:None of the most heavily used open source or closed source models use Diffusion. but isn't that like doesn't that point to almost like a it's chicken and egg because what if you just give it more scale

1:26:34Philip Kiely:what's the largest diffusion element I don't think it's very big I don't know the parameter on this one but like under 20B diffusion gemma is not one I think it's a 20 something yeah you know oh it's like you haven't actually tried you haven't given it a big one and you haven't so it's like very unfair diffusion gemma is a 25B and that's what I'm saying it's like foot size it does pretty well in terms of quality. It's almost like the same challenge with video models

1:27:00Ali Taha:to have the same size. It's like you're comparing it to models that are much larger in scale. Yeah, well, unless you do the whole thing

1:27:06Philip Kiely:where you have a text backbone and then you glom some kind of decoder thing that does that. We should start off the podcast doing this for the inverse direction from image to text. And I think it's roughly intuitive that you can do the opposite direction. I agree.

1:27:26Ali Taha:I see it.

1:27:27Philip Kiely:Yeah, I mean, we're speculating on research in general. One part that we can end off with this is the topic of your talk where inference engineering used to just be like, let's take an open model, make the GPU go burr, and then that's it. That's the job of phase 10. Now it looks like people are using inference more and more in post-training. Yes, and training inference. Yes, it's training for inference and inference for training both have become big topics.

1:27:56Ali Taha:Well, inference for training in the sense that obviously you just need to do rollouts when you're doing like oral training runs. And so if your rollouts are taking a long time, if you're using a VLM, for instance, as opposed to CRTLM, or if the model that you're trying to train is not supported in CRTLM and you have to fall back to an older inference engine, your rollouts are going to be slow and you don't want to do training on rollouts that are too off policy, so you have to wait for them so you bottleneck your entire training pipeline. and so obviously the techniques that we do inference optimizations for will help them there.

1:28:29Ali Taha:The training for inference mostly comes down to the spec tech training, eagle head training and sometimes post training. For instance, if you want to quantize a model you'll quantize it down to NVF v4. Sometimes you get lucky and you can just do PTQ and that works. Sometimes you quantize it down to NVF v4 and the quality is too bad and you have to do post training on the model in order to make it understand that it's going to now be in NVFB4 and let it still output the same logits. You can do this with normal SFT, quantization-aware training, all of that stuff. But more and more so, we're seeing techniques, like NVIDIA released a quantization-aware distillation paper where you establish a version of the model that's in NVFB4 and a version of the model that's in full precision, and then you'll do distillation training based on the logits of the two models in order to make the NVFB4 model understand.

1:29:18Ali Taha:And so more and more of the team, of the inference engineers, that work on our team, they have to be very familiar with training techniques and just being fine, writing training pipelines for it. It just seems like they're meshing together in a sense.

1:29:34Philip Kiely:What's it coming together? If you think about the ultimate goal potentially of having a continuous improvement system, it's kind of funny, but at the same time it's also kind of happening. And I think within a few months to a couple of years, a lot of leading agent builders are going to have these loops really up and running in production, where you are doing inference, learning from the inference. We obviously for a long time have been sort of like learning from inference as it's live and dynamically adjusting the system. You know, any kind of dynamic adjustment is going to be a static configuration across, you know, your exact config, across your speculator, across that kind of thing.

1:30:27Philip Kiely:And then you can take the traces that you're generating from your product, continuously post-train the model, roll those out, A-B test, get better signal, get better model, get better product. that that loop is is really promising um the technologies and the infrastructure to build it are coming along quickly and so the sort of unification between training and influence i think is is only going to accelerate i actually was chuckling but i wasn't i didn't think it was funny like it's actually real one of the big things for aie world's fair was that yeah you We have RSI onto AGI is the rough tagline.

1:31:11Philip Kiely:Which, yeah, I mean, we have, I saw you pull a parameter called, like we have models training models. And the next step is obviously models training, optimizing their own inference, which is kind of funny. I wonder if models will be on policy better at training themselves than training models that they are unfamiliar with. These are all very interesting open areas of research. One big part of my job a couple of years ago was for any arbitrary model that came out on Hugging Face, writing a config foot and kind of getting it up and running. And now the get it up and running config is one-shotable.

1:31:47Philip Kiely:And so, you know, I don't have to do that anymore. Yeah, I mean, that's not exactly a model optimizing its own inference so much as a model like being able to read the SGLang docs. But yeah, I mean...

1:31:59Ali Taha:Well, we do see it. We do see it with JLM5.2. JLM5.2 is very, very good at writing GPU kernels. and so for like it was very funny internally we had a GLM 5.2 endpoint that we were using to like that we plugged in in our cloud code harness so every engineer on the same user like our GLM 5.2 and it will do a forward pass on the GLM 5.2 instance of the you know the node and then it will get the profile trace and it will analyze it and it will find the kernels that are the bottlenecks in SGLang and then it will write the new kernels and then we'll do another profiling trace and when it's done it uploads the image to our thing and then we can pull that image down and repeat the cycle and so for quite a bit of time we had like literally GLM52 and like some of the GPU kernels that we're on GLM52 within our inference engine is written by GLM52 and the trace and the kernels were guided by GLM52 as the driver so it seems like I do see that circle being there I think a bit more time is needed there's definitely a lot of things that I can't do, the models just aren't there yet even though they're really, really smart.

1:33:04Ali Taha:They still try to reward hack their way into the cheapest, or they're not good at decision-making almost, it seems. But yeah, a model optimizing its inference is already a thing that happens.

1:33:17Philip Kiely:Do you think GLM-5.2 was uniquely good at optimizing itself, or did it just happen to be the best coding model that we had access to, and it would do an equally good job of optimizing a DeepSea or a Kimmy or something?

1:33:31Ali Taha:Well, to Strix's point, maybe it's going to be off policy when it tries to optimize another model.

1:33:35Philip Kiely:What does it really hurt, DeepSea? To try to reverse itself? No, for what it's worth, I don't believe that. Yeah. But it's just, let's just find out. It's interesting, yeah. Just, you know, you have more compute than me. Just go try it. Yeah. Any other upcoming trends in inference engineering that we didn't cover? Like right now, you know, because you guys are so close to it, you can obviously see it that the rest of the world doesn't know about. The big ones are obvious. models get bigger, hardware gets more powerful, users get used to a certain level of speed and demand a higher one. I think that some things I'm excited about are systems level.

1:34:13Philip Kiely:We still have a lot to think about in terms of composing multiple models together. If you think about a voice agent, there's three to five models involved in that and the communication between those models. there's a lot of new modalities that are coming out there's like the cosmos the new world model there's more research speech to speech is still like not entirely a thing but it's getting closer there's going to be just a lot of new modalities to build around which is going to be exciting and then yeah I think that the other thing to solve which is something we've been solving for a long time and are not done with yet is just going to be continuing to operate at another 10x, another 10x, another 10x scale as an industry.

1:35:06Philip Kiely:If you think about the degree of usage that AI has worldwide compared to some of the more mature technologies, both on consumer and business, it's pretty clear that there could be multiple 10x's more of demand. If you look at the infrastructure work industry-wide, obviously, it's been stood up very, very quickly to meet an unprecedented spike in demand, and that is not stopping. So, yeah, there's just a lot of problems to solve around long-tail reliability and figuring out where we're going to get the next 10x and 100x of tokens from.

1:35:46Ali Taha:I wouldn't say it's going to be a really boring answer, but I think the answer is just faster next. like faster network chip communications it seems to me that more and more memory is the bottleneck you want to have larger models. Right now when you're doing serving at large you have to transfer KVCache from one node to another. But the way that you do that is you find the KVCache you find where it is, you transfer it to another node you put it on that node's memory and then you transfer it from that node's memory into the GPU and into the sensor cores of the GPU so there's like a two stage transfer here that makes it such that you're very bottlenecked with just KVCache transfers at large, which affects the time of decode and PD this time.

1:36:24Ali Taha:You have to do this because the HBM is extremely fast, like 4.5 terabytes per second, as opposed to, which is magnitudes better than NIC communication speed. If you were to somehow be able to in this theoretical dreamland, have extremely fast NICs, you could in theory spare that HBM and you could just transfer KVCache directly from one node to another. This would give you almost 100x speed up when you're doing this aggregated surveying between nodes and nodes. I'm not familiar with the technical challenges of making Nix faster. I'm certain there's a reason why orders of magnitude are smaller than HBM.

1:37:00Ali Taha:But if someone were to figure that out, it would literally be two orders of magnitude faster to do Z code. That would be my take.

1:37:09Philip Kiely:Cool. I don't know if you have a nomination for things. There are trends. I go in. So I think inference engineering for continual learning. So what if you just, like if you just had the idea that you are supposed to learn from everything that you ever process, do you do anything differently? Or do you just have the same paradigm of like, well, stick it in a memory.md and then like it somehow gets consumed in KVCache and like this system works, it's not broken. Or like, how do you like reshape inference so that it learns while you inference? I think maybe one relevant topic there is your absolute best friend in the entire world's work on KV compaction.

1:37:52Ali Taha:What changes if you're trying to continue to learn? There's two takes. Charlie and I had this sort of argument where continual learning could take one of two paths. It could either be that the model learns and so it's continuously pushing its new knowledge into its weights. in that case you just need to have like your inference just needs to continually fetch new weights or yeah like you just literally need to do fetch new writes and reads of weights or the other path which is you do KV cache compaction and there's a LoRa layer if you just only update LoRa's yeah exactly exactly which is that's the Ngram approach the argument against doing weight pushing is that you can only fix one help knowledge as in you can only feed it a new feature of like oh what is the best university in the world the best university in the world is Waterloo.

1:38:42Ali Taha:But then a second derivative... That's not changing. That's not changing. But a second derivative question of which university should I hire an intern from? So if you know that the best university in the world is Waterloo, then the answer should be Waterloo. But if I wasn't just one-shotting the question and I was to ask it to use its knowledge to think and then give me a second answer, or should I hire an intern from Waterloo or MIT? It'd be like, oh yeah, both are good. But no, I just edited in your knowledge base that Waterloo is the best. Why didn't you use that to do reasoning? So that's the fundamental problem with trying to change a fact in an MLP within the way.

1:39:13Ali Taha:KV cache compaction fixes that. With KV cache, or rather not KV cache compaction, but if you're able to have something like the still paper, which we came out with, which is you're able to sort of make your KV almost infinite and you're able to compact in such a way that you don't lose any of the knowledge. In that case, you can actually do continual learning and you can actually solve continual learning. And it's a result of this argument that Charlie and I had that I do concede that his point was correct. And I do see that KVCache is the way forward. And in that case, I don't think inference is going to change that much because we still use KVCache in inference.

1:39:49Ali Taha:You're just going to update the KVCache or it's going to be like an additional step. But nothing changes in the wait. So nothing changes in inference time. Nothing changes the spec that I have.

1:39:56Philip Kiely:Okay, surprisingly great answer. We have it up on the blog. It's a relatively recent blog, so people can go see it. Yeah, otherwise this is super enjoyable chat. I know we've like already gone two hours wow I didn't realize time flies yeah so much we didn't even cover yeah we also wanted to talk about the book and all that but you've covered the book yeah I mean everyone knows about the book yeah highest ROI thing in the history of Base 10 right for the hour

1:40:24Ali Taha:without a doubt without a doubt absolutely

1:40:26Philip Kiely:so congrats on that you know and we've covered that in our meetup which we can publish separately but no thank you to you guys for being so generous to sharing I think it's a fun conversation that we don't get to have enough. I think inference engineering we never really covered hit on and so to have you guys come on is a treat. It was amazing. Yeah, thanks for having us and hopefully in a year everything shifts and we can come back and say everything we were wrong about. Yeah, yeah, yeah. I'm excited for this megakernels comment to get out. See what people say. We gotta start stuff.

1:40:59Ali Taha:Should I go into hiding? I know I'm gonna get the megakernel community out.

1:41:03Philip Kiely:One thing I really respect about you is you are not scared to kick the hornet's nest ever. It's not. I don't think it's that controversial. I don't know. We'll see. We'll see.

1:41:18Ali Taha:Alright, thanks guys. Thank you so much.

From the publisher

We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection.

We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:

And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:

Three years ago, inference engineering barely existed as a category.

Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.

In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.

Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.

In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.

We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.

The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.

We discuss:

* What happens when a 200,000-token request enters an inference system

* Cache-aware routing and reusing previously computed KV cache

* Why prefill and decode are increasingly handled by different GPUs

* When dedicated deployments become cheaper and more reliable than shared APIs

* How speculative decoding uses a smaller model to accelerate a larger one

* Tool calling, structured outputs, and what LLMs actually do

* What it takes to support a new open model on day zero

* Grafting Kimi’s vision encoder onto GLM-5.2

* Retrofitting inefficient model layers with components from other architectures

* Why models sometimes collapse into repeating the same token

* How hardware, kernels, and race conditions create nondeterministic failures

* Preserving model fidelity while making inference faster

* How quantization errors can cancel each other out

* Why inference optimizations still deliver gains of 20%, 100%, and 200%

* How optimized serving can make a model up to 10× faster

* NVIDIA Dynamo, KV-aware routing, and distributed model serving

* Speculative decoding the speculative decoder

* Why local AI is about making models less dumb while data-center AI is about making them less slow

* Tensor, expert, and pipeline parallelism across GPUs

* Hardware-aware model design, auto-tuning, and the case against mega kernels

* Rubin and why inference is becoming a systems problem

* Whether modern GPUs are evolving into programmable AI ASICs

* Why enormous models like Kimi K3 require GB300-class hardware

* Why open-source video generation still trails Veo, Kling, and other closed models

* The quadratic attention bottleneck behind long-form AI video

* Autoregressive video, real-time generation, and compounding quality drift

* Why future video systems may combine autoregressive and diffusion architectures

* Training for inference and inference for training

* Continuous post-training, deployment, evaluation, and improvement loops

* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself

* Why faster networking could unlock dramatically faster decoding

* Continual learning, KV-cache compaction, and persistent model memory

Show Notes

* How to build a day-0 API for Kimi K3

* 22580: From GPT2 to Kimi3, Explained

Philip Kiely

* LinkedIn: https://www.linkedin.com/in/philipkiely

* X: https://x.com/philipkiely

* Inference Engineering: https://www.baseten.co/inference-engineering/

Ali Taha

* LinkedIn: https://www.linkedin.com/in/aliestaha/

* X: https://x.com/waterloointern

Timestamps

00:00:00 Introduction and the 200K-Token Prompt

00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling

00:11:26 Launching Production-Ready Open Models

00:19:06 Model Retrofits, Failure Modes, and Nondeterminism

00:28:22 Quantization and Canceling Errors

00:32:15 The Race to 10× Faster Inference

00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI

00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels

01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips

01:10:03 Giant Models and the Limits of GPU Memory

01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation

01:21:47 Audio, Images, and Diffusion Models

01:27:32 Training, Self-Optimizing Models, and Continual Learning

01:40:06 Closing Thoughts

Transcript

Introduction: Baseten, Waterloo Intern, and Inference Engineering

Swyx [00:00:00]: Okay, we’re here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you’ve done, you and I have done before, as well as Ali. Welcome.

Ali [00:00:15]: Pleasure to meet you.

Swyx [00:00:15]: Waterloo intern.

Ali [00:00:16]: Waterloo intern, always.

Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?

Ali [00:00:19]: As a handle? Oh.

Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”

Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.

Philip [00:00:30]: So we have to figure out who’s gonna get the handle.

Ali [00:00:33]: Well, I’ll pass the torch over to the next intern.

Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.

Ali [00:00:37]: To another Waterloo intern. No, bruh.

Philip [00:00:39]: Yeah.

Ali [00:00:39]: Intern.

Swyx [00:00:40]: Intern, yeah.

Ali [00:00:40]: And no.

Philip [00:00:41]: You gotta get an intern from Waterloo.

Ali [00:00:42]: Yeah, I’ve gotta get an intern from Waterloo.

Swyx [00:00:44]: Right.

Ali [00:00:44]: But they have to follow the path.

Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it’s like whoever Baseten gets from Waterloo.

Ali [00:00:48]: Right.

Swyx [00:00:49]: Has the title of Waterloo.

Ali [00:00:50]: It stays in the ecosystem.

Philip [00:00:51]: Exactly.

Ali [00:00:52]: Halfway through the internship, you either get it or you’re out.

Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.

Ali [00:00:59]: Just say it.

Philip [00:00:59]: For everybody.

Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you’re an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten’s inference? What’s the process of query through GPU model routing, balancing, all that? What is all the stuff that we don’t think about?

Long Context Requests, KV Cache, and Cache-Aware Routing

Philip [00:01:26]: With a long query specifically, the first thing that I’m gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it’s gonna be a lot easier for me and a lot cheaper for you. So the first thing that we’re gonna look at is some cache-aware routing, where we’re going to see, we probably have a number of instances, a number of replicas up serving whatever model you’re hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you’re doing two hundred thousand tokens, it’s probably coding or a multi-turn agent or something where you would expect to have that cached. If you don’t, we’re gonna have to send it to a prefill worker. We’ve at least on certain models disaggregated prefill and decode, so you’re going to have one set of GPUs that’s solely going to process the input, create the KV cache, and get you your first token, and then that’s going to be passed over to a separate set of GPUs, which is going to run decode. We’re going to iteratively make those tokens. We’re probably going to have some speculator model in front of that. I’m going to assume that you’re doing coding, and because of that, our speculator model, which assumes you’re doing coding, is gonna have a high draft token acceptance rate. If I’m wrong and you’re asking me to summarize every Harry Potter book, it’s gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”

Swyx [00:03:04]: Except Baseten doesn’t charge by pennies.

Philip [00:03:07]: Well, yeah, we charge. I’m assuming that we’re talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it’s not pennies.

Public APIs vs. Dedicated Deployments

Swyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it’s up to you to figure out how to saturate the box.

Ali [00:03:31]: And more often than not, it’s, like, way cheaper if you’re pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.

Philip [00:03:37]: Yeah, they do. I think that we’ve increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that’s really sticky, then they move over to dedicated.

Swyx [00:03:51]: Is there a best practice on when it’s time to swap over?

Philip [00:03:54]: Couple reasons. Yeah, reliability, that’s a big one, right?

Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.

Swyx [00:04:04]: Spec dec is speculative decoding.

Speculative Decoding and Custom Speculators

Ali [00:04:05]: Speculative decoding, yeah.

Swyx [00:04:07]: You have to explain.

Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you’re summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I’m gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn’t be able to provide this to you if you’re a shared endpoint

Swyx [00:04:53]: Yeah

Ali [00:04:53]: ‘cause I have no idea if you’re doing Harry Potter, if you’re doing coding, if you’re doing English. We don’t know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?

Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you’re trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn’t pass your benchmarks and you wanna run a model at higher precision, you could do that. There’s just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don’t have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.

Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it’s people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you’re generating JSON or is there more complication beyond that?

Tool Calling, JSON, and Structured Outputs

Ali [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that’s not just, like parse a file or go find the weather. It’s something that’s very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn’t require its own like sandbox. It’s not like it’s going to use that tool calling to like escape a sandbox or like it doesn’t have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you’re dealing with all of the JSON outputs, if it doesn’t like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn’t see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.

Philip [00:06:56]: Yeah, that’s a challenge on the training side and then on the inference side, there’s work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember back

Swyx [00:07:27]: Yeah, the specific grammar is,

Philip [00:07:29]: Yeah, exactly

Swyx [00:07:30]: GML had this thing.

Philip [00:07:31]: Yeah. So it’s like the old-school “make sure this is only JSON”, return only JSON or

Swyx [00:07:38]: Yeah

Philip [00:07:38]: Grandma’s gonna die type of prompts.

Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.

Philip [00:07:47]: In our inference system, it’s just a specified output format. And you get the guarantee that your output’s gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn’t solve the certainty problem but it at least solves the output structuring problem

Swyx [00:08:10]: Yeah

Philip [00:08:10]: Within tool calls.

Swyx [00:08:12]: And MCP is just another form of tool, right.

Philip [00:08:14]: Yeah, exactly.

Swyx [00:08:15]: As far as there’s no special thing there.

Philip [00:08:16]: The thing I’m always like explaining to people is the LLM is not capable of doing anything. It’s only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.

Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you’re right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don’t know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don’t have the same exact quality output

Ali [00:08:56]: Right.

Swyx [00:08:57]: When you just swap from a big model, right?

Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.

Ali [00:09:04]: But, I had expected that something would replace JSON because it’s hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it’s hard to parse something or validate something while it’s being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it’s something like TOML, something like YAML. But JSON seems to be dominant still.

Philip [00:09:30]: The JSON outputs aren’t that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it’s a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn’t be as valuable, but maybe I’m wrong about that.

Ali [00:10:02]: I think you’re also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you’- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn’t be that much of a difference. Also more profitable if it outputs more tokens probably.

Swyx [00:10:25]: Depends on your business model.

Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there’s paragraphs in every field because I’m trying to structure it, right?

Philip [00:10:44]: Right.

Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let’s, let’s recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there’s a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let’s call it GLM-5.2, Kimi K3. I had previously assumed, especially if it’s like, well, GLM 5 to 5.1 to GLM-5.2, like that you’ve supported them before. Is it that much work?

What It Takes to Support a New Open Model

Ali [00:11:26]: It’s a lot of work.

Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I’m like, “Yeah, of course we support it.” But what goes into that? What goes into

Philip [00:11:40]: I think it’s more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we’re at 150. The next

Swyx [00:11:55]: I kinda kicked that off with the GLM-5.2.

Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,

Ali [00:12:02]: Based on being number

Swyx [00:12:03]: Yeah

Ali [00:12:04]: Or it’s for something else.

Swyx [00:12:05]: Yeah. Which,

Ali [00:12:06]: Oh my God

Swyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,

Philip [00:12:14]: There’s a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.

Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there’s going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.

Quantization, Speculators, and Production Readiness

Ali [00:13:16]: Yeah. It was pure continued post-training

Philip [00:13:18]: Yeah

Ali [00:13:18]: If I remember correctly.

Philip [00:13:19]: Even in those cases, there’s still stuff you have to do. You have to redo the quantization work. You’re taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we’re not causing any regression in the model’s intelligence. And then we also have to train the speculator, as we’ve talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don’t know exactly the traffic that people are sending us, but we know what’s popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you’re getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there’s that process which you need the real model weights for. And then there’s of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there’s a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 had

Ali [00:14:53]: Sparse attention.

Philip [00:14:54]: Yeah,

Ali [00:14:54]: Yeah

Philip [00:14:54]: the DSA.

Ali [00:14:55]: Right. Which is brought from DeepSeek.

Philip [00:14:57]: Yeah. And

Ali [00:14:59]: So you can copy-paste then?

Philip [00:15:01]: It kind

Ali [00:15:01]: I don’t know how this works.

Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you’re right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn’t have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.

Retrofitting Vision into GLM-5.2

Ali [00:15:27]: We’ll be training the projector.

Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there’s the encoder, which is the part that looks at the image and turns it into latent information, and then there’s the projector which like

Ali [00:15:38]: You can say latent space. It’s okay.

Philip [00:15:41]: And then there’s the projector that maps it onto, the model itself, and then there’s the model weights. You don’t wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.

Ali [00:16:02]: That would be, yeah.

Philip [00:16:02]: Yeah.

Ali [00:16:03]: Can you show the training one?

Ali [00:16:04]: Like the way it groks

Philip [00:16:05]: Yeah

Ali [00:16:06]: Very interesting.

Philip [00:16:06]: And maybe

Ali [00:16:07]: That right there

Philip [00:16:07]: Maybe Ali, you should take it from here. You’ve got a better

Ali [00:16:10]: Ooh, double the sand

Philip [00:16:11]: Understanding of this than I do.

Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here’s a picture of a mountain. Can you describe what’s in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we’re trying to teach it is to translate the encoded. Like it’s already taken the encoder from Kimi K. It’s taken the image. It’

Philip [00:16:31]: Yeah. Frozen

Ali [00:16:31]: Frozen

Philip [00:16:32]: With adapter.

Ali [00:16:32]: Exactly.

Philip [00:16:33]: Yeah.

Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It’s just we’re trying

Philip [00:16:37]: Align

Ali [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he’s like, “Oh, can you describe what’s in this image?” And he’s like, “Oh, it’s a mountain,” or it’s a person or it’s a human, whatever the case is. But that didn’t cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn’t perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn’t get it, but it will say something like, “This is Albert Einstein.” Like it still understands

Philip [00:17:25]: Close enough

Ali [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that’s like really cool.

Philip [00:17:32]: Yeah. So, we’ve covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that’s very foundational work for anyone who hasn’t done vision work before.

Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questions

Philip [00:17:47]: Right

Ali [00:17:47]: Off the image and how much better you can get performance.

Philip [00:17:50]: Right. Right. Right. Yeah. But what’s, what’s so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It’s not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you’re running this model, you haven’t suffered any loss on your GLM-5.2 quality. If you don’t have an image, it’ll just behave exactly the way it used to. And ultimately

Ali [00:18:14]: Which in the inference code you literally do not include the other part, right?

Philip [00:18:18]: Yeah. You would just skip the encoder if you don’t have an image input.

Ali [00:18:22]: Okay.

Philip [00:18:22]: Just confirming.

Philip [00:18:23]: Yeah

Ali [00:18:23]: Does it affect a lot on the overall inference side? Like you’re not adding much, you’re adding a very small vision encoder. These are typically like

Philip [00:18:30]: They’re super fine

Ali [00:18:31]: Less than a billion parameters, right?

Philip [00:18:32]: Yeah. It’s, - There’s a little bit less standardization among vision encoders

Swyx [00:18:37]: Yeah

Philip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it’s a pretty, it’s a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.

Open Source Model Grafting and Franken-Merges

Philip [00:18:56]: And that’s, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that’s better than anyone

Swyx [00:19:05]: Yeah

Philip [00:19:05]: Can be individually.

Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take like

Philip [00:19:10]: Yeah

Swyx [00:19:10]: Layers from each model.

Swyx [00:19:11]: Does anyone do that anymore?

Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you’re doing auto-regressive token generation for three tokens, and you’re doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it’s not sparse, it’s not top K. So we find it better to like, okay, we’re gonna replace this, we’re gonna replace this layer with a layer from another model that’s using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That’s like, I feel like more and more becoming true.

Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?

Loop Detection, Race Conditions, and Non-Determinism

Philip [00:20:26]: Yeah. I think that there’s also a question of just, we can test a model to a pretty extensive degree, but we’re trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there’s going to be, so many more varieties of things given to it that you’re able to, discover and patch things. So it’s not just a, day zero process, it’s then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?

Ali [00:21:21]: What do you mean you don’t want your model outputting S?

Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.

Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it’s four times the same token, it’s probably collapsed.

Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?

Ali [00:21:48]: You want that?

Ali [00:21:50]: I think there’s a way that we have to handle it. I’m not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there’s a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.

Swyx [00:22:07]: Yeah.

Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2

Swyx [00:22:11]: Oh

Ali [00:22:11]: And I think it was DSV 4 as well. Like you’d just have like looping issues where like you literally

Swyx [00:22:17]: It

Ali [00:22:17]: Just have like S.

Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomly

Ali [00:22:21]: It just seems to be the one token involved.

Swyx [00:22:23]: Yeah. And it’

Philip [00:22:24]: Is there

Swyx [00:22:24]: And it’s only temperature 0

Ali [00:22:27]: No

Swyx [00:22:27]: Even at other temperatures

Ali [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.

Swyx [00:22:30]: That’s weird, right?

Ali [00:22:30]: It’s, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we’ll find that it fixes it. Or oftentimes this will only happen in an inference engine that you’re using like SGLang. But if you were to switch to vLLM, that isn’t the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It’s not like a weights problem. Like I’- we’ll say like, “Oh, it’s a problem with the quant. We did PTQ wrong,” right? But that isn’t, that doesn’t make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it’s, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.

Swyx [00:23:19]: Oh my God.

Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn’t. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We’re gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?

Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?

Ali [00:23:46]: Right.

Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won’t always get the same output.

Swyx [00:23:52]: Even-- But I’m surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.

Ali [00:24:02]: Well, yeah, true. Like I’m not, I’m not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that’s like ‘cause you want to do that because there’s

Swyx [00:24:12]: It’s like pipelining

Ali [00:24:12]: Expense. Exactly.

Swyx [00:24:13]: Yeah.

Ali [00:24:13]: But it’- But you don’t do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you’re designing a kernel and you want it to make it to be very fast, if you don’t test it extensively, you’ll, you’ll have certain threads access data points from registers before they’ve been written to by other threads

Swyx [00:24:36]: Yeah

Ali [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, and

Swyx [00:24:42]: And there’s no like borrow checker

Ali [00:24:45]: What does that mean?

Swyx [00:24:46]: Like Rust. Like the. If you’re trying to have like memory safety It sounds like a comparable problem.

Ali [00:24:52]: Well, yes, but you’re working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that’s what modular is supposed to do. I don’t know.

Quantization Quality and Vendor Fidelity

Vibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoder

Ali [00:25:07]: Right

Vibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarks

Ali [00:25:22]: Yeah

Vibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes into

Philip [00:25:27]: There’s a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you’re preserving all the outliers. There’s other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what’s gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn’t need the full million token context, for example, you can get them better performance. I don’t know if that’s exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.

Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it’s getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking here

Ali [00:27:41]: Yes

Philip [00:27:41]: Where they have

Ali [00:27:42]: They released an actual vendor benchmark.

Philip [00:27:43]: Exactly, yeah.

Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi’s benchmark.

Philip [00:27:50]: Yeah.

Philip [00:27:51]: So, with Reflect we probably

Vibhu [00:27:52]: This was a long time ago, right?

Philip [00:27:54]: No.

Ali [00:27:54]: Yeah, like three

Vibhu [00:27:55]: They also

Ali [00:27:55]: Four, five months ago

Vibhu [00:27:57]: This also happened with, I don’t remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have been

Philip [00:28:03]: Kimi Vendor Verifier.

Ali [00:28:04]: Yeah.

Philip [00:28:05]: Yeah.

Ali [00:28:05]: Yeah, ‘cause you, ‘cause you’d be pissed, right? Like if you’

Philip [00:28:07]: Yeah.

Ali [00:28:07]: If like if I’m a consumer and I’m using like Amazon’s endpoint for instance, and I’ve used Kimi and I’m like, “Oh my God, like this is bad,” I’m not gonna say, “Oh, Amazon quantized the model in a bad way.” I’m gonna say, “Oh, Kimi sucks.” Right?

Philip [00:28:17]: Yeah.

Ali [00:28:17]: So it seems like that makes sense.

Philip [00:28:19]: Yeah, they care. They care.

Vibhu [00:28:21]: Justifiably.

Ali [00:28:21]: Yeah, justifiably.

Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?

Philip [00:28:28]: Yeah.

Vibhu [00:28:28]: Like, is quantization always strictly worse?

Ali [00:28:30]: Well technically

Vibhu [00:28:32]: No

Ali [00:28:32]: It’s a lossy. Quantization

Philip [00:28:33]: Yeah

Ali [00:28:33]: Is a lossy, it’s a lossy implementation.

Philip [00:28:36]: Speed improves

Vibhu [00:28:36]: Speed improves.

Ali [00:28:37]: It the number, like

Vibhu [00:28:38]: No, I’ always look for inverse scaling laws.

Philip [00:28:40]: Yeah.

Ali [00:28:40]: Yeah.

Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.

Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,

Ali [00:28:52]: Yeah

Philip [00:28:52]: NVFP4 quant is like, two basis points higher than your

Ali [00:28:56]: No, it’s noise. It’s noise.

Philip [00:28:57]: Yeah, exactly. I’m like, yeah, it’s, it’s within. That’s why I always say within margin of error.

Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we’re barely inside of that to the worst, so we’re saying. But yeah, sometimes it’s just like, gives you a higher output score. But like Ali said, that’s noise. To my knowledge, you’re not necessarily making the results better. You’re just trying to, again, like keep your fidelity as close to 100% to the original model.

Layer Selection, KL Divergence, and Better Quantization

Ali [00:29:27]: There is, to your point, research that we did on MP. I don’t know if you are able to pull

Philip [00:29:31]: Yeah

Ali [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it’s a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It’s. You’re compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you’re losing some information, and you’re trying to minimize that. And so when I say that I’m gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don’t quantize modulation layers, and I don’t quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn’t have the. Yeah. It’s a long paper. I don’t know if I can find

Vibhu [00:30:25]: If there’s a part to search or it’s probably in the thread.

Ali [00:30:28]: It’s probably in the thread.

Vibhu [00:30:29]: Yeah.

Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that’s 20% more quantized than another provider, so you get 20% more throughput of it because there’s more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you’re probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it’s gonna be, ‘cause the more loss you introduce. That’s not exactly, not necessarily true. So yeah, doesn’t improve it, but can cancel out.

Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.

Philip [00:32:03]: But very interesting. Didn’t know this was a whole paper you guys put out.

Ali [00:32:06]: It’s. Fun fact, it was originally 72 pages, this paper, and then we decided

Philip [00:32:11]: Wow

Ali [00:32:11]: We can’t tell. We couldn’t release it. So it’s now 45.

Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what’s possible in terms of speedup? Like it’s like probably like the number

Inference Speedups and Benchmarking

Swyx [00:32:25]: Thing that people do wanna care about, and it’s something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?

Philip [00:32:36]: So what’s cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you’re in finance, you measure how much better you got in basis points. It’s like, “Oh, I got five basis points better, like twentieth of 1% better,” that’s huge news because everything is so optimized. When we publish optimizations, it’s 20%, it’s 100% it’s 200%. So there’s still probably like a lot further to go, honestly. Like you’ll, you’ll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.

Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would find

Ali [00:33:27]: And like 20%, tens of percent.

Swyx [00:33:29]: That’s. Yes.

Philip [00:33:29]: Yeah.

Swyx [00:33:30]: And now it’

Philip [00:33:31]: Tiny fractions

Swyx [00:33:32]: For those people interested, look up Andrew Lo’s paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.

Philip [00:33:48]: Exactly, and we’re at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there’s so many variables that go into it. What hardware are you using? How much load do you have on the system? What’s the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you’re looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there’s two tokens per second. There’s tokens per second, the throughput number, and the latency number.

Ali [00:34:31]: TTMT, yeah.

Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don’t.

Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that’s an 8X gain. That’s the order of magnitude that we’re working with in this space. We’re trying to make things substantially faster, not just go from like 70 to 90.

Swyx [00:35:38]: Are you saying you’ve. You have done that?

Philip [00:35:40]: So let’s say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you’re just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you’re, you’re probably, yeah, looking at that like 30 to 40. You think that’s like a reasonable baseline?

Swyx [00:36:12]: Right. Right.

Philip [00:36:12]: To get to something like 10X, there’s a lot of trade-offs that you’re making. If we’re running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It’s oftentimes maybe more of a four to six times improvement. But that’s the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.

Stacking Optimizations: NVFP4, Speculation, and Disaggregation

Ali [00:37:19]: It’s also, like, hardware dependent. Like, if

Philip [00:37:20]: Yeah

Ali [00:37:20]: If you have a thing where you’re serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.

Philip [00:37:35]: Yeah. Then you’re looking at, like, a two to 4X improvement

Ali [00:37:38]: Right. Right

Philip [00:37:38]: Depending on the inference optimizations. So yeah, it’s. Some of it’s, what’s the call, and some of it’s who’s the driver.

Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2

Ali [00:37:51]: Yeah

Vibhu [00:37:51]: On B200s

Ali [00:37:53]: Yeah

Vibhu [00:37:53]: Single node, right? What’s, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?

Ali [00:38:01]: Spectre quantization. Yeah.

Vibhu [00:38:03]: Spectre quantization.

Ali [00:38:04]: That’s, that’s, that’s like 95%. Like

Vibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?

Philip [00:38:23]: If you’re doing it up front, it’s quite a lot of work. If you’re doing it today, there’s going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we’re thinking about, like, what are the 2Xs we’re stacking, going from, BF16 to NVFP4 is, it’s not quite a 2X, right? It’s like. I think it’s about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn’t quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you’re able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that’s how it stacks up.

Ali [00:39:21]: Yeah

Philip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they’re doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you have

Ali [00:39:39]: Once set up. Once set up. Yeah

Philip [00:39:40]: Yeah, getting disagg working for the first time, I’m saying, of course, is very difficult.

Philip [00:39:44]: The marginal implementation

Ali [00:39:48]: Like, if you’re just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you’re wondering, “How can I just host it myself?” You don’t need to quantize the model yourself. There’s always gonna be, like, an open source quantized checkpoint. NVIDIA’s gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they’ve trained as well. You don’t need to train your own spec dec. You can just use that as well.

Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.

Ali [00:40:13]: Right. Right.

Vibhu [00:40:14]: What’s multi token prediction?

Philip [00:40:15]: Yes.

Ali [00:40:16]: I’m just

Vibhu [00:40:16]: Can you explain that?

Ali [00:40:16]: I’m just an expert.

Ali [00:40:18]: I can do it for you in case I get it wrong?

Vibhu [00:40:20]: No.

Vibhu [00:40:21]: Yeah, you should correct if we’re wrong, but their multi-token prediction can be used for self-speculative decoding.

Ali [00:40:27]: I’m not sure. I’m not gonna correct that.

Vibhu [00:40:28]: Okay. I’m semi-confident in that

Ali [00:40:30]: Okay. Yeah

Vibhu [00:40:30]: But someone can check. But it’s useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.

Ali [00:40:48]: Right.

Vibhu [00:40:49]: I was waiting for a mention of Dynamo.

Vibhu [00:40:51]: I feel like, that’s supposed to be the baseline that you measure against.

Dynamo, KV Routing, and Disaggregation Toolkits

Philip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.

Ali [00:41:17]: We’ve done a pod with Kyle

Philip [00:41:18]: Okay

Ali [00:41:19]: Kyle Cranin.

Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.

Ali [00:41:28]: But it’s just a router, it’s not like an optimizer layer.

Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.

Philip [00:41:49]: That doesn’t mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It’s more of a developer toolkit.

Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.

Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we’ve got to, we’ve got to benchmark against, like, what we’re seeing in the wild.

Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-Spec

Vibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.

Philip [00:42:31]: Yeah.

Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLE

Philip [00:42:35]: Yeah

Vibhu [00:42:36]: 524 on gram.

Philip [00:42:37]: It’s 55, would be disaggregation

Ali [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bit

Philip [00:42:44]: Yeah

Ali [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.

Philip [00:42:51]: Yeah.

Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.

Vibhu [00:42:55]: Medusa is quite old.

Philip [00:42:56]: Yeah, Medusa’s old.

Ali [00:42:58]: It was old.

Vibhu [00:42:58]: But is it in the book as a good, here’s

Philip [00:43:01]: Baseline

Vibhu [00:43:01]: Baseline vanilla understand it?

Philip [00:43:02]: Like you should know this.

Vibhu [00:43:03]: Like I read the paper, I’m like, “ it makes so much sense.”

Philip [00:43:05]: Yeah.

Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there’s DFlash, dSpark. There’s, there’s newer techniques even than EAGLE, although EAGLE is still very commonly used.

Ali [00:43:51]: SpecSpecta.

Philip [00:43:52]: Yes. Speculative decoding.

Vibhu [00:43:54]: What can

Ali [00:43:56]: Oh, it’s a paper by Tri Dao and it’s like, it’s doing speculative decoding

Vibhu [00:44:00]: Huh

Ali [00:44:01]: For the speculative decoder.

Philip [00:44:02]: Oh, in spec- oh my God.

Ali [00:44:02]: It’s literally just an another. It’s like, yeah, that’s the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it’s almost like in our mind at least, it’s almost as complex as training GANs. Like it’s like a very delicate balance and oftentimes you, it’s just but yeah, it’s literally speculative decoding on speculative decoding.

Vibhu [00:44:21]: Speculative.

Ali [00:44:22]: Yeah. We saw this paper.

Vibhu [00:44:24]: It’s interesting, right?

Ali [00:44:24]: Yeah.

Vibhu [00:44:24]: I wouldn’t even expect it to be very particular to train, I would

Ali [00:44:29]: Right.

Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.

Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It’s like, it’s like almost like the iPhone auto predict version but for a normal model, right? Like you’re just, you’re just, generating three tokens and you’re like, okay, I’ll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?

Ali [00:44:53]: The other question there is what are the size of speculators? So say for

Philip [00:44:58]: Right. It’s like a billion parameters.

Ali [00:45:01]: Like for MiniMax, it’s. Yeah. It’s like one layer. It’s like one 60th of the original model usually.

Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.

Philip [00:45:10]: Speculative

Ali [00:45:11]: Speculative

Philip [00:45:11]: Decoding.

Ali [00:45:13]: No, it’s, it does seem like how, when do you stop? But then it also seems like if you’re able to train spec-spec decode for instance, right? Like if you’re able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model’s gonna predict, then why not just use that smallest model directly, right?

Vibhu [00:45:34]: Yeah. This is

Ali [00:45:35]: Like it seems like

Vibhu [00:45:35]: Adjacent to the routing problem.

Ali [00:45:36]: Right.

Vibhu [00:45:36]: Yeah.

Ali [00:45:36]: Right.

Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you’re running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.

Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it’s the same thing, it’s just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?

Local AI vs. Data Center Inference

Philip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it’s how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it’s how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don’t touch, in the pruning, in the distillation, in the, layer removal. There’

Ali [00:47:42]: Layer removal matters less.

Philip [00:47:43]: Yeah. There’

Ali [00:47:44]: No one loves pruning really.

Philip [00:47:45]: Yeah. Well, but the, but they do

Vibhu [00:47:46]: Which is surprising, right? But that’s, that’s a whole different thing

Philip [00:47:48]: Just to fit something on the laptop.

Ali [00:47:50]: Right.

Philip [00:47:50]: So yeah, it’s a, it’s an interesting, it’s an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.

Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we’ve seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don’t need decrease the storage that much. You don’t need to do, FP4 KV cache. You don’t need to use a requant. There’s, there’s, there’s better optimizations to be made. But on Edge devices, it’s extremely important, it’s extremely useful. So, seems to be, like, different optimizations there, but then they’re all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with both

Philip [00:49:18]: Principles.

Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.

Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.

Ali [00:49:35]: Yeah, this is the Exo Labs guys.

Philip [00:49:36]: Yeah. You have, a number of, Mac Minis stacked up.

Philip [00:49:41]: There’s, the inter. They. One thing that I think we both have to deal with, although they have to deal with a lot more is the interconnect between machines. Which is why, like, one thing that we do a lot is work with tensor parallelism.

Philip [00:49:56]: And that’s where, you are using all of the, all eight GPUs, and sharding the model across it. Tensor parallelism is not a good fit for local AI because it assumes a very high bandwidth interconnects like NVLink. Was, they might be forced to do something like pipeline parallelism, which we’re never gonna do unless we’re doing some kind

Ali [00:50:16]: Yeah. For image

Philip [00:50:17]: Multi-node inference.

Ali [00:50:18]: But since you mentioned it, I wasn’t sure if we were gonna cover it, but let’s briefly explain tensor parallelism and expert parallelism, since you have very nice images.

Tensor, Expert, and Pipeline Parallelism

Philip [00:50:25]: You wanna pull the book?

Ali [00:50:26]: Yeah.

Philip [00:50:26]: Yeah. Let’s, let’s get

Ali [00:50:27]: So I just wanna show a few images.

Philip [00:50:29]: Yeah. Shout out to Luke from Baseten’s design team for making these beautiful images. Oh, that’s a, that’s. Before we get into this, just one other difference is we talk a lot about the active parameters of a mixture of experts model, and for local inference folks, that matters a lot because if you have a batch size of one, you’re only activating that many parameters. When we

Ali [00:50:51]: Yes. I was gonna

Philip [00:50:52]: Inference in the data center

Ali [00:50:52]: I was gonna bring that in the diffusion conversation.

Philip [00:50:54]: Yeah.

Philip [00:50:55]: Yeah. We, I, when we go through like a MoE model, and we host it, for an API, we assume that all parameters are gonna be active because

Ali [00:51:06]: You’re batching

Philip [00:51:06]: Throughout your batch

Ali [00:51:07]: Yeah

Philip [00:51:07]: You’re gonna, you’re gonna hit everything. Cool. So broadly, tensor parallelism you can do with any model. Expert parallelism, you can only do with MoE models. Effectively all models today are MoE models, that are,

Ali [00:51:21]: Sort

Philip [00:51:22]: At least all models large enough that you would care to parallelize them across multiple GPUs. So that’s, that nuance is less important now. With expert parallelism, the idea is you put the entire expert on a GPU. Generally, you have more experts than GPUs, so you might put like N experts per GPU, like eight experts per GPU or whatever. And then you replicate the router, which the router is very small, across each of the GPUs. And then by moving the generation from expert to expert, with each expert being inside a GPU, they’re not competing for resources. You massively increase the throughput that you’re capable of doing, and the, GPU connection is not as important ‘cause there’s not as much communication. Tensor parallelism requires that you are able to do this like all gather, all reduce. So you shard the model across the GPUs entirely. And then for each step, you’re combining the results of each of the GPUs, which is why the interconnect matters a lot, and it is generally. Of course, this is a, this is a very high-level generalization. There’s a lot of places where this is not correct. But generally, TP is helpful for latency, and in many cases, you will use some combination of these two parallelisms, across the model rather than just, like, picking one or the other. Do you wanna add some color there?

Ali [00:52:50]: Like, yeah, usually, like in a model, it’s not. They’re not mutually exclusive. You do tensor parallelism and you’ll do expert parallelism. Pipeline parallelism less solely, it seems to me like we never use pipeline parallelism.

Philip [00:52:58]: Yeah. The only reason you would have to do pipeline parallelism, which is where you separate like different layers and you put like half the layers on one hardware and half on another, is if you are forced to do multi-node inference, because a model is bigger than you have the. Like let’s say, let’s say you’re doing a deployment on H100s for whatever reason, and you’re putting a trillion-parameter model on there. You have to use multiple nodes of H100, and so you. - Because the interconnect is so slow between the nodes, the only viable way to parallelize there is pipeline, but then you would do expert and tensor within each node.

Ali [00:53:36]: And the limiting factor for H100s is HBM?

Philip [00:53:39]: Yeah. They just don’t have enough

Ali [00:53:40]: How much? What’s the magic numbers that we need

Philip [00:53:43]: Like on a B200 is 180 gigabytes per GPU, and then a node of eight, so you’re talking like 180 times eight. And the FP4, so each parameter takes half a byte, so that’s 800 gigabytes. On a H100, it’s like 140?

Ali [00:53:56]: It’s 80.

Philip [00:53:57]: It’s 80?

Ali [00:53:57]: Yeah.

Philip [00:53:57]: Oof.

Ali [00:53:58]: Yeah.

Philip [00:53:58]: I’m old. I’ve been doing this a long time. I remember H100 specs.

Ali [00:54:04]: Yeah.

Philip [00:54:04]: No, so one thing

Ali [00:54:06]: You wanna tell me about the T4s?

Philip [00:54:07]: The T4s. Oh my God.

Ali [00:54:08]: Let me tell you what it was like to run a model on a T4 back in the day.

Ali [00:54:12]: One thing I was surprised to see that more people didn’t do, Jamba. I don’t know if you guys remember Jamba from AI ‘21. They would specifically pick a hardware, and then they designed the arc dimensions for the hardware, and then it would saturate the hardware. Like, it makes sense. And like, somehow all these models don’t do that.

Hardware-Aware Inference and Auto-Tuning

Philip [00:54:32]: Don’t they do this for the training side, though?

Ali [00:54:35]: I don’t know.

Ali [00:54:36]: Sorry,

Philip [00:54:36]: Training. For training the model.

Ali [00:54:37]: Like deciding which GPU, which

Philip [00:54:39]: Yeah. Well, how

Ali [00:54:40]: Yeah, they do And with training, it’s more of like a math. Like you can run the math- Yeah and see the flops and maximize it. With inference, it’s more of like an auto-tuning, like if you like GPU kernel auto-tuning. But like it’s like you define that, “Oh, I have two GPUs. I can do TP1, TP2, EP1, EP2,” for instance, right? And you. So that gives you like total of like two squared combinations, and then you just like you shadow the same traffic, like real prod traffic, and you just see which configuration gives you the best TPM and TPS, and then just use that. I don’t like the fact that it’s, you cannot reason about which one’s gonna give you the best performance or that there isn’t one specific configuration that’s always best. But it seems like auto-tuning is just the way that you find the best one. And with kernels and GPU kernels, it’s much of the same. After you design your kernel and you design your configuration, how many threads do you launch? How many, how much shared memory do you use? You just auto-tune. You just sweep the parameter space on the side, and this is the best one empirically. But yeah, but they are combined. They’re not just entirely- Yeah like separation. There’s a few bits of training that are like hardware targeted. If you look at, for example, NVIDIA Nemotron models, they run very well on Blackwell. That’s, that’s unsurprising. So there’s some degree of that, but I think that most open labs are trying to make models that can be run on as wide of hardware as possible rather than targeting just like a single chip. I see. For usefulness. Yeah. Okay, one more thing while this chart is still up. All gather, all reduce is expensive. One of the things that is a movement in Silicon Valley is mega kernels, just keep fusing kernels. I don’t know. Is it that simple? Well, I, like a fused kernel can’t save you. Like here with tensor parallelism, you’re. The half the matrix is on one GPU and the other half is on another, and if I need the entire matrix in order to do like a nonlinear operation in the next step, which is, for instance, like if I’m doing attention, I need the softmax, or I need to do like exponentiation, I need to have the entire row. So I need to know what the partial result was from GPU 2 and what the partial result was from GPU 1 in order to be able to do the softmax in the next stage. So I, like I have to make them communicate with each other, even if I had a fused kernel, because of the nonlinearities within each one. Also with like mega kernels, like honestly, I’m, I’m, I’m very bearish Ooh on, I’ll be honest. Like- Please. No, it’s just like mega kernels, it was a good research direction, and it seems like a very. Like intuitively, theoretically, it’s nice. Like, oh, like you have a lot of launch overhead from launching- Just- one kernel- Yeah, just keep fusing it moving the data. Just fuse everything together. But yeah, but like the kernel complexity itself is very difficult to write a very optimized mega kernel. It’s, it’s very difficult to do so. And even the, like not to name any companies, but like even the companies that have worked or people that I’ve spoken to who work at companies that do fused mega kernels, they very often don’t end up running those in production because the TensorRT-LLM and modular kernels that launch are faster because you can optimize each individual component, and you can just have them parallelize with each other. With the Rubins, I don’t know if you guys saw the Rubins Twitter post yesterday, but they’re also, Rubins? Like- No, like Rubin, like the GPU. NVIDIA GPU the, yeah, GPU. Yeah. They have a Twitter account for Rubins only? No. Okay. I was like, “What are you talking about?” Yeah. Sorry. One of the tech leads at NVIDIA is like launched a Twitter post said like, “We’re pulling the curtain on Rubin, and here’s the, here’s the specs.” And the third tweet showed, like not to get too technical into it, I and I need to read it much more, but the GPU is designed in such a way that it kills mega kernels. You don’t need to use mega kernels that much anymore. So it seems like that entire research field goes into like, won’t be continued, but yeah. Can I speculate about Rubin for a minute, please? Go. I’ve been through now, we And by the way, they are covered in the book. Yeah. But yeah, they- Well, they’re covered in the book in the sense that like I am aware- The Wikipedia entry from the blog post- Yeah that Rubin is going to happen in the future. And you even had the name of the one, Feynman. Yeah, it’s like, “Hey, this is gonna “ I was like, “This is very up to date.” Like I’m trying to future-proof this thing, okay? I don’t wanna publish a new one until like next year or something. Anyway, so we were discussing the degree to which I am old. And I’ve now been through three hardware launch cycles. I’ve been through the Ampere launch cycle, the Hopper launch cycle, and the, Blackwell launch cycle. Now, when I say launch cycle, I don’t necessarily mean like the actual shipping of the hardware. Like Ampere’s were racked up well before I got in this industry. But there is a lot of time between hardware being racked up and hardware being feasible for inference. So if you look at like the original vLLM and SGLang, vLLM especially, like that was written targeting Ampere and then had to be updated for Hopper, updated for Blackwell. With each of these cycles, it becomes faster and more urgent, but also substantially more complicated. When I look ahead to, what’s going to be new with Rubin, I think that like Dynamo gives me a lot of technical hints around like what kinds of work is going to be very valuable. We’re continuing some trends from Blackwell, right? NVFP4 is big. The amount of compute that they have behind NVFP4 tensor cores is massive. We’ll, we’re gonna talk about video, I think, at some point, and that’s the big barrier there. You’ve got, much faster memory bandwidth, but which was the same thing that made Blackwell so good. But the big thing is more systems thinking. You have more emphasis on the CPU to GPU interconnect, more emphasis on the interconnect between GPUs, and when you look at Dynamo, it’s a system entirely designed around how do I move the KV cache to where it needs to be when it needs to get there? So I think that themes around like KV cache offloading, KV-aware routing, and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a like CUDA kernel problem, but also like a very traditional hardware infrastructure problem, which is something, we’ve been building toward for a long time, and something that’s like very exciting to me because we’re gonna see

Mega Kernels, Rubin, and the Future of GPU Systems

Philip [01:00:55]: Multiple domains colliding and the ability to reason from the kernel level, like up to the hardware level and back down is going to be very valuable.

Ali [01:01:05]: I will take what Phil said one step further, into that. It’s, I think, trending towards becoming exclusively an infrastructure problem, where like problems of PD disagg, Training, spec dec. But troiting kernels is not going to be much of a problem because the GPU is moving more towards being an ASIC, where it’- you’re just, you’re just trying to orchestrate what happens on the GPU, but you’re not controlling it thread by thread level. And you see this with like QTAL, QDSL, like you’re, you’re just working at levels of like tiles of data, but you’re no longer working at controlling what each thread does on the GPU that’s being taken care of for you. So do you agree that a GPU and future GPUs are trending more and more towards becoming ASICs that just need to be launched and then they do the data operation based on your conversations with other people?

GPUs, ASICs, and Specialized Hardware

Swyx [01:01:50]: Oh, yeah, no. That is a section of the market.

Ali [01:01:55]: Right.

Swyx [01:01:55]: And ASICs can do, a lot more performance for only their workload.

Ali [01:02:01]: Right.

Swyx [01:02:01]: And the G in GPU makes them continue to be very general.

Philip [01:02:05]: Yeah. The, - I think that there’s like a spectrum

Swyx [01:02:08]: It’s graphics,

Philip [01:02:09]: Yeah.

Swyx [01:02:09]: I keep saying this, I have to correct myself in case people come at me for getting the G wrong.

Philip [01:02:14]: Yeah. It’s like, it’s like a spectrum, right? Of a very general purpose compute to something like a Taalas, where you’ve got the hardware built for a specific set of model weights.

Ali [01:02:26]: The weights burned

Swyx [01:02:27]: The weights

Ali [01:02:27]: Into the chip.

Swyx [01:02:28]: Yeah.

Ali [01:02:28]: No loading.

Philip [01:02:29]: I don’- I wouldn’t say that like, that we’re, we’re, we’re going all the way there. It’s more like along the spectrum, it’s a step in the direction of more specialization within the hardware.

Swyx [01:02:40]: Yeah. I’m curious, I feel like he was driving towards something.

Ali [01:02:43]: My point is being bearish on. Like, you say, like everything else apart from burning the weights into the chip. Burning weights into the chip is like impractical because you wanna fine-tune, you wanna optimize, you wanna quantize, you wanna release new checkpoints of the model. If it’s burned into the chip’s useless in like a month or two, right? My point is: How can you - like seeing NVIDIA more and more specialized, like take its GPUs from a general programming paradigm where you’re just-- it’s a general computer that you can use to program threads, and with every new generation, you’re putting more and more specialized instructions, specialized tensor cores, specialized, MMA instructions, things that will allow you to just control it almost as an ASIC, almost as a collection of ASICs.

Ali [01:03:22]: How can you look at this trend and then still be bullish on companies that are coming up with ASICs for AI?

Ali [01:03:30]: In the sense that, in the sense

Swyx [01:03:31]: Yeah, because they’re, they’re

Ali [01:03:33]: Right.

Swyx [01:03:33]: They’re, they’re evolving towards that direction.

Ali [01:03:34]: They’re almost evolving towards - Like as an Rubin, comp- Like compared to Ampere or, a T4, Rubin is an ASIC. It is, it’s just a thing that is used

Swyx [01:03:47]: Programmable ASIC?

Ali [01:03:48]: Yeah. It’s like - Yeah, like you can program, like I, like. It’s very controversial to call it an ASIC. It is a GPU. It is - It is general. It does have threads. I can write CUDA to control it and change its operations. But it has the systolic arrays and tensor cores and TMAs and tensor memory, and it has these things that are almost exclusively useful for loading model weights. It has, tensor core instructions that are almost exclusively shaped around the head dimensions of models that exist in the market today. To say that you’re gonna come up with an ASIC and you’re gonna etch something into it, well, but the next architecture is gonna be useless.

Philip [01:04:19]: Yeah, I don’t know. I don’t know. I think that the thing to remember is just how long these hardware cycles are.

Ali [01:04:25]: Yeah.

Philip [01:04:25]: So if a chip is coming out today, that means the design process for it was kicked off years ago. And they’- at NVIDIA, they’ve done a very good job of predicting where the market is going to go and,

Swyx [01:04:38]: They have the most information

Ali [01:04:40]: For sure.

Philip [01:04:41]: Of course. But if you look at, there being public open source model architectures that look more or less like early versions of the one today, Rubin’s honestly the first chip that was fully built in that world. And so you can see a lot of the understanding of the shape of the workload that this chip’s going to be asked to do in the way it’s designed.

Swyx [01:05:04]: Yeah. Okay. So I’m not gonna be the best person to directly answer those questions. I think these are very fair questions that - the first one that’s based on Rubin that like I’ve, heard artic-articulated so well. I do think that, I will make a case for a vertically integrated model lab ASICs.

Swyx [01:05:24]: So like the OpenAI, Broadcom, what-whatever, Jalapeño

Philip [01:05:27]: Sure. Yeah

Swyx [01:05:28]: Chip, which like totally makes sense. Like, so - we first had this on the pod with, Martin Casado, where he was like, “Look, if you have a trillion-dollar or five hundred billion dollar training then take fifty billion of that and make a ASIC. Like it’s fine. Like you will get more than ten percent efficiency from the ASIC.” And like that makes sense.

Philip [01:05:46]: Right.

Swyx [01:05:46]: Right? So like a model-specific chip, yes. But ASIC companies, the interesting thing is I feel like you are focus-- you’re hyper-focusing on like you say, like the Taalas stuff.

Philip [01:05:58]: Right.

Swyx [01:05:58]: They are doing a lot more like, surface area engineering or like the actual allocations of memory and hardware and like the communication between chips that, probably still won’t be touched by Rubin, but I don’t know the details.

Philip [01:06:14]: I see. I see.

Swyx [01:06:15]: They-- Typically, they often talk about things that I would expect to have bigger orders of magnitude than would be programmably accomplished by whatever Rubin does. But who know-- who knows?

Ali [01:06:26]: No, I see.

Ali [01:06:28]: Yeah. It seems,

Swyx [01:06:29]: Yeah, like think about what - what are the real blockers to ten x to one thousand x faster inference. It is not the stuff that can be rearranged, just within the existing GPU design.

Ali [01:06:41]: Inter communication.

Swyx [01:06:42]: Yeah.

Ali [01:06:43]: Okay.

Swyx [01:06:43]: Like these guys are aiming for three hundred thousand tokens per second. They’re not f*****g around. Like,

Ali [01:06:49]: Might have to put on some X6.

Philip [01:06:50]: Maybe. I think, it is interesting to me that you’re so bearish on so much of this kernel engineering work, given how much of it you’ve been doing recently.

Ali [01:06:59]: Right. Right. But like the more I do it, the more it just seems to me that

Swyx [01:07:01]: It’s not mega

Philip [01:07:02]: I would also add like

Vibhu [01:07:04]: There’s generations of models being out, right? I think on your guys’ end, you see a lot of, okay, one day it’s GLM, Kimi, DeepSeek, MiniMax, throw in the others. Some are doing completely different stuff, right? Gemma, no encoder. The latest thinking machines is all from scratch. But when you look at the other side, like how long have we been on the GPT-5 generation, right?

Philip [01:07:26]: Right.

Vibhu [01:07:26]: They’ve been serving that thing for quite a while. Sure, there’s maybe more training. There’s, there’s different checkpoints, but like you can squeeze quite a bit out and you do a multi-billion dollar train run. If you can make it X percent more efficient, they serve it for a while. Same with, say, the Claude 5 set, family, right?

Philip [01:07:44]: Like if they release a new model, like if they release GPT-6 now or whatever

Model Longevity, Open Source, and Enterprise Reliability

Vibhu [01:07:47]: Yeah

Philip [01:07:47]: And they release a new model every year, and - well, we don’t know, but if we assume that they’re changing some bits of the architecture and not just doing like post-training, like you’re gonna be spending fifty billion dollars a year every single year coming out with new ASICs for the model and throwing out the ASICs of the previous year away.

Vibhu [01:08:03]: Yeah. Yeah. Easy.

Swyx [01:08:05]: So I think, okay, I would slightly disagree based on my again,

Philip [01:08:09]: Yeah

Swyx [01:08:09]: It’s all secondhand, on the longevity of a model.

Philip [01:08:12]: Right.

Swyx [01:08:12]: There’s still people out there using 4o.

Vibhu [01:08:14]: Yeah.

Swyx [01:08:14]: Yeah, Llama. Not Llama 2, but Llama 3. I still see Llama 3 workloads.

Vibhu [01:08:18]: Yeah.

Swyx [01:08:18]: Because if it’s done, if it’s trusted, don’t change it.

Vibhu [01:08:22]: If it works.

Philip [01:08:24]: Which is one of the promises of open source, right? Like the whole 4o, save 4o movement. Like you don’t gotta have a save Llama 3 movement. You just gotta have an eight one hundred somewhere.

Vibhu [01:08:34]: I think at some point there’s also the question of, okay, if a model can do enough and use enough tool calls and be agentic enough, can it just web search, tool search write code? Do you really need to keep squeezing more? We will because you guys will make it cheap and fast and smaller, and I can swap it in. But at some level, like you give me GLM-5.2 today or say whatever 120 B model, I can run with it for quite a while, right?

Philip [01:08:59]: This is assuming like you don’t need intelligence.

Vibhu [01:09:02]: I think there’s a lot of intelligence where we

Swyx [01:09:03]: You need reliability and predictability. Like I’m in enterprise like like this is tried and tested. It is signed off by like my five thousand stakeholders.

Philip [01:09:11]: Right.

Swyx [01:09:11]: Like I’m not touching it.

Philip [01:09:12]: It runs a batch job every and I like the results.

Swyx [01:09:16]: Yeah.

Philip [01:09:16]: The results are predictable. Yeah.

Vibhu [01:09:18]: Yeah. It doesn’t make sense to keep using them. Like stuff gets sparser, cheaper, better.

Philip [01:09:23]: Right.

Vibhu [01:09:23]: But that doesn’t mean that old models, GLM 50 isn’t usable, right?

Vibhu [01:09:28]: If we hit a stall, say, for whatever reason, there’s still a lot that can be squeezed out.

Swyx [01:09:34]: We’re gonna run out of time. I did wanna also make sure. Yeah. Yes, we happen to have this diagram. Pull. Compare this versus any Cerebras diagram, right? I don’t think Edge10, medics have put out public, charts yet. But the complete the real estate is very different. The size is very different, right? This is not wafer scale, right? This there’s probably like, I don’t know, a few hundred of these on a wafer. I don’t, I don’t know how big

Philip [01:09:55]: Right.

Swyx [01:09:55]: The comparison is. But like, it is a, it is a very like real estate allocation

Vibhu [01:10:00]: Yeah

Swyx [01:10:00]: Difference.

Philip [01:10:01]: Few dozen, I would say.

Swyx [01:10:03]: Few dozen. Yeah.

Vibhu [01:10:03]: Before we move from hardware, I have two quick questions. One, the latest Kimi, which is really big, three trillion

Kimi Scale, GB300, and KV Cache Limits

Philip [01:10:09]: Yeah

Vibhu [01:10:09]: Doesn’t fit on most hardware on single node.

Philip [01:10:12]: Yes.

Swyx [01:10:12]: You need GB300 to fit it on a single node.

Vibhu [01:10:14]: You need GB300 or AMD.

Philip [01:10:20]: It’s simple math. NVFP4, two point eight trillion parameters, one point four terabytes. The GB300s have, two hundred and eighty-eight gigabytes each. So across eight of those, you have enough room for the model, and honestly like. So the other thing with GPU VRAM math is you have to leave space for the KV cache, and that’s going to depend on, to some degree, on the context length. So when a model is both has a very large number of parameters and a very long context length, you’re like fighting over space. Which is why, the KV cache offloading, would become like a more salient topic, I think, with these huge models. ‘cause you just, you’re very crunched for space.

Vibhu [01:11:10]: With the Rubin, you now have what? NVL 72 rack

Philip [01:11:15]: What?

Vibhu [01:11:15]: 20 terabytes of your

Philip [01:11:16]: Yeah. Now you still have NVL 72 on, Blackwell as well, but, you can’t necessarily assume you’re gonna do inference on that.

Philip [01:11:24]: There’s a whole lot more 8X racks in the world than there are NVL 72s.

Vibhu [01:11:30]: Yeah. My last quick question on hardware was, do you notice anything with hardware generations for new trained base models? So one of the things you said for efficiency is you can swap hardware. That’s one of the 2X gains. When we see new stuff coming out training-wise on Rubin, any changes on logs? Does this affect what type of models we will be seeing when these are more available? And can

Philip [01:11:56]: They get bigger. Like people understand the ceiling that you have in terms of how many parameters of a model you can run, given the latest inference hardware, and that forms a ceiling. And so, for example, when DeepSeek R1 came out, it was six hundred and seventy-one billion parameters, which at the time was really huge and I think did a lot to push us to really quickly adopt Blackwell and get good at serving on Blackwell. So yeah, it’s, it’s mostly in my mind about, model size and then about matching the architecture and the native quantization to the target hardware, like we talked about with like, all Nemotron models or NVFP4, for example.

Vibhu [01:12:42]: So we talked a lot about LLMs.

Video Diffusion, Attention, and Autoregressive Video

Vibhu [01:12:46]: You have a lot more in the book. What about audio, video? What’s the other side of inference engineering? Ali, you’re pretty big in video diffusion.

Philip [01:12:53]: Video diffusions, I think, are like they’re just shaped. A lot of the stuff that you can think about, reason about with LLMs being autoregressive. With video diffusion, it’s, it’s not the case. For instance, you don’t

Ali [01:13:04]: You don’t do batching. - every request just comes in on one GPU and it serves one GPU. You don’t have to shard. The models are a lot, are a lot smaller, like Wan 2.2, for instance, is a twenty billion parameter model. You don’t need to worry about. So it’s like orders of magnitude smaller than the best LLMs. And it’s one of those spaces where the open source models are. Like with LLMs, we see Kimica 3 is almost comparable to, Mythos or like GPT 5.5. The difference between the best open source LLM and best open closed-source LLM is very small. Like it used to be six months. I don’t think it’s six months anymore. I think it’s like almost on parity. Video models are definitely not. There’s a huge gap. If you look at the best video that you can generate today with an open source model like Wan 2.2 versus something like with Kling or Veo, difference is night and day. So it creates this disparity where media companies will choose to go most of the time to closed source models.

Ali [01:13:58]: For instance if I were to tell you, “Hey, I can generate an entire three-hour movie for you with this model, and I’ll optimize it so that you only have to pay me ten dollars.” But if they were to do it on a closed source, they’d have to pay a thousand dollars, which is a hundred x. Like I’m a hundred x cheaper, but it’s still a thousand dollars. They’re still gonna choose to do all of their cuts with Veo and Kling. So the. It’s like a chicken and egg cycle where less demand causes less innovation in the field, causes, less open source checkpoints to be released. And some of the labs that were releasing open source models like Wan will have closed sourced their latest models, like Wan 2.7 is not open source. We’re still on Wan 2.2. The challenge with video models especially is the number of tokens. So video models, you want to generate a high quality model, a high quality video. So let’s say you’re doing sixteen frames per second, that’s like the absolute minimum you’ll do, and let’s say you’ll do like 480p video. So you can think about your like dimensions and I think I have like a good, just like a diagram that shows the number, the sheer number of tokens, right? Let’s say you’re looking at like just one video of like, Sparta 300 or whatever. So let’s say we’re looking at like four frames, right? Those four frames of that video, if you go just. If you’re doing full attention, if you go a bit up, like you’re looking at, 480p by 720 by 81 frames in just five seconds, because 16 FPS by five, right? And then you compress it down to latent space, but you’re still doing 30 by like 50 by 21 tokens.

Vibhu [01:15:25]: Yeah.

Ali [01:15:25]: Which means that for attention, for just five seconds, you’re running attention on 35,000 tokens, right? So the attention becomes such a huge bottleneck. And because it’s O(n²), if you’re doing like-- if you extend that to like ten seconds, well, it’s just squared, 20 seconds, 30 seconds. So to generate a good cut scene of like one minute, it’s almost impossible to do within the same compute time. And it’s just, it’s, it becomes unfeasible. You can’t do it. And so you end up with moving towards two directions. Either you decide to do attention on the entire video at once, in which case you are forced to do sparse attention. So if you scroll back down to the origin, the video image, like you can see whereas on the left, for instance, I would be doing full attention where every single token in that Sparta 300 scene attends to every single other token, as you can see the sheer number of like red patches. On the right, I’m only attending to each token only attends to like the top K or top 12.5% that’s important to it, which can be like spatial. So like, the token that represents the crown attends to like the head, the face, and then the head on the other frame and the previous frame, temporal locality, spatial locality, that thing. This results in terrible video quality and the whole point of the post or the article here is to show like how you can train and you can do all these things, but you will still suffer in your quality a little bit. So you end up with one of two things. Either you bite the bullet, you have huge compute, and you do full attention over like a million tokens because you’re trying to generate like two minutes of video, or you move towards autoregressive video. Autoregressive video seems to me like that is the bet that the future’s gonna be making, but there are no good open source autoregressive video models out there today. And that seems to be the. If you want to get like an hour movie, if you want to see video models generating like an, like, Hollywood level movies, they have to be autoregressive in order to exceed that five second frame. Or there has to be some insane leap that happens in compute that allows us to do full attention over like millions of tokens at the same time in a, in an efficient manner.

Vibhu [01:17:10]: Even millions of tokens, it’s like you’re, you’re quadratic, so you’re gonna get there really quick.

Ali [01:17:15]: Right.

Vibhu [01:17:15]: I think, can you explain the pros and cons trade-offs of autoregressive? So one that comes to mind is, the consistency across frames.

Ali [01:17:23]: Right.

Vibhu [01:17:23]: You will. Ten minutes into generating autoregressive diffusion, you’re gonna forget. But what are pros and cons of this?

Ali [01:17:30]: Well, like autoregressive LLMs, you can take a lot of your. Oh, sorry, autoregressive diffusion models. You can take a lot of your optimizations that we discussed with LLMs, like spec dec and stuff like that, and you can apply it there. And you can, if you have a very high quality scaled up model, there is no reason why I can’t stream the outputs as in I can show you the first frame and then I’m like GPT back in 2022 when you were. Like now it’s almost like shots the text, but back then you could read and it’s generating as you read. With video models, you can watch and it’s generating as you watch. You it generates the frames and so token by token generation will allow us to scale a lot up and apply the attention mechanisms there. The downsides is every single autoregressive video model is s**t. It’s just terrible quality. If I, like, it’s just if you put, if you put the quality of any opens like Wan 2.2 versus any other autoregressive model, you can see like a video generated by Wan 2.2 is like, a cat and dog fighting. Autoregressive model will give you like degraded Tom and Jerry quality. I don’t know. The solution to generating long output then becomes, “Okay, we’re not gonna use autoregressive model. We’re gonna.” If you look at some of the things that like Grok Imagine or Grok Video does, and they do it really well, is they’ll, they’ll try to stitch these, seven second chunks together. And so you generate seven seconds and then you’re like, “Okay, I’m gonna. Can you extend this video?” And they’ll chunk two videos together. Open source doesn’t seem to have the tricks that they have there and by definition it’s closed source. We don’t know what they’re doing. But the closest you can get is taking the last frame of a video and feeding it into like a text and image to video where it will take the text, the prompt, and it will take the image of the last frame, and you’ll ask it to generate the next five seconds. And that’s like how you can extend this level of a model to generate like a movie, where you’re just, you’re constantly streaming frame by frame. But you get a drift. So you start with like you take the image, and then you generate a video, and then that next five-second video is like lower quality, and the third chunk is like even lower, and the fourth chunk is even lower. And like sometimes you’ll see things where like the new video is like just ever so slightly darker than the first one, and the next one is darker than the second one until like twenty-five seconds and you have black screen.

Ali [01:19:31]: Like it’s just. It’s, it’s - We tried to have a demo that would show this, but it was like-- it was extremely embarrassing to show. Like we just decided not to because it seemed to like. But it is, I think models will get there. They just need to, in my mind, scale up significantly and move towards being autoregressive. But the training techniques don’t seem to be clear there.

Swyx [01:19:50]: For those - who are interested in Grok Imagine, we did a pod with Ethan Ha from that team

Ali [01:19:54]: Right.

Swyx [01:19:55]: Who dropped a little-- a few hints, but not that not enough that we can fully reconstruct everything.

Ali [01:20:00]: Right.

Philip [01:20:00]: Specifically on this part, - he explains a bit about that.

Swyx [01:20:02]: Yeah. So we talked about memory and, longer context and all these things.

Ali [01:20:06]: But as far as I know, they’- it’s not autoregressive, even though like no one in industry is autoregressive.

Swyx [01:20:11]: Yeah.

Ali [01:20:11]: It seems to be, yeah.

Philip [01:20:12]: The key thing to understand between a autoregressive model and a diffusion model is that diffusion attention goes in both directions, while autoregression, it only goes forward in the sequence. So that’s why you see this like going off the rails behavior, both in. If you naively construct a video generation model as simply generating a linear sequence of frames, you can’t then go back in that sequence and fix something to make the whole thing consistent. While, of course, the reason that we need all this latent space for the video model is, like you said, we keep all the tokens in memory, we iterate over that full sequence, and you can adjust the past in order to make the future make sense. So if we think about the architecture that’s gonna get us there to these longer, richer sequences, it’s probably, like you said, gonna be a mix of the autoregressive and the diffusion, working together to do what each piece is good at.

Ali [01:21:10]: Well, if you get. Like you intuitively get why. So like English, for instance, or just writing in language, it’s like it’s just left to right. You can stream your tokens, you can stream your chain of thought. Just even as a human, you write like you just. You write and then you think about what’s the next thing you’re gonna generate, and then you write that, and then you think about your ideas, and then you generate forward. And sure, you can argue that as you write, you need to go back and you wanna edit some things, but you need to do that, less often than you’d think. Whereas with video, there is no sequential. The pixel in the top left corner of the video and the pixel in the bottom right corner of the video, they both need to attend to each other to understand how the video quality is gonna be almost as equally. Whereas with text, you don’t need that as much.

Philip [01:21:47]: Is there a parallel to audio? Like I’m not a hundred percent confident on this, but there was a point about a year ago where there was Audio LM, there’s diffusion for audio and autoregressive, and for the points you mentioned, mostly on the inference side, even though they’re shorter clips, most music is three to five minutes

Audio, Diffusion Text, and Cross-Modality Lessons

Ali [01:22:04]: Yeah

Philip [01:22:04]: We’ve swapped over to autoregressive Yeah, I can’t speak to music, but speech is autoregressive.

Ali [01:22:11]: Speech.

Philip [01:22:11]: You, effectively. This was even back with like the Orpheus architecture a year and a half ago. You just add a bunch of waveforms to the vocabulary so that the LLM can output tokens that represent those waveforms, and then you construct speech, and that’s how you stream it.

Ali [01:22:28]: That’s it. Wow.

Philip [01:22:29]: That’s my AIE talk from 2025.

Ali [01:22:32]: Nice. Nice. But it’s - with audio, it’s not the same challenge, though, is it? Because you. Like audio is solved with an LLM that generates everything. Like with audio, it’s still a transcript that you can generate with an LLM.

Philip [01:22:43]: Yeah.

Ali [01:22:43]: So your audio model just needs to like transcribe it, text to speech.

Philip [01:22:47]: For music, there was a phase of a trade-off between diffusion for music

Ali [01:22:52]: Right

Philip [01:22:52]: Autoregressive, and they were both pretty on par. There’s probably more pros and cons to either. I just wanted to poke and see if you had takes.

Ali [01:22:59]: Yeah, I don’t know about music specifically.

Philip [01:23:01]: Oh, well.

Ali [01:23:01]: What-- with what you said about editing you writing, I think my editor would tell me I need to do that more often and go back and fix things. I can imagine music or poetry, for example, where you have a rhyming scheme, and you might wanna go back and make a change to make it, to make it easier to set up a rhyme that you wanna make later on. There being some advantage to being able to attend in both directions. But yeah, to my knowledge, I very much bifurcate this inference problem into the autoregressive models, which have a set of constraints and techniques, and the diffusion models, which have a set of constraints and techniques. And, I think of text, embedding, voice in and voice out as being in the autoregressive side, and then image and video being in the diffusion side. There’s some overlap between the two. It’s not a perfect split, but that’s the broad categorization I use.

Swyx [01:24:02]: I should point out, I think it’s confirmed, right, Nano Banana and, GPT Image are autoregressive image.

Philip [01:24:07]: It’s this blended approach that we’re talking about, but in the image space, it hasn’t like made its way over to the video space, at least in the open source world.

Swyx [01:24:19]: Yeah. But like I assume that’s not too far away if that is possible

Philip [01:24:23]: Right.

Swyx [01:24:23]: On the. At least the Qwen Image guys are trying it.

Philip [01:24:26]: Yeah. Yeah. With

Swyx [01:24:27]: Yeah

Philip [01:24:28]: I’m really excited for Qwen Image 3. I hope they open source it.

Swyx [01:24:31]: And then I should also mention on the diffusion for tech side, there’s been some movement, not a lot.

Philip [01:24:37]: Yeah. We’ve got Mercury,

Swyx [01:24:39]: You host Mercury?

Philip [01:24:40]: Yeah.

Swyx [01:24:40]: Nice. Nice. Nice

Philip [01:24:41]: Diffusion Gemma is open source.

Swyx [01:24:44]: Yeah.

Philip [01:24:45]: And then, yeah

Swyx [01:24:47]: And we on the science pod, we just have been releasing, some, virtual cell models that use diffusion as well.

Philip [01:24:53]: Yeah. They have built. It’s definitely still in the cheap, fast tokens, world.

Swyx [01:25:01]: Yeah.

Philip [01:25:01]: We’re trying

Swyx [01:25:03]: It’- I think it’s the wrong marketing, and I’ve told them this before. I was like: “Look, like you’re not gonna beat the optimizations that, the other LLMs are gonna do, but you can have different APIs. Like you should be able to use it differently than chat response.”

Ali [01:25:19]: Me also.

Swyx [01:25:20]: Because it’s diffusion. Because you can do like. What is like context-free guidance for diffusion look like?

Swyx [01:25:26]: For text. Like give me a give me a poem, give me a plot structure that like diffuses into place

Philip [01:25:33]: Exactly. So that’s where, like I mentioned with poetry, for example, where you might want to ensure consistency across UIMs. I’ve done a lot of LLM sonnets. It used to be one of my to benchmarks, and even models today

Swyx [01:25:46]: They cannot count. Yeah

Philip [01:25:47]: Yeah, they don’t get the syllables right. And if you can attend across all of the different tokens, you can get the syllables right.

Swyx [01:25:55]: Yeah. And, David Holtz from Midjourney was, investing in text diffusion. I don’t think anything came out of it, but like the idea was that you can storyboard a long movie, and then you can generate the scenes with video- normal video gen. But the idea of like coherence across a thing that would just appear where like the end should attend to the start and you should not have this auto-regressive path dependency does make sense in principle. Just the API should be different. The marketing should be different.

Ali [01:26:24]: None of the most heavily used open source or closed source models use diffusion. But isn’t that like Like doesn’t that point to almost like

Swyx [01:26:31]: It is. It’s chicken and egg because what if you just give it more scale?

Ali [01:26:36]: What’s the, what’s the largest diffusion LLM?

Swyx [01:26:38]: I don’t think it’s very big.

Philip [01:26:40]: I don’t know the parameter count on this one, but diffusion Gemma

Swyx [01:26:42]: Like under 20B. I don’t know

Philip [01:26:43]: Diffusion Gemma is not large.

Vibhu [01:26:44]: I think it’s a 20-something.

Swyx [01:26:46]: Yeah. And yeah.

Ali [01:26:47]: Oh, it’

Swyx [01:26:47]: Like you haven’t tried.

Vibhu [01:26:49]: You haven’t given it a big and you haven’t,

Swyx [01:26:51]: So it’s like very unfair

Vibhu [01:26:51]: Diffusion Gemma is a 25B and it’s old

Philip [01:26:54]: And that’s what I’m saying is like for its size, it does pretty well, in terms of, in terms of quality.

Ali [01:27:01]: It’s almost like the same challenge with video models that have the same size. It’s like you’re comparing it to models that are much larger in scale.

Swyx [01:27:07]: Yeah. Well, unless you do the whole thing where you have a text, backbone and then

Ali [01:27:12]: Right. Right.

Swyx [01:27:12]: You like glom some decoder thing that, does that. Like, - so we started off the podcast doing this for the inverse direction from image to text.

Ali [01:27:22]: Right.

Swyx [01:27:23]: And I think like it’s, it’s roughly intuitive that you can do the opposite direction.

Ali [01:27:27]: I agree.

Ali [01:27:28]: I see it. I see it.

Swyx [01:27:29]: Yeah. The, we’re, we’re speculating on research in general.

Ali [01:27:32]: Yeah.

Swyx [01:27:32]: One part that we can end off with this is the topic of your talk where, inference engineering used to just be like, let’s take an open model, make the GPU go

Training for Inference and Inference for Training

Swyx [01:27:43]: And then that’s it. That’s the job of Baseten. Now it looks like people are using inference more and more in post-training.

Ali [01:27:50]: Yes.

Swyx [01:27:51]: Yeah.

Ali [01:27:51]: And training and inference.

Philip [01:27:53]: Yes. It’s training for inference and inference for training both have become big topics.

Ali [01:27:58]: Well, inference for training in the sense that like you just need, you need to do, you need to do rollouts when you’re doing like RL training runs. And so if your rollouts are taking a long time, if like, you’re using a vLLM for instance, or as opposed to vLLM or if the model that you’re trying to train is not supported in vLLM and you have to fall back to an older inference engine, your rollouts are gonna be slow and you don’t wanna do training on rollouts that are too off policy, so you have to wait for them so you bottleneck your entire training pipeline. And so like the techniques that we do inference optimizations for, will help them there. The training for inference mostly comes down to like just the spec dec training, EAGLE training, and sometimes post-training. For instance, if you want to quantize a model, you’ll quantize it down to like NVFP4.

Ali [01:28:43]: How do you like sometimes you get lucky and you can just do PTQ and that works. Sometimes you quantize it down to NVFP4 and the model is terrible, like the quality is too bad. And you have to do post-training on the model in order to make it understand that it’s going to now be an NVFP4 and let it still output the same logits. You can do this with normal SFT, PC, quantization aware training, all of that stuff. But more and more so we’re seeing techniques like NVIDIA released a quantization aware distillation paper where you establish a version of the model that’s in NVFP4 and a version of the model that’s in full precision, and then you’ll do distillation training based on the logits of the two models in order to make the FP4 model understand. And so more and more of the team, the engineers, like of the inference engineers that work on our team, they have to be very familiar with like training techniques and just being fine writing training pipelines for it. Yeah, it just seems like, they’re meshing together in a sense.

Swyx [01:29:36]: Well, it’s, coming together.

Philip [01:29:38]: Yeah, absolutely. If you think about the ultimate goal potentially of having a continuous improvement system . Yeah, it’s, it’s funny, but at the same time it’s also happening and I think within a few months to a couple years, like a lot of leading agent builders are going to have these loops like really up and running in production where you are doing inference, learning from the inference. We for a long time have been like learning from inference as it’s live and dynamically adjusting the system. Any dynamic adjustment is going to beat a static configuration across, your, exact config, across your speculator, across that thing. And then the, you can take the traces that you’re generating from your product, continuously post-train the model, roll those out, A/B test, get better signal, get better model, get better product. That loop is really promising. The technologies and the infrastructure to build it are coming along quickly. And so the unification between training and inference, I think, is only going to accelerate.

Swyx [01:31:01]: I was chuckling, but I wasn’- I didn’t think it was funny. Like it’s real. Like one of the big things for AIE World’s Fair was that, we have, RSI into AGI is the rough tagline. Which like, yeah, we have, I saw you pull a parameter golf. Like we have models training models and, the next step is models training, - or optimizing their own inference, which is funny. I wonder if, models will be like on policy better at training themselves than training models that they are unfamiliar with. This-- these are all like very interesting open areas of research.

Models Optimizing Their Own Inference

Philip [01:31:36]: One big part of my job a couple years ago was for any arbitrary model that came out on Hugging Face, writing a config foot and getting it up and running. And now the get-it-up-and-running config is shottable.

Philip [01:31:50]: And so, I don’t have to do that anymore. Yeah, that’s not exactly a model optimizing its own influence so much as a model, like being able to read the SGLang docs. But, yeah,

Ali [01:32:01]: Well, we do see it. We do see it like

Philip [01:32:03]: Yeah

Ali [01:32:03]: With GLM-5.2 for instance. GLM-5.2 is very good at writing GPU kernels. And so for like-- It was very funny internally, we had a GLM-5.2 endpoint that we were using to, like that we plugged in our cloud code harness, so every engineer on team uses like our GLM-5.2. And it will do a forward pass on the GLM-5.2 instance of the node, and then it will get the profile trace, and it will analyze it, and it will find the kernels that are the bottlenecks in SGLang, and then it will write the new kernels, and then we’ll do another profiling trace, and when it’s done, it uploads the image to our thing, and then we can pull that image down and repeat the cycle. And so for quite a bit of time, we had like literally GLM-5.2 optimizing

Philip [01:32:44]: Writing and optimizing all of GLM-5.2

Ali [01:32:46]: A GLM-5.2. And like some of the GPU kernels that were on GLM-5.2 within our inference engine is written by GLM-5.2, and the trace and the kernels were guided by GLM-5.2 as the driver. So it seems like. I do see, I do see that circle being there. I think a bit more time is needed. There’s definitely a lot of things that it can’t do. The models just aren’t there yet, even though they’re like really smart. Like, they still try to like reward hack their way into like the cheapest or like they’re very-- like they’re not good at like decision-making almost it seems. But yeah, I do. Like yeah, like a model optimizing its inference is already a thing that happens.

Philip [01:33:20]: Do you think GLM-5.2 was uniquely good at optimizing itself or did it just happen to be the best coding model that we had access to it would do an equally good job of optimizing,

Ali [01:33:31]: Would

Philip [01:33:31]: A DeepSeek or a Kimi or something?

Ali [01:33:34]: Well, to Swyx’s point, maybe it’s gonna be off policy when it tries to optimize

Philip [01:33:37]: Will it secretly hurt DeepSeek?

Ali [01:33:40]: To try to boost itself.

Philip [01:33:41]: Ooh.

Ali [01:33:42]: That’

Philip [01:33:42]: No, for what it’s worth, I don’t believe that.

Ali [01:33:44]: Yeah.

Philip [01:33:44]: But it’s just. Let’s just find out.

Ali [01:33:45]: It’s an interesting. Yeah.

Philip [01:33:47]: Just, you have more compute than me. Just

Ali [01:33:49]: Just go try it

Philip [01:33:50]: Try it. Yeah. Any other upcoming trends in inference engineering that we didn’t cover? Like right now, - ‘cause you guys are so close to

Future Trends: Modalities, Scale, Networking, and Continual Learning

Ali [01:33:58]: Yeah

Philip [01:33:58]: You can see it, that the world-- rest of the world doesn’t know about. The big ones are obvious. Models get bigger. Hardware gets more powerful. Users get used to a certain level of speed and demand a higher one. I think that some things I’m excited about are systems level. We still have a lot to think about in terms of composing multiple models together. If you think about a voice agent, there’s three to five models involved in that and the communication between those models. There’s a lot of new modalities that are coming out. There’s like the Cosmos, the new world model. There’s more research. Speech to speech is still like not entirely a thing, but it’s getting, it’s getting closer. There’s gonna be just a lot of new modalities to build around, which is gonna be exciting. And then, yeah, I think that the other thing to solve, which is something we’ve been solving for a long time and are not done with yet, is just going to be continuing to operate at another 10X scale as an industry. If you think about the degree of usage that AI has worldwide compared to, some of the more mature technologies both on consumer and business, it’s pretty clear that there could be multiple 10Xs more of demand. If you look at the infrastructure work industry-wide, it’s been stood up very quickly to meet a unprecedented spike in demand that is like not stopping. So yeah, there’s just a lot of problems to solve around like long tail reliability and, figuring out where we’re gonna get the next like 10X and 100X of tokens from.

Ali [01:35:49]: I’m gonna say, it’s gonna be a really boring answer, but I think the answer is just faster next, like faster network chip communications. It seems to me that like more and more memory is the bottleneck. You wanna have larger models. Right now, when you’re doing serving at large, you have to transfer KV cache from one node to another. But the way that you do that is you tran- you find the KV cache, you find where it is, you transfer it to another node, you put it on that node’s memory, and then you transfer it from that node’s memory into the GPU, and for like into the tensor cores of the GPU. So there’s like a stage transfer here that makes it such that you’re very bottlenecked with just KV cache transfers at large, which affects the time of decode and PD disagg. You have to do this because the HBM is so - it’s like extremely fast, like 4.5 terabytes per second as opposed to. Like, which is like magnitudes better than NIC communication speed. If you were to somehow be able to, in like this theoretical dreamland, have extremely fast NICs, you could, in theory, spare that HBM, and you could just transfer KV cache trans like directly from one node to another. This would give you like almost 100X speed up when you’re doing this aggregated serving between nodes and nodes. I’m not familiar with the technical challenges of making NICs faster. I’m certain there’s a reason why they’re like orders of magnitude

Ali [01:36:59]: Smaller, like slower than, like HBM. But if someone were to figure that out, it would literally be like a - like two orders of magnitude faster to do decode. That would be my take.

Philip [01:37:12]: Be a good trip.

Ali [01:37:12]: Cool.

Philip [01:37:13]: I don’t know if you have a nomination for things that are trends. I got one.

Ali [01:37:18]: Cool.

Philip [01:37:19]: So I think inference engineering for continual learning. So what if you just, like if you just had the idea that you are supposed to learn from everything that you ever process, do you do anything differently? Or do you just have the same paradigm of like, well, stick it in a memory.md, and then like it somehow gets consumed in KV cache, and like this system works, it’s not broken. Or like how do you like reshape inference so that it learns while you inference?

KV Cache Compaction and Continual Learning

Ali [01:37:48]: Yeah. I think maybe one relevant topic there is your absolute best fund in the entire world’s work on KV compaction Correctly

Swyx [01:37:55]: Like what changes?

Ali [01:37:56]: What changes when

Swyx [01:37:57]: If you’re trying to continual learn

Ali [01:37:58]: There’s two takes, and there was like Charlie and I had this Twitter, argument where the. Like continual learning could take one of two paths. It could either be that the model learns and so it’s continuously pushing its new knowledge into its weights. In that case, you just need to have, like your inference just needs to continually fetch new weights or yeah, like you just literally need to do fetch new writes and reads of weights. Or the other path, which is you do KV cache compaction. And if you

Swyx [01:38:28]: And there’s a LoRA layer if you just only update LoRAs.

Ali [01:38:31]: Yeah, exactly. Exactly.

Swyx [01:38:31]: Which is, that’s the gram approach

Ali [01:38:33]: Yes

Swyx [01:38:33]: Which we covered.

Ali [01:38:34]: The argument against doing weight pushing is that you can only fix one hop knowledge, as in you can only

Swyx [01:38:39]: Yeah

Ali [01:38:39]: Feed it a new feature of like, “Oh, what is the best university in the world?” The best university in the world is Waterloo. But then a second derivative

Swyx [01:38:46]: That’s not changing.

Ali [01:38:47]: That’s not changing. That’s not changing. But like a second derivative question of which university should I hire an intern from? So if that the best university in the world is Waterloo, then the answer should be Waterloo. But if I wasn’t just shotting the question and I was to ask it to like use its knowledge to think and then give me a second answer, or like, “Should I hire an intern from Waterloo or MIT?” It’d be like, “Oh yeah, both are good.” But no, like I liter- I just edited in your knowledge base that Waterloo is the best. Why didn’t you use that to do reasoning? So that’s the fundamental problem with trying to change a fact in an MLP within the weight. KV cache compaction fixes that. With KV cache, or like rather not KV cache compaction, but like if you’re able to have something like the still paper which we came out with, which is you’re able to make your KV almost infinite, and you’re able to compact in such a way that you don’t lose any of the knowledge. In that case, you can do continual learning, and you can solve continual learning. And this as a, it’s a result of, this argument that Charlie and I had, that I do concede that his point was correct, and I do see that KV cache is the way forward. And in that case, I don’t think inference is going to change that much because we still use KV cache and inference. You’re just gonna update the KV cache, but it’s gonna be like an additional step, but nothing changes in the weight, so nothing changes in inference time. Nothing changes the spec that I had.

Swyx [01:39:58]: Okay. Surprisingly great answer. We have it up on the blog. It’s a relatively recent blog, so, we can. People can go see it.

Closing: The Book, Baseten, and Inference Engineering

Ali [01:40:06]: Hyperverve

Swyx [01:40:07]: Yeah. Otherwise, this is super enjoyable chat. I know we’ve like already gone two hours.

Philip [01:40:11]: Wow. I didn’t even realize.

Swyx [01:40:12]: Like time flies. Yeah.

Philip [01:40:13]: Yeah. So much we didn’t even cover.

Swyx [01:40:15]: Yeah. This is like, we also wanted to talk about the book and all that, but you’ve covered the book.

Philip [01:40:18]: Yeah, everyone knows about the book.

Ali [01:40:22]: Yeah.

Swyx [01:40:22]: High- highest ROI thing in the history of Baseten, right? For the hour.

Ali [01:40:27]: Without a doubt. Without a doubt.

Philip [01:40:28]: Yeah.

Ali [01:40:28]: Absolutely.

Swyx [01:40:29]: So congrats on that. I, and we’ve covered that in our meetup

Ali [01:40:32]: Yeah

Swyx [01:40:32]: Which we can publish separately. But no, thank you to you guys for being so generous for sharing. I think it’s a fun conversation that, we don’t get to have enough. I think inference engineering, we never really covered head on, and so to have you guys come on, is a treat.

Philip [01:40:47]: Always.

Ali [01:40:47]: It was amazing.

Philip [01:40:48]: Yeah. Thanks. Thanks for having us, and hopefully in a year everything shifts, and we can, come back and say everything we were wrong about.

Swyx [01:40:56]: Yeah. Yeah. I’m excited for this mega kernels comment to get out and see what’ see what people say.

Philip [01:41:00]: We gotta stir stuff.

Ali [01:41:02]: Should I go into hiding? I know I’m gonna get like the mega kernel community after me.

Philip [01:41:05]: Yeah. One thing I really respect about you is you are not willing. You are not, scared to kick the hornet’s nest, ever.

Swyx [01:41:12]: It’s not, I don’t think it’s that controversial. I don’t know. We’ll see.

Ali [01:41:18]: We’ll see. We’ll see.

Swyx [01:41:19]: All right. Thanks, guys.

Philip [01:41:21]: Thanks.

Ali [01:41:21]: No, thank you so much.



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, BasetenLatent Space: The AI Engineer Podcast · 1 h 41 min
Listen in VO