973: AI Systems Performance Engineering, with Chris Fregly

10 Mar 2026 · 1 h 12 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI systems performance engineering—how to co-optimize hardware, OS/driver stack, CUDA/PyTorch, and algorithms for efficient GPU training and especially inference (including KV cache, prefill/decode, multi-node/disaggregated inference).

Guest

Chris Fregly, principal AI systems performance specialist; worked at AWS, Databricks, and Netflix (Emmy-winning work at Netflix); now O’Reilly author of AI Systems Performance Engineering (1,060 pages) and two prior AWS-focused books.

Key claims

NVIDIA documentation is “a total disaster” and forums are often unhelpful; performance gains come from “mechanical sympathy” (hardware/software/algorithm co-design) and from profiling beyond PyTorch profiler into GPU low-level counters (occupancy, streaming multiprocessors, separate pipelines for SFUs and tensor cores, memory bandwidth). Examples: DeepSeek R1 reportedly trained for <$6M vs OpenAI models costing hundreds of millions; attributed to undocumented/under-documented NVIDIA tricks, algorithm/software/hardware co-design, custom storage layers, and treating communication bandwidth as scarce. Notable examples/tools: 175-item optimization checklist; GitHub repo with ~700–800 examples; Zymtrace for runtime profiling; advice to benchmark on the target machine (e.g., Blackwell/B200), lock GPU clocks, and avoid network file systems.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Journey to Writing the Book

0:39 to 2:00

Chris discusses his writing process and the inspiration behind his book.

“This episode of Super Data Science is made possible by Anthropic, Cisco, Excel Data, and the Open Data Science Conference.”

The Cost of Writing: A Starbucks Story

2:00 to 2:20

Chris shares how he spent thousands at Starbucks while writing his book.

The Challenges of Information Access

2:20 to 4:20

Chris explains the difficulties he faced in accessing quality information for his book.

“understood every aspect of the NVIDIA stack.”

Performance Engineering Insights

4:20 to 6:40

A deep dive into performance engineering choices that lead to cost reductions in AI model training.

“Blackwell was more a spec and had not even been, you know, taped out yet, I don't think, by NVIDIA.”

Communication Bandwidth vs. Compute Power

6:40 to 9:00

Discussion on the significance of communication bandwidth in AI systems.

“And what's really interesting, of course, is that due to the Chinese export restrictions, US export to China, China, in theory, cannot get a hold of the best chips.”

Understanding the Hardware-Software Relationship

9:00 to 11:30

Chris emphasizes the importance of understanding both hardware and software in AI engineering.

“that storage layer being extremely critical because it is a bottleneck.”

The Future of AI Systems Performance

11:30 to 14:00

Thoughts on the evolving landscape of AI hardware and software performance.

“I feel like to probably make the most of performance engineering for AI systems, you probably would.”

Co-Design of Hardware and Software

14:00 to 15:09

Learn about the relationship between hardware improvements and software algorithms in AI systems.

“Think of this hardware software co-design.”

Horizontal Scaling of Intelligence

15:09 to 17:44

Discover the concept of agents sharing knowledge across a network and its implications for AI.

“Agents sharing knowledge across a network, coordinating on common intent, reasoning together.”

Mechanical Sympathy in AI Systems

17:44 to 19:15

Understand the importance of co-optimization in hardware, software, and algorithms.

“And yeah, so I don't know if there's anything more to say about that term.”
Show all 24 chapters

Attention to GPU Metrics

19:15 to 21:01

Learn why monitoring GPU metrics is crucial for optimizing AI performance.

“And this is more than just temperature, which can certainly affect things more so than CPUs.”

Deep Dive into PyTorch Profiler

21:01 to 22:49

Explore the limitations of the PyTorch profiler and what additional metrics are crucial.

“PyTorch is by far the most popular library today for putting the pieces together of an AI system.”

Understanding GPU Architecture

22:49 to 25:18

Gain insights into how GPU architecture impacts AI performance and the significance of specialized function units.

“Yeah, the GPU does expose a lot of counters, and there's really the two or three main tools are at the kernel level.”

Practical AI System Optimization

25:18 to 28:00

Learn about practical steps and tools for optimizing AI systems effectively.

“these things and make sure that they're all fully utilized all at once.”

AI System Optimization Challenges

28:00 to 29:19

Discussion on the difficulties of optimizing AI systems and using a book checklist.

“because it basically got too hot and then stopped processing as quickly.”

Key Optimization Tips for AI Systems

29:20 to 38:07

Tips for optimizing AI systems, discussing GPU configurations, software versions, and memory management.

“of the book, which obviously I do, I give that to Cursor.”

Understanding Inference in AI

38:08 to 42:00

Overview of inference in AI, including multi-node inference and caching mechanisms.

“And yeah, of course, check out Chris's book, AI Systems Performance Engineering, to get the full list of 175, the full thousand pages on how you could be optimizing your systems.”

Understanding Disaggregated PD and KVCache

42:00 to 43:06

Learn about disaggregated pre-filled decode and its significance in AI engineering.

“based on how big the GPU RAM is, sometimes on disk.”

The Impact of AI Coding Assistants on Workflow

43:06 to 45:58

Discover how AI coding assistants are revolutionizing programming workflows and performance optimization.

“There are your interview tips of the episode.”

Shifting Coding Practices with AI Assistance

45:58 to 50:53

Explore the changing practices in code review and testing with AI code generation.

“until people realize that, you know, Codex is actually better for other stuff.”

Investing in AI Hardware and Startups

50:53 to 55:48

Gain insights into the future of AI hardware and investment strategies in the sector.

“I guess I need to stop looking at the code.”

Investing in AI Technologies

56:00 to 1:02:21

Learn about investment strategies and market opportunities in AI tech.

“And so it's a really great complement to the Nvidia chip, which suffers from this, you know, memory bandwidth limitation.”

Book Recommendation: Founding Sales

1:02:21 to 1:04:24

Discover why 'Founding Sales' is crucial for early-stage founders.

“And these days, it's a relatively old book.”

Chris Fregly's Expertise and Social Media

1:04:24 to 1:06:39

Find out how to follow Chris Fregly and learn more about his work.

“And you'll read something and say, oh, that's exactly what I was thinking.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:My guest today spent$6 ,000 at Starbucks writing a 1 ,000-page book that NVIDIA's own documentation couldn't provide, and what he found under the hood of GPU computing will update everything you know about engineering AI systems. Welcome to episode number 973 of the Super Data Science Podcast. I'm your host, Jon Krohn. Today's guest, Chris Fregly, has worked as a principal engineer and AI systems performance specialist at AWS, Databricks, and Netflix, and now he's recorded his vast knowledge into his third O 'Reilly book, a massive, invaluable tome called AI Systems Performance Engineering. In today's episode, he distills the book's thousand pages down to its most valuable takeaways.

0:39Jon Krohn:Enjoy. This episode of Super Data Science is made possible by Anthropic, Cisco, Excel Data, and the Open Data Science Conference. Chris Fregley, welcome to the Super Data Science Podcast. I've been trying to get a conversation with you for like a decade, and it's finally now happened. Yeah, welcome to the show. Yeah, man. Thanks for having, uh, yeah. I, I heard you had one of my friends, Ancha on a recent podcast and she, yeah, she said the same thing that you were trying. I was like, let's do it. So the second I found out, I sent you an email, but Ancha was an episode number 963. She's absolutely brilliant.

1:17Jon Krohn:As you know, you've coauthored books with her in the past. You're a three-time O 'Reilly book author. And in this episode, we're going to talk mostly about your latest book. It's called AI Systems Performance Engineering. And then it's got a really long subtitle, Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch. And it's not just the subtitle that's long. This book is a thousand pages. It's actually, it's including the indices and stuff. It's a thousand and sixty pages. And there's photos of you. If people go to your LinkedIn profile and they see your banner at the top, there's photos of you with the book.

1:58Jon Krohn:And I love that you're holding the book sideways so that people can see what a thousand page book looks like. It's a banger. Yeah. Thousand pages. It's funny. I went back through all my Starbucks bills over the last year because I spent, yeah, I basically, when I left Amazon, I just started working on this book to make sure I understood every aspect of the NVIDIA stack. And yeah, I'm sure we'll get to it. But I went to Starbucks pretty much every day for a year and my bill was over like$5 ,000 or$6 ,000, I think. Yeah, just on lattes alone. And you go to Starbucks to read or to write as well? To write.

2:43Jon Krohn:To write, eh? You write in the coffee shop. Yeah. I used to write at bars. Well, I still do write at bars. I write anywhere I can. And I'm not good at sitting in a quiet room. I'm good at chaos around me. And then, you know, yeah. So Ancha, by the way, is not that way. And so she and I have very, very different writing styles, very different work styles. Yeah. She, yeah, she prefers. I've never actually seen him there, but for years I lived in the West Village of Manhattan and supposedly Malcolm Gladwell would frequent this coffee shop that was like next door to my apartment. And for the same reason, he would kind of, Malcolm Gladwell would come in for a two two or three hours a day and just write.

3:24Jon Krohn:So it works for some people. We're one of the best selling authors of our time. So cool strategy. And yeah, is it all of that coffee that resulted in it being a 1000 page book? How does that, you know, I guess, first of all, tell us why you felt you needed to write this book, AI assistance, performance engineering. And did you expect it was going to be a thousand pages when you started? How did that? Absolutely not. No, So the book started off, I had proposed it. So it's my third time. So proposals can be a little bit looser with O 'Reilly because they've already vetted me and I've already got book sales behind me and they've seen the quality of my work.

4:04So really, it was scoping out what I thought was going to be mostly PyTorch, maybe one chapter on CUDA and then the rest on the hardware, you know, GPU hardware, Blackwell. And, you know, keep in mind when I started writing this, Blackwell was more a spec and had not even been, you know, taped out yet, I don't think, by NVIDIA. So I was just going off a lot of information. And then, no, it started off, the proposal was 300 pages. And the purest motivation was when I would be working with customers back at Amazon, I was more focused on the NVIDIA stack back at Amazon than I was training. I had a little training experience, but mostly NVIDIA customers.

4:56And I was having a hard time getting information. And I was going through NVIDIA forums. I was getting, which are not that good, by the way. There's a lot of unanswered threads on the NVIDIA site. The NVIDIA documentation is a total disaster. trying to find things from Twitter, from X. Like I was getting more information from X and actual practitioners. And so that was actually the signal for the first two books as well. The first one was really about data science and specifically SageMaker and that whole ecosystem, which there's a million products within AWS. There's Glue and there's EMR and there's Redshift and all that.

5:39And so the inspiration has always been, it's really hard for me to find the information. And I know there's a lot of other people that are trying to find this information. So let's get this into a book form.

5:54Jon Krohn:Yeah, it's been hailed as the missing manual for the industry, a comprehensive encyclopedia that teaches you how to co-optimize hardware, software, and algorithms together for the world's most powerful AI models. Let's start off with chapter one. So chapter one of the book highlights how the brute force approach to AI is now being challenged by high efficiency engineering. So you talk about, for example, DeepSeek R1, which is now it's been about a year since it was released, and it allegedly cost less than six million dollars to train. While comparable models, you know, the O1, the O3 out of OpenAI, for example, supposedly cost hundreds of millions of dollars.

6:35Jon Krohn:And so what were the kinds of performance engineering choices that made that 10 to 20x cost reduction possible? Yeah. And what's really interesting, of course, is that due to the Chinese export restrictions, US export to China, China, in theory, cannot get a hold of the best chips. So they had to find clever solutions. You know, one that stands out is they found essentially undocumented or an underdocumented. I would say all of NVIDIA stuff is really underdocumented. So, but they found one in particular way to bypass a, you know, particular cache or short circuit, I would say, just, you know, keeping it at a high level.

7:25which led to huge, huge performance gains in this sort of limited capacity chip. And it's something, someone asked the OpenAI folks if they have discovered the same thing. And Sam Altman and company said that they're positive that DeepSeek, or that they, being OpenAI has found all of the same tricks that DeepSeq has found without getting into details. So these things aren't necessarily that sort of revolutionary. It's just no one talks about them. The labs have them as their secret sauce. And it just so happens that DeepSeq was open source and they talked about this. A lot of other innovations though, I mean, DeepSeq had what was called the open source week that I mentioned in the first chapter.

8:14And really, it was not just sort of CUDA type innovations, hardware innovations, but it was co-designing, which is the term that is starting to be used quite a bit. They figured out ways that they can modify algorithms, they can modify the actual software, and they can tune it specifically for their set of hardware. Also included in that hardware is the storage system. So they actually wrote their own storage layer, which is very interesting. This is something Jensen has just started talking about more and more. I think it was last year's GTC, where it's the first time that you really hear Jensen and the NVIDIA folks talking about the storage layer, that storage layer being extremely critical because it is a bottleneck.

9:07Basically picturing all the different bottlenecks, memory access. And compute is basically on a single machine or a single cluster, we've got a lot of compute. It's just all the other bottlenecks, which is like retrieving data from memory from storage.

9:27Jon Krohn:Yep. Right. Yeah. So Jensen, of course, Jensen Huang, the CEO of NVIDIA, probably all of our listeners already know that, but filling in one tiny little blank there. And going a little bit more into that deep seek, Seek paper and DeepSeek R1 release, they treated communication bandwidth. You were implying it there, but communication bandwidth as the scarce resource. So for a developer today, an AI engineer today who feels that they're compute poor, but maybe they actually aren't, maybe they're bandwidth poor, how can they fix that? How can they optimize that on their GPU instead of just buying more GPU hours.

10:03Yeah. The most critical metric or characteristic of a GPU system these days is the memory bandwidth, right? So a lot of times people focus on flops and teraflops and tensor cores and quantization, you know, without getting into too much detail there, you really want to look at from generation to generation, pay attention to the memory bandwidth. How fast can I get the weights of these 500 billion parameter models, one trillion parameter models? I need to load those into the GPU, into those registers, into those caches. And that's 100 % dictated by the memory bandwidth. with. And so one thing I always encourage teams that I work with, customers that I advise, companies I advise is nail that layer down.

11:10You have to do tricks, you have to overlap and keep the compute busy while you're loading data in. I mean, these are relatively well known concepts these days. It's just taking it from concept down to the hardware level. There's a lot of layers in between that are not well documented. And that's what my book focuses on.

11:33Jon Krohn:Really cool. I feel like to probably make the most of performance engineering for AI systems, you probably would. I mean, it's going to be tough in a podcast episode like this in an hour to convey to people all the things that they need to know or even the key things that they need to know? Because I imagine a lot of it is specific to the circumstances that they're in. Would you say that's correct? Right. At the end of the day, everything is workload dependent. You're typically fixed in the hardware that you can get, the budget that you have, right? This was in particular DeepSeek, didn't even have the same GPUs that the Western labs have.

12:18Um, you know, uh, um, and so trying to, and I was just looking at the book here. Yeah. After like you write a thousand pages, you can't even like really remember where anything is. So the secret to this book, by the way, real quick is that this is actually three different books and multiple times I had. So, yeah, I said there was only going to be one or two chapters on CUDA. Now there's, I think eight chapters on CUDA. And the reason is I couldn't find any single resource. Like there's a lot of CUDA books out there, but they're written by people for specific use cases like fluid dynamics and things that just aren't really relevant to the LLM world.

13:02There's certainly overlap in the underlying concepts, but being able to tie, here's a model, here's a DeepSeq model, or here's GPT OSS, by, you know, which is the open source version out of OpenAI. You know, how do I map that in a mechanical sympathetic manner? So this is one term that I've loved for years and years since my old Java day is trying to optimize JVMs and things is this concept of mechanical sympathy. And this is also ties into a piece of advice that I always give to my friends and specifically and specifically my friend's children is if you're in a CS program right now and thinking that your skills are going to not going to be relevant, really dive deep into the hardware and try to understand your surroundings.

13:56Now, the hardware is going to change, of course, by the time they graduate and stuff, but at least like don't stop at the software layer. Think of this hardware software co-design. The algorithms haven't changed that much. There's There's been tweaks over the last few years, different layer normalizers and, yeah, like different parts of the transformer. But what is sure to change is the hardware is going to get faster, the memory bandwidth is going to change, and you really need to be able to think in terms of that whole stack. And that's why this book exploded into a thousand pages was because every layer that I would uncover, I'm like, oh man, this is going to be way better if I can explain how all this works together.

15:08Jon Krohn:horizontally. Agents sharing knowledge across a network, coordinating on common intent, reasoning together. The infrastructure for that second horizontal axis doesn't exist yet. Outshift by Cisco is formalizing it. They call it the internet of cognition. They're publishing the architecture and building reference implementations. Read Scaling Out Super Intelligence. We've got a link to that in the show notes. Then check out episode number 961. In it, Dr. Vijoy Pandey, the head of Outshift by Cisco, walks through how horizontal scaling of intelligence works and why it matters. Right. So one of the sub books in your thousand page book, AI Systems Performance Engineering, is CUDA.

15:48Jon Krohn:What are the other two? Right. So think of it like, so there's, I think, three or four chapters on PyTorch. So think of it like the three parts of the co-design are the hardware, the software, and then the algorithms. And so really, the first few chapters kind of set up the hardware, set up the operating system, because all these things matter. One small change at the operating system can make or break and cost you millions of dollars. Right. And so I've got a whole chapter dedicated to that. And so, yeah, you know, basically all of those different layers. So hardware, software on the software side, there's there's basically a whole book that's in here.

16:45the last, I think, six or seven chapters that are on inference. I do cover training. However, I decided to really, really dive deeper into inference because of where everything's at with the current. I'm trying to keep this book relevant for the next couple of years. And so it's really, it's inference hardware.

17:08Jon Krohn:And inference is where, you can correct me if I'm wrong on this, But I believe about when you think about how much the workloads that a given GPU is handling in its lifetime, 99 % of that is going to be handling inference, not training. Absolutely. Yeah. Even for, you know, post-training, for the reinforcement learning, a huge portion of that is actually doing the, yeah, like the inference part of it. And then just a little bit of weight adjustments. And absolutely. Yeah, and so this co-optimization that you've been talking about of hardware, software, and algorithms, those three pieces together, it's a concept that you refer to in your book as mechanical sympathy.

17:49Jon Krohn:And yeah, so I don't know if there's anything more to say about that term. Yeah, it's a term. Martin Thompson, a famous computer scientist, software engineer, many years ago, described it. He was making this analogy to Jackie Stewart, who is a British Formula One champion. in. And Jackie Stewart was famous for not just being a driver, but also able to construct and, you know, build the cars. And that was really something very unique to, you know, things back then. And yeah, I would say, as even today, I don't think a lot of drivers are, you know, really capable of like, turning the wrench, right.

18:38And, you know, making these adjustments, they can give the feedback, which is what the constructors need to make changes throughout the race and things, but they're not necessarily the ones that are able. And it's a huge advantage when you're able. So a really big sort of emphasis, and this comes through with the GitHub repo for this book, by the way, which has about 700 examples and I think is about to reach maybe 800 by this weekend. I'm constantly updating the GitHub repo, which is documented in the book. But something I very much emphasize that I don't see teams doing is paying attention to the GPU metrics.

19:22And this is more than just temperature, which can certainly affect things more so than CPUs. but look at the low-level counters. It's something, NVIDIA has tools, but they're very cryptic. They're command line tools. They've got a UI, but it's janky and it's outdated. And I know that they're working on something better. But actually, one of the startups that I advise called Zymtrace, they've sort of cracked this. And they're sort of like, if you're familiar with AppDynamics, they are sort of the, you know, runtime, the kind of modern version of like AppDynamics, where it can look at the CPU, the GPU, it can look at PyTorch, give you exact stack traces.

20:06This is exactly what I was about to build, by the way, after, you know, writing this book was like realizing, you know, because part of me writing a book is figuring out like, what are the boundaries of this space? And can I find people that can fill in these things? And So it's a very untapped market and being able to, in real time, make adjustments, turn the wrench. And so profiling down to that level is extremely important. And not just stopping at the PyTorch profiler, which is what all the PyTorch community. it's so funny because it's just one or two more commands away, but the PyTorch folks seem to just stop at the PyTorch profiler.

20:52But that is not enough to really, really dive in and make these changes.

20:56Jon Krohn:To give our listeners who aren't aware, who maybe aren't hands-on practitioners of AI, PyTorch is by far the most popular library today for putting the pieces together of an AI system. And what is it, Chris, what's the PyTorch profiler? What do you see in that and why isn't it enough? Yeah, PyTorch profiler, you know, the reason the PyTorch community likes it is because it's Python friendly. It's using terms that are familiar to, you know, PyTorch users. It's showing you where time is being spent in your training job or in your, you know, inference or basically any PyTorch job that you're running, code base that you're running.

21:41But it really only gives you a few different aspects of, you know, it's the tip of the iceberg, you know, when it comes to trying to actually really understand the system and really achieve mechanical, you know, understanding of the whole system. So, and it'll show you sort of high level that the CPU is being used here and that the GPU now comes in and there could be better overlap here where you could be moving data at the same time. But there's like, you know, probably 50 to 60 more critical or just as critical metrics that are needed to really move things forward. And so that's why I haven't really seen a whole lot of like progress in terms of performance or that type of progress is really limited to the frontier labs or folks like DeepSeek that are taking the time to really understand what's beneath the covers.

22:38Jon Krohn:What are the kind of, you mentioned there, kind of a dozen or so metrics that are critical, that are difficult to see? What are kind of like the key ones that we might not be aware of and how could we possibly see them? Sure. Yeah, the GPU does expose a lot of counters, and there's really the two or three main tools are at the kernel level. So kernel is like what it's the program that the GPU actually runs, right? And when you program to a GPU, you have to realize that part of it's happening on the CPU to kind of prepare the data, then the data has to ship over to the GPU. And now the GPU has a whole bunch of constructs and hardware optimizations that have been built up for the last 10 years specific to AI.

23:31So there's things like the streaming multiprocessor, right? Think of it like a core or a set of cores that you would tune for parallelism on the CPU. Well, the GPU has that same thing called streaming multiprocessors. There's the amount of time being spent on those streaming multiprocessors. There's a percentage that's called the occupancy. And it's basically all about the utilization of the hardware. But the other really key critical thing is that there's separate pipelines. there's instruction pipelines that power all of these different portions of the GPU. So the GPU has specialized chips for soft Macs, which is used super heavily in like transformer world.

24:27And so what we've seen is since transformers, you know, 2016, 2017, once NVIDIA like found out that this is what researchers are using, they started building in these specialized function units there, which short, so that acronym is then SFU for specialized function unit to handle those transformer specific operations. And then of course, there's also tensor cores, which are specialized cores that can do floating 0.8, floating 0.16. So like specialized operations that of course are, common in transformers. Those all have separate pipelines. Trying to move memory in and out, that's another pipeline.

25:17And so you really have to understand how can I best parallelize all of these things and make sure that they're all fully utilized all at once.

25:27Jon Krohn:Really cool. I don't have any experience with that level of depth. And in fact, that leads me to my next question because most data scientists, AI engineers, we stay at the Python level. At what point does an AI system scale make it mandatory to go down the stack into CUDA, C++, Triton, or I don't know, SFU? For sure. Yeah. And one of the things I've been really trying to do with my GitHub repo is like I've got MCP tools that you can, you know, if if you're inside of Cursor, inside of Cloud Code, you can use these MCP, right? But the issue is that you have to be directly on the Blockwell B200 machine to run these tools.

26:19And so you have to shell in. So like my particular development flow is, I shell right in with Cursor or with like VS Code, with the different plugins, Codex and Cloud Code, I shell in directly to the B200 that I'm actually going to run this on. And I start chatting basically with the system and say, okay, tell me about what's your memory bandwidth, run this particular PyTorch code or this inference job and tell me where is time being spent. And then tell me which lines of code. And so I'm able to actually chat with my system. And so that's super powerful. And it can also detect things like when the GPU is getting warm.

27:08And so I could lock the GPU clocks, for example, to get a consistent benchmark. One thing about the GPU sort of ecosystem is, it's like these things are purposely running way hotter than a CPU would. If you recall, maybe being a kid, you were building your own machine or you have friends that were building machine and they would always try to overclock, you know, either the CPU or the GPU. Yeah, right. So the NVIDIA folks have decided just to like, yeah, just overclock it themselves. But, you know, if you're trying to run a benchmark or trying to see if a change that you made had an impact, you have to basically cap that overclocking because if you go up too hot, it's going to throttle and then your change will appear to not have impact because it basically got too hot and then stopped processing as quickly.

28:08So, yeah. It's still very hard even to go from my book, which has all of this information in one place, to like a day-to-day practical thing, which is where the GitHub repo, which is where these different startups are really focusing on.

Read the full transcript

28:27Jon Krohn:Right. And so earlier in this episode, I said how it would be very difficult in 60 minutes in an hour-long podcast episode to convey all the things that somebody needs to know in order to optimize their AI system. Your book actually ends with a 175-item optimization checklist, which kind of goes to prove my point that if we took a 60-minute episode, it would be able to spend 20 seconds on each of those items in the checklist, which obviously is not going to convey any meaningful information. But if a listener had to take a handful of those, maybe like three of those as the kinds of optimizations that would give them the best return on their time, investing their time in that, is it easy to kind of say that these handful are the kinds of key things that they need to be doing?

29:19I'll give you a little tip, which is if you have the PDF of the book, which obviously I do, I give that to Cursor. I give that to CloudCode. And when I'm analyzing a system, right? Because these models don't have all of these things at their fingertips. There's a lack of training data, for example. So one thing I do is I actually put the book in context and it's a thousand pages. So you need a big model that has a big context window and ask it, what are the top 10 things I should be looking at? And it'll start to run its checks and everything, assuming that you have the right tools installed, which is the NVIDIA Insight, and the NVIDIA Compute.

30:18So that's kind of longer term what you would do, which is try to give enough context. Now, I think this will change because I know that the labs, I know firsthand that the labs are training more and more on hardware specific for AMD, because they themselves have this interest in their models getting better at like optimizing kernels. But to be very systematic about these things, I'm certainly going to talk through a few of these, but make sure there's some sort of like before and after profiling best practices, things like locking the GPU clocks and making sure it doesn't hit peak. But then there's really having like a methodical way of going through.

31:11So in my GitHub repo, for example, I've got this harness that will do things consistently before and then after. And if there's any inconsistencies, it'll fail the comparison and say, hey, this is not apples to apples. So it's extremely important that you have a very, very methodical way of comparing changes. But yeah, things to keep an eye on, Well, yeah, I'll start small and then kind of go bigger. On a single GPU, obviously you really want to understand if you're on Blackwell, has extremely different characteristics than Hopper, then we'll have in the future Vera, or I think it's Ruben is the next generation, and then Feynman is the next generation after that.

32:03But separate from the GPU, there's a lot of things that can go wrong at the actual operating system level. And with the GPU driver, for example, there's the famous CUDA toolkit and then the NVIDIA driver. And yes, anyone that's tried to enter this world has spent, I'm sure, countless hours trying to get these things synced up and the versions all set up properly. Once you get past that, and I do want to emphasize that you should always be on the latest. You know, I don't always upgrade for the sake of like upgrading in my life, just sort of generally. But I've noticed in the GPU world, it's extremely critical to stay up to date.

32:48That does cause problems because sometimes the, you know, PyTorch doesn't keep up with the latest CUDA 13.1, you know, or, but you really, really have to be aware of the versions in this world. But yeah, I would keep an eye on memory fragmentation, extremely important. You want to disable swap, for example. These are things that if you just get a basic Linux system, these things are all turned on because they're optimized for CPUs and just sort of general purpose workloads and a lot of different tuning that goes in. This was actually one of my favorite things about the SageMaker

33:32sort of ecosystem is that there was a lot of work put in. Yeah, so SageMaker is the Amazon, or it's a service within AWS that offers you GPUs. And specifically, there's a service called SageMaker HyperPod, which is pretty bare bones in a good way. It's got Kubernetes installed. It's got all of these things tuned. You've got a lot of people focusing on this. There's a lot of customers that are tuning, that are helping Amazon with the right configuration for all of these things. You almost always want to occupy the entire GPU. So you don't want to share like you would in a CPU land where you can split it up and virtualize it.

34:19You want to get as close to bare metal as possible. Containers have gotten a lot better throughout the years with GPUs. So we're seeing a lot more folks use Kubernetes and Docker for these kinds of things. But I would almost encourage to stick with bare metal. IO optimization, I think is the number one killer here. This is the lowest hanging fruit. Use the fastest disks that you can get. Use local disk. Don't try to use any network file system. It's convenient to use a network file system, but it will almost always cause you problems. I had said earlier that Jensen, the CEO of NVIDIA, has been focusing a lot on storage recently.

35:06Probably not as much as I think he should be going forward, but he does have a lot of partners that are working with, and there's something called GPU direct storage, which can load data from disk right into the GPU without going through CPUs, without creating copy buffers, and those kinds of things. So that would impact the sort of data pipeline. And when you scale, of course, to multiple nodes, multiple GPUs, which is these days just a given, you really can't do anything that complex on a single GPU anymore. So then you start to optimize the network and then you have to make sure that you're using direct GPU to GPU communication and use things like NVLink, right?

35:55The NVIDIA interconnect that's been optimized for the... So one piece of advice is when you're looking around for a GPU cloud provider, I guess these days people call them NeoClouds, like CoreWeave and stuff, try to find one that has bought into the entire reference architecture, you know, like by NVIDIA. If you, you know, there's a group of, I guess, cloud incumbents that have their own networking, that have their own stuff that's been built up, you know, throughout the years. And they are basically patching and bypassing the sort of natural way that the NVIDIA GPUs were designed. And that's actually, in my opinion, the highest value proposition for these Neo clouds is they're not tied to all of those sort of CPU legacy ways of doing things.

36:55They have bought into the NVIDIA ecosystem. In some senses, they're favored by NVIDIA, I would posit, which seems natural, right? because you're going to favor customers that buy a lot more from you. Because keep in mind, one of the most pivotal aspects or the most pivotal acquisition by NVIDIA was this company Mellanox. So they bought this company Mellanox back in, I think, 2019. I think it closed in maybe 2020. They're purely networking. And so now like NVIDIA since 2019 has gone from just a GPU company to an AI systems company. And so if you buy into that whole thing, there's a lot of efficiencies that happen, even in the network hardware can actually do calculations with these like Mellanox routers that, yeah, that like NVIDIA purchase.

37:53So, yep.

37:54Jon Krohn:Fantastic. Thank you for all of those tips, kind of giving us, distilling for us the most valuable of the 175 items provided in your book. I think that was like four of them. Yeah. And that's great. I mean, what a nice thing to be able to do in this episode. And yeah, of course, check out Chris's book, AI Systems Performance Engineering, to get the full list of 175, the full thousand pages on how you could be optimizing your systems. um let's let's shift gears now a little bit chris to talking about inference specifically uh you know we talked earlier in this episode about how inference is most of what's happening at scales with ai models training is an important part but inference is most of what we're doing with gpus these days and so you talk in the book about multi-node inference and disaggregated architectures.

38:46Jon Krohn:What the heck are those? Yeah. Welcome to another chapter that was supposed to be a single chapter or a topic that was supposed to be a single chapter. Yeah. So this is one of, this is one of the three books, one of the three sub books. Exactly. Inference, there's a lot going on in inference that folks probably on the surface don't realize. So stated like simply with inference, we're doing just a forward pass, where we've got the weights, the weights, which are the parameters in this, let's say it's a trillion parameter model. We just need to compute the answer, boom, just go forward. So on paper, this seems extremely simple.

39:29In reality, transformer-based inference, there's a lot of places for things to slow down. And so there's a lot of caching that needs to happen so that we're not recomputing things on every, you know. So think of when you go to ChatGPT or to Cloud, you type in a query. It's like, it is going to look at all of the context that you've passed in. It's going to compare everything to like everything else that's in there. it's going to find things that are relevant. And it would have to do that the same way every single time on every single query, for example. And yes, even within the exact same query, it's doing that many times because it's generating the next token and the next token.

40:18And so there's something called a KVCache. Funny enough, I was going to put KVCache on my license plate because I just shipped my car from another place. And somebody in California has already taken the word KVCache on their license plate. So I'm like, that's how popular KVCache is.

40:36Jon Krohn:You can probably in any other state in the US still get that. Totally, right? So someone beat me to it. But think of it just like any sort of cache system where you're pre-computing. But on a single query, these caches can be humongous. I mean, especially with people that are uploading documents, that are trying to scan things, there's a lot of computations happening. So, now you've got a case. So, there's building the cache, which is step one, and then there's actually generating the next token, which is step two. And so, and, you know, yeah, getting, right, like starting to actually generate the response.

41:27So those are called pre-fill is the first step where it's like generating all of this, these computations that's extremely, extremely compute heavy. And so that's where you potentially want a node or a set of nodes in your cluster that are configured differently. So memory bandwidth, it's still important, but it's not as big of a deal as the second step, which is actually generating those tokens. Because that second step has to move all the weights back and forth from your GPU RAM, sometimes CPU RAM, based on how big the GPU RAM is, sometimes on disk. So these are very different configurations.

42:13And so you will often hear people talk about disaggregated PD, which is disaggregated pre-filled decode. I would guarantee you if you are going for like an interview in 2026 and need to study something for your interview, it is going to be about this pre -filled decode. It's going to be about KVCache. So if you're going to talk to any of these big labs in any sort of engineering capacity study, I can't remember what chapter it is in my book because it was supposed to be a half a chapter actually, and it ended up being, I think, three or four. So there's trying to understand it, and then there's trying to optimize it.

43:01Jon Krohn:And pre-fill, decode, disaggregation, and KV cache. There you go. There are your interview tips of the episode. Speaking of tips and ways that people can be getting ahead more quickly, you mentioned to me prior to us recording that you've been loving using AI coding assistants. And so, you know, tools like Cursor, Cloud Code, Codex. How has your workflow changed as a result of these kinds of tools? How does that impact in particular AI systems performance, optimization? and yeah, what do our listeners need to know about where this is going? For sure. I tweeted about this a couple of weeks ago.

43:44If you're manually writing code in this year, 2026, you are way behind. And I would not have said that if I didn't spend all of last year watching this sort of evolution because I started off in my sort of legacy ways, right, where I was writing every line of code. And I vowed for 2026 to only use these coding assistants and to see how far I could get. Now, I've done a couple of projects for some friends and for some portfolio companies. And just to make sure that, you know. So the short of it is you can fire off about 10 to 15 different things, like different aspects of either features that you're trying to build or bugs that you're trying to find.

44:41I'm personally using these tools right now to do a lot of optimizations and to look at different aspects. And so I'll say, take a look at the occupancy percentage for this particular kernel that I'm trying to optimize. At the same time, I'm having another GPU that has the same code that's analyzing a different part of it. And so think of a four GPU system and you can run separate experiments, each one running on a different GPU or even four separate GPUs, like separate nodes that have their own memory and stuff. But yeah, my personal workflow, I'm going to be a little controversial and say I prefer codecs.

45:28Yeah, I prefer the open AI. One thing, and again, very, very workload specific, I would say. I think if I was doing UI stuff, I would probably maybe prefer Claude right now. Claude's great with the UI. It's kind of my little dirty secret that I use Codex. And so while all the other developers are flooding, right, the anthropic GPUs and the inference stacks, I've got kind of my little group of people that just use Codex and I still have really good performance until people realize that, you know, Codex is actually better for other stuff. But for right now, I'm enjoying good performance and a fair amount of, you know, token.

46:08But the one thing I would recommend for AI systems, performance would be be on the machine. Don't try to do this stuff on your MacBook and hope that it works or even on a personal, like NVIDIA GPU, like an RTX. Yeah, I don't know, 50. Like, I've literally never used any of the personal GPUs because to me, the profile is so different. The hardware is so different. The memory bandwidth completely different. So if you are working on an on an AI project and you are ultimately going to deploy on a Blackwell or on a hopper, you have to be on that machine during development. Don't try to take any shortcuts.

46:50I recently tried the DGX Spark, which is the little mini one that got a lot of buzz, you know, the end of last year. There's so many things disabled on that, that none of my benchmarking was even working because it's just completely missing a lot of, you know, core hardware components and then, you know, software components on top. The driver isn't the same. It's, you know, a lot of different things.

47:14Jon Krohn:Nice. Really appreciate your insights on what you're doing with the coding assistance today. When you're working on those, Chris, do you still review all of the code before it goes to production? I did. I was once, yeah, yeah. I was once like you or once like the question that you're asking. In fact, I was just working with, with yeah, I was just working with someone this weekend, trying to get them a little bit up to speed. It's a good friend of mine. And I'm like, look, you're doing things like you're going to be outdated within the, you know, by St. Paddy's Day. Yeah, it was the joke specifically.

47:53And I was watching their workflow. And, you know, these assistants, they like to write little snippets of Python code or, you know, any or bash scripts to get stuff done. And this person was literally reading every single line of the bash script. Like it's happening during the thinking process and they would go up and they would expand it and they would be looking and they try to find the script on disc. And I'm like, dude, you have to let go. So the short answer is I did and it's exhausting and I've learned to step back and there's like temporary scripts that these things write to write the final code.

48:41And you could do that and you could be very disciplined and review all the lines of code. You could write tests. You can have the LLM actually write the tests as well too. I've actually even sort of let go of unit tests. And this is very controversial for all of my test-driven friends and folks that are listening. I focus on the evals. So I clearly set up and provide examples that I know are correct. Those are sort of the collaboration set. And then I use the LLM to actually judge the quality. I have it constantly running. So I've got evals that are always running in the background in a separate tab while I'm, you know, building the software.

49:34So short answer is no, I stopped doing that. And it really slows things down. And it's hard for people to do. But I'm shipping code a lot faster. I'm, you know, with kernel optimization, you have to be very careful. these models want to hack the reward, or it's called reward hacking, and just get the fastest thing possible. And if that means zeroing out everything to make the computations a lot faster, it's going to do it. And so one thing I spent a lot of time on with my GitHub repo is these correctness checks and correctness verifications. And these pop up every single month, someone like release some small little startup releases something that they have achieved, you know, 50x speed up.

50:30And the first thing I do is take that code and I put it into my harness and boom, I see that they're not using CUDA streams properly or that they have that it's not their fault. It's just that they've given the LLM too much freedom. And so we have to get better about what it means to review. And I don't think even going through all the code that you would be able to catch these things, because I assume that they went through and looked at the code, but not unless you're actually running the code on the machine, and you've got all of the profiling, you know, set up where it can show you how many streams are being used and, you know, all the different aspects.

51:11If you don't have correctness checks, if that's in the form of evals for your sort of end user application or, you know, performance metrics and like, real, you know, correctness checks within your, your performance harness, then you're not doing it the right way.

51:26Jon Krohn:Yeah, really cool. I guess I need to stop looking at the code. Yeah, do you still do that? I'm living in the stone age. So my last semi technical question for you, before we start wrapping up, there's so many interesting things about you. We actually, we've really just scratched the surface here. We focused on your latest book, which again, you know, if people like this episode, you've got to get Chris's latest book, AI systems, performance engineering. Um, but in addition to that, you've done tons of other interesting things. So your previous two books were both, uh, they're both on, they were like on AWS books.

52:04Jon Krohn:It was like data science on AWS. Uh, I think that might've been with Ante Bart. that was the one that was with her. And I think you ended up at AWS because you were the chief product officer at a company called Pipeline AI that was acquired. And you've also, you've worked in principal engineering roles at Databricks and Netflix, including getting an Emmy at Netflix for the work that you've done. But the thing that I want to talk about, the only thing that we have time to talk about now is your role as an investor. So you've been an investor or advisor in companies like XAI, which is now at the time of recording, it was just acquired by SpaceX.

52:43Jon Krohn:You invested in Grok, G-R-O-Q, which makes LPUs, language processing units. And that firm was acquired by NVIDIA. And so, yeah, I'd love to understand your thinking on what's going on with hardware. Do you believe, for example, that the future of AI performance lies in specialized LPU hardware like Grok builds or in better software orchestration on top of general purpose GPUs? And you can also answer this question more broadly with just generally where you think Harper is going and where people should be putting their money? I'm in a fortunate position where I do have extra cash and I can take some chances.

53:32These are all secondary market shares. They're not public companies or some are about to IPO and things. On the investor side of it, I look for, this is very trite these days, but you have to look at the founders. Yes, Elon, he's an absolute nut job, of course, and he's very controversial. And my sister hates Teslas because of him and all this kind of stuff. So I can't even talk about some of these investments when I'm home for Christmas and things. But you have to look at their ability to fundraise. Right now, the secondary market is where all the action is. And so if you're still investing in Apple and things like you're, you know, you're like behind the times like this is not, not where the money where the money is currently being made, especially because there's been such a dry spell with these IPOs.

54:35So, you know, hopefully that changes. Yeah, I'd say my best investments so far have been Databricks. Of course, I was an early employee there back in the day, and I know those founders very well. You have to look at folks that can raise money, that can convince people that they're on the right path, that have their stuff together. Where are things going? There's so many different types of chips coming out. In fact, I'm about to go to lunch with some folks at Etched, which is another little player and similar to Grok with a Q. It's good to keep an eye on these folks to see sort of generally where things are moving.

55:19I mentioned before that the biggest bottleneck is this memory bandwidth. This is something that's that like this is the main reason why the NVIDIA folks purchased with a Q or did their little weird acquisition, you know, licensing deal is the, like LPU is, has all of that memory, like on the same chip, right? Like right next to the actual, right? Like the accelerator itself. And so that then decrease and that's really good for like inference. Cause like I said, when you're doing inference, there's that a decode step that has to move a lot of weights around that's very, very memory bandwidth bound.

56:01And so it's a really great complement to the Nvidia chip, which suffers from this, you know, memory bandwidth limitation. And it's just a fundamental architecture design that dates back to their graphics days. And so instead of building a cycle of, you know, five or six years out where where they have the chip they ended up. So, and also with like XAI, it's interesting because I passed up on SpaceX a few years ago, as I didn't see the leap that Elon is currently talking to. I couldn't, I didn't have enough vision for that. I saw that I liked XAI in that, you know, Grok with a K, which I actually personally don't really use that much.

56:50I should because I'm an investor. But I liked the work ethic of those guys. I actually met with them quite a bit last year while I was writing the book. They're big backers of SGLang, which is like VLLM, which is an inference engine, but it's SGLang. It's another project that kind of came out of Berkeley around the same time, VLLM, that does things a little bit differently. So we had some conversations. Some of those conversations actually are part of the book now. What's the specific question again? I talked my way out of that. Yeah, I had a question around, it's basically, where is this going?

57:35Jon Krohn:Where are the opportunities for investing? And of course, we've also got to give the disclaimer that we are not investment advisors. But I think it's just something that people are interested in. You've got a really interesting being in the Bay Area with the kinds of companies you've invested in, you have kind of an inside baseball perspective on where things could be going. You've already given us some of the tips around things like, you know, your advice is that the secondary market is where there's a lot more money to be made than in the public markets, your apples, for example. But yeah, and you also kind of already answered the question about whether the future lies in specialized LPU hardware or more general purpose GPUs.

58:13Jon Krohn:So I think you've covered that, Chris. The folks over at Together AI, the person, I think he's the chief scientist, one of the co-founders of this company called Together AI, is a big fan of general purpose GPUs in that it gives you more flexibility to support different types of algorithms. Not all the world is going to be transformers. I mean, this is going to change, it just has to change because we need a better. And all of these chips that are specialized for transformers may not be the best chips going forward, which was quite honestly a tiny bit of concern when I made the Grok with the Q investment is, oh crap, if something changes and another type of model starts to dominate, like another type of algorithm, you know, where is that going to leave folks like Grok with a Q?

59:16But so, you know, fortunately I've got shares on, you know, both sides of it. I've got public NVIDIA and I've got, you know, private things as well too. So you have to kind of have a balance there. But yeah, and I'm the worst investor in the world. I sold Netflix when it was low and I bought it when it was high and then it went low again and I lost a bunch of money and, you know, things like that. But so don't take advice from me. Great, great. I just have to be getting lucky.

59:43Jon Krohn:Buy high, sell low is the takeaway advice from Chris Fregler today on investing. Great. Yeah, certainly look at the addressable market. If you're familiar with Sequoia, it has a very famous pitch deck. And I'm very familiar with this pitch deck because I spent many years fundraising for my startups. And when I go into even a job, if I start talking to folks about potentially joining, even as an advisor or any type of role, if I decide to go back into engineering full time or whatnot, you want to stick. there's a famous, I think it's 12 or 13 slides. And it seems so, so silly to only have 13 slides to make or break, you know, millions of dollars, but you know, think of like, yeah, try to find that slide deck.

1:00:37It's, it's out on, um, yeah.

1:00:40Jon Krohn:And it's, it's the subject of countless blog posts out there having the experience of building these decks myself. It's, you're not going to have a hard time finding, you know, guides to, to your, you know, dozen page deck. And I think the key takeaway, maybe this is exactly where you're going. I don't mean to take the words out of your mouth, but if you can't explain in that short number of slides, why this is a compelling investment, it's probably not a good investment. Yep. And value proposition to me, that is the most important slide. I mean, yeah, obviously team and, you know, yeah, like assuming team and all that stuff's there, but if you can't clearly explain why now and why, you know, yeah, why am I the right fit for this?

1:01:22And, and, you know, what's the real value? Like, that's always like, to me, the hardest slide is like, yeah, how do I quantify this? And, but yeah, I mean, always look at, you know, the addressable market. And also, you know, these days, it comes down to the, like, business model and the pricing, gone are the days of this SaaS-based seat pricing. Now it's all about value, right? So like value-based pricing and yeah, they call it outcome-based pricing, right? Which is very obvious for things like customer service and stuff, but yeah, trying to think about that. So you really want to look for founders that are in tune with that and especially repeat founders that might not have that have traditional have been part of traditional uh you know pricing models business models yeah you really gotta suss that out great thanks for

1:02:15Jon Krohn:all of those inside baseball tips chris on investing and yeah so now we'll start wrapping up the episode you already know you've been warned by antir barth that i always ask for a book recommendation at the end and don't let you use your own book and so i think you have come prepared Yeah, one of my favorite books is called Founding Sales. And these days, it's a relatively old book. It's from August of 2020. Founding Sales, the early stage go-to market handbook. And I'll tell you why this is an important book and was so critical to me. So, the author goes through his evolution in learning about go-to-market GTM.

1:02:59I used to laugh when my friends that came out of business school would talk about go-to-market. It was so far from the things that I would think about on a day-to-day basis. And then I became a founder. And then I was like, oh crap, I have to figure this out. So there was a point in time and I still, I'm fortunate that I live somewhat near Stanford. So I take a lot of weekend classes and I don't have any kind of executive degree or anything. I just have an undergrad from Northwestern, but I can soak up like what is valuable to me. So really thinking through what does go-to market mean specifically in a SaaS environment And not all the things will be relevant to LLMs, of course.

1:03:44This is an older book. But it really gets you to see the pain points that he had gone through and the different pivots that he made and how his competition really affected his route through this. So I point this out because I think this audience is mostly technical, right? And if you ever do think about becoming a founder, yeah, check out Founding Sales because it's a nice sort of engineering friendly way, like logical way to think through how to come up with your real true value proposition for your ideas. And you'll read something and say, oh, that's exactly what I was thinking. And now I see how it's bad, how it's not really going to convince anyone to use this product or to give me money.

1:04:37Jon Krohn:Nice. That's a really great tip. And I like how it follows along nicely with actually the arc of this episode as well, where we started off really technical, really heavy in the weeds on AI systems optimization. And then we kind of ended with commercial discussions and a commercial book recommendation. So perfectly on theme, Chris, in the end. You didn't know that's where I was going with questions. So pretty cool. And yeah, then my final, final question before I let you go is how can people listen to you after this episode? Obviously, your book, AI Systems Performance Engineering is the way to go.

1:05:10Jon Krohn:I'll have a link to buy that book as well as to the GitHub repo for the book. Beyond that, how should people be following your thoughts after this episode? Yeah, I'm C. Fregly, C-F-R-E-G-L-Y, pretty much everywhere. I recently discovered I have a second cousin named Chris Fregley who does lots of van conversions. He's in tech. He's a Kubernetes guy and we actually live pretty close to each other. So we've been connecting. He's half my age, which is real fun to get his perspective on the tech industry. But we're both in San Francisco. So, but yeah, I'm the original C. Fregley. Yes. Okay. I see you saying like, watch out for that guy.

1:05:56Jon Krohn:Don't follow him. If there's van conversions you've gone, it's the wrong guy. Totally. That's the most like unique thing is I don't do van conversions. I prefer to live in a house, you know, condo. Yeah. So see Fregley on Twitter, see Fregley GitHub. See, or yeah, Chris Fregley LinkedIn. Perfect. Yeah. We'll of course have links to all of those in the show notes. Chris, such an honor to have you on the show after the kind of decade of being aware of you. You're such a rock star. Thank you for taking the time. And yeah, next book comes out, you've got an open invitation. I am swearing off book writing.

1:06:34Yeah, which I said after the first one, after the second one, but yeah, we'll see. So, hey, John. Yeah, man, great chatting.

1:06:42Jon Krohn:What a great episode today with Chris Fregley in it. He covered how DeepSeek R1 achieved 10 to 20X training cost reduction over comparable Western models by co-designing hardware, software, and algorithms together, a recurring topic throughout the episode. He talked about how memory bandwidth, not flops or tensor cores, is the single most critical metric to pay attention to when evaluating GPU performance from generation to generation. He talked about mechanical sympathy, a term borrowed from Formula One that means understanding the full hardware-software stack, and Chris argues it's the most valuable skill for anyone studying computer science or related jobs today.

1:07:19Jon Krohn:He talked about how the PyTorch profiler only shows the tip of the iceberg. There are 50 to 60 additional GPU level metrics involving streaming multiprocessors, occupancy, specialized function units, and instruction pipelines that are essential for real optimization. And he spilled the beans that he's abandoned line-by-line code review in favor of continuous evals and correctness harnesses and says that if you're still manually writing code in 2026, you're way behind. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Chris Fregley's social media profiles, as well as my own at superdatascience.com slash 973.

1:08:01Jon Krohn:All right. Thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, partnerships manager, Natalie Zajski, researcher, Serge Masise, writer, Dr. Zara Karche, and our founder, Kirill Aramanko. Thanks to them all for producing another super episode for us today, for enabling that super team to create this free podcast for you. We're deeply grateful to our sponsors. You can support the show by checking out our sponsors links, which are in the show notes. And if you'd ever like to sponsor the show yourself, you can get the details on how at johnkrone.com slash podcast.

1:08:35Jon Krohn:Otherwise, share this episode with people who would want to listen to it, review it on your favorite podcasting platform or on YouTube. Subscribe, obviously, if you're not a subscriber, but most importantly, just keep on tuning in. I'm so grateful to have you listening, and I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there, and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

From the publisher

No one should be manually writing code in 2026, thinks Chris Fregly, Jon Krohn’s guest on this week’s episode. In this interview about Chris’ latest book, AI Systems Performance Engineering, he explains why it’s so important to consider memory bandwidth when evaluating GPU performance, that understanding the full hardware software stack is the most valuable skill for anyone working in AI development, and which shortcuts we still shouldn’t ever take when writing code, even though we might be outsourcing a great deal to generative AI.

This episode is brought to you by the ⁠⁠Cisco, by Acceldata and by ⁠ODSC, the Open Data Science Conference⁠.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/973⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(03:39) Why Chris wrote AI Systems Performance Engineering 

(21:39) Essential coding metrics 

(37:24) The importance of inference when coding

(42:11) How to manage workflows while using AI agents

(51:37) Where and how to invest in the AI market

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
973: AI Systems Performance Engineering, with Chris FreglySuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 12 min
Listen in VO