How to Engineer AI Inference Systems with Philip Kiely - #766

30 Apr 2026 · 55 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Inference engineering for AI systems—how to design, optimize, and operate end-to-end systems from user request to GPU execution, emphasizing fast “day zero” support for new model architectures and rapid research-to-production timelines.

Guest

Philip Kiely, head of AI education at Base 10. Background: joined Base 10 in 2022 (before public ChatGPT); worked on inference-focused ML problems, from early CPU/T4 GPU setups and XGBoost to larger models (e.g., Whisper’s 1B parameters) requiring many GPUs. Has 4+ years at Base 10 and experience across the shift from predictive to generative and from internal ML to revenue-critical inference.

Key claims

Inference is the stickiest and most important AI workload; it requires broad expertise (CUDA/PyTorch, quantization/speculation/KV cache, distributed systems) plus ownership of the full user experience. Applied inference research can reach production in hours (example: PoloQuant implemented as a CUDA kernel in 31 hours). Companies should treat inference as a spectrum of trade-offs (latency/throughput, batching, speculation, quantization, KV cache sensitivity).

Notable examples

PoloQuant; Whisper; KV cache/prefix sensitivity; named entity recognition runtime (Base 10 claims 1ms vs 500ms with frontier LLMs); Talos ASIC demo (~16,000 tokens/sec); GPU generation longevity (L4/L40, Hopper staying power) and 2026 predictions (inference disaggregation, more specialized hardware).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Philip Kiley's Journey to Base 10

0:46 to 2:56

Philip shares his background and path into inference engineering.

“Welcome to another episode of the TwiML AI podcast.”

The Evolution of Inference in AI

2:57 to 4:53

Discussion on how inference has changed and its importance in AI.

“And I've been at it for more than four years now.”

Complexities of Inference Systems

4:54 to 7:02

Exploration of the challenges involved in building effective inference systems.

“And then as, you know, true generative models came around, we were already, you know, thinking about an inference first view of the world.”

Rapid Research to Production in Inference

7:03 to 11:41

Insight on how quickly research translates into production within inference.

“I think there's, as an inference system, you have to own the entire experience.”

The Broader Impact of Inference Engineering

11:42 to 13:32

Discussing who should care about inference engineering and its growing importance.

“To be clear, trading with a D, not training.”

The Growing Demand for Inference Engineers

14:04 to 15:49

Explore the increasing need for inference engineers and the impact on AI applications.

“Well, one reason to care about it is that it's a really interesting field that you might consider working in.”

Understanding Inference and Its Implications

15:49 to 19:51

Learn how inference knowledge influences product strategy and user experience.

“You know, so let's talk about like when and why and how like knowing about inference impacts, you know, the decisions I'm making and the strategy or approach that I'm taking.”

Navigating Inference Providers and Options

19:51 to 21:45

Discover the spectrum of inference solutions from closed models to dedicated services.

“A lot of the other examples, do they, you know, to take advantage of, you know, things like things that you mentioned, like, you know, changing batch sizes or, you know, different tiers of service?”

Transitioning to In-House Inference Solutions

21:45 to 27:15

Understand the journey and considerations for moving from external to internal inference systems.

“But as companies of all kinds, especially like vertical AI native companies are achieving the massive scale that they are, everyone is thinking about taking control over their AI, owning their intelligence.”

The Lifecycle of GPUs in Inference

27:15 to 28:07

Examine the lifespan and relevance of GPUs in the evolving field of inference.

Show all 17 chapters

The Evolution of GPU Generations for Inference

28:07 to 31:18

Discussion on the lifecycle and usage of different GPU generations in AI inference.

“like haven't really been all the way through blackwell because blackwell isn't fully rolled out yet.”

The Future of Inference Engineering

31:18 to 36:46

Exploration of the complexities and future of inference engineering in AI with the impact of AI tools.

“but I am very bullish on the sort of valuable lifespan of GPU generations, especially starting with Hopper.”

Optimizing Inference for Multimodal Models

36:46 to 42:03

Insights on optimizing inference for multimodal models and the need for specialized systems.

“We talked a little bit about this idea of like adding constraints as a way to get to higher performing systems or more efficient systems.”

Optimizing AI Inference Systems

42:03 to 44:38

Learn about the importance of specialized inference systems for AI workloads.

“as a specialized model, maybe a highly optimized small LLN can do that task in 500 milliseconds.”

Modality-Specific Considerations

44:38 to 47:16

Discover how different AI models require unique performance goals and architectures.

“But if the VLM is any example, then, you know, each of these modalities has its own quirks that you need to address.”

The Role of Software and Open Source

47:16 to 50:00

Understand the significance of open-source software in developing AI inference engines.

“different performance goal and having a different performance goal completely changes the way you architect your system.”

Hardware Specialization Trends

50:00 to 52:36

Explore how hardware specialization is evolving to meet AI demands.

“I think that what we're seeing, you know, maybe 2026 is the year of a lot of things.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I mean, it might be the fastest timeline in the world. If you think about medicine, for example, it can take decades for research to reach a pharmacy. If you think about, you know, physics or engineering, it can take years to apply a new concept. Even within AI, if you want to train a model off of a new technique, it can still take weeks or months to find the exact right way to sort of express that technique. But with influence, the timeline is often ours. A new model architecture comes out. You have to figure out how to support it day zero.

0:46Sam Charrington:All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Philip Kiley. Philip is head of AI education at Base 10. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Philip, great to see you again and welcome to the podcast. Hey, Sam. Thanks for having me on. Absolutely. Absolutely. We met at the recent GTC conference where you were signing copies of that book that you have there over your, I guess, your left shoulder. Yeah. And I've got one here on my desk. Hey, there you go.

1:25There you go. That was a really fun conference. I felt like a rock star, you know, just like handing out. We also had, honestly, I think part of it was we had free ice cream at the booth as well. And I gave away a lot of books, but I did run the numbers at the end of the show. The ice cream was more popular.

1:45Sam Charrington:you know i i thought uh it'd be good to get you on and talk through like you know a little bit about you know inference engineering you know what what you're writing about in the book but also kind of what you're seeing beyond the book and um what is really kind of making a difference from you know inference perspective nowadays so before we get to that though i'd love to have you you talk a little bit about, you know, your journey, like what brought you to, you know, base 10 and inference engineering? So back in 2022, this was 10 months before chat GPT was launched to the public. I had a really strong and novel thesis about the AI industry.

2:31That strong and novel thesis was I need a job. and so i i applied for a bunch of jobs and i ended up choosing between base 10 and a laundry delivery startup um and it was like a really tough decision uh because hey it's it's a good the other one's a good company they're still around um but ultimately i i decided to go work on some some really interesting problems in the ml space with a tiny startup at the time called Base 10. And I've been at it for more than four years now. I've had a really fortunate position to be able to learn on the job and experience the rise of the AI industry firsthand, going from predictive to generative modeling, going from modeling being something mostly done for internal workflows to modeling, you know, being part of the cost of goods sold, modeling being in the revenue path for AI native companies and going from like very, very small use cases because of that to very, very large and complex ones.

3:40Sam Charrington:You know, one thing that's interesting about your experience, and I'm thinking about Base10 in particular, like Base10, as your timing indicates, predated ChatGPT. And they were, you know, one of, you know, a great number of companies at the time kind of competing around this MLOps, you know, set of ideas. And with the rise of generative AI, a lot of those companies, you know, got acquired. A lot of those companies went away. But of those companies that are, you know, still around and doing well, like a lot of them are in inference. Why is that? So inference is, I believe, the most important workload in AI.

4:29And it's definitely the stickiest. That's been one of the reasons we've treated it as our wedge into the AI infrastructure market. Inference is very, very complicated. And doing it well is really critical to any company that actually relies on generative models to power their product. so we started as an inference company but we were just influencing very different models i was you know taking cpus and i was taking some you know t4 gpus or a10gs when i was feeling real fancy and running you know xg boost models on them running very early like wave 2 vec and gptj type models and honestly just like having a lot of fun with it uh we were building simple classifiers dashboards that kind of stuff and found over time that like as the models were getting a little bit bigger and more complicated like i remember when whisper came out that was huge because oh man now the model is a billion parameters like can you imagine one billion parameters in just one model like i can't believe i'm going to spend all my parameters in one place but this model now it's like it's really capable but now it needs gpus you need a lot of them all of a sudden and it's hard to get it running it's hard to make it fast enough there's a lot of interesting problems all of a sudden in in running models and so i think it it just kind of naturally occurred that like the most complicated piece of the broader ml serving stack became inference as these models got larger.

6:08And then as, you know, true generative models came around, we were already, you know, thinking about an inference first view of the world. And it felt natural to just extend that to larger and larger and more capable models.

6:21Sam Charrington:You referenced both model serving and inference as a component of model serving. Like, where do you draw the line, you know, between the beginning and end of inference and the broader model surveying set of requirements? I probably use those terms a little too loosely and a little too interchangeably because the technical definition would be that like model serving is everything that happens end to end from a user request to a user getting the response and inference is the piece that happens on the GPU. But when I think about inference as a discipline, I am thinking about end to end. I think there's, as an inference system, you have to own the entire experience.

7:09And so to do that, you have to think not only about what's happening on the GPU, but what's happening every step of the way on the user journey.

7:17Sam Charrington:Yeah, so let's talk a little bit more deeply about like what makes inference difficult. You know, clearly there, you know, we know the models are getting bigger. We know we don't have enough compute in the form of CPU, GPU, etc. You know, talk a little bit about some of the other challenges that you see with regards to doing inference well. Yeah, so inference is a really fun and difficult topic. And I'm going to go into metaphor territory for a second here, which is that I'm a martial artist. I have been my entire life. Are you familiar with like UFC and MMA and all those kind of things? so you know there you can't just be an expert in one thing and expect to become a champion you can't just be a great wrestler you can't just be a great boxer the idea is that you have a lot of different skills each one of which can take a lifetime to master and somehow you have to be excellent at all of them in order to be a well-rounded mixed martial artist and you know in my background, I'm certainly no UFC champion, but I've done striking of various disciplines.

8:25I've done wrestling in various disciplines, and I've started to understand how those things blend together. I think inference is very similar. You need to have a wide range of expertise on a lot of very different complicated topics to build a truly effective inference system. There's what happens on the GPU, right? There's understanding CUDA level programming and PyTorch on top of that, and then the inference engines on top of that. There's being able to take and apply the research, the understanding of different quantization techniques, different speculation algorithms, different KV cache reuse mechanisms.

9:01How do you parallelize models across GPUs? How do you disaggregate across different types of hardware? So there's all that applied research. And then on top of that, there are all the traditional problems of running very large scale distributed systems. Anything that you're familiar with from the last 20 years of cloud architecture, as well as understanding how to take very, very large volume workloads that are coming from around the world and handle them appropriately. So every piece of backend web infrastructure, every piece of GPU programming, it all comes together in inference in Windows and packages that are very demanding.

9:43You often have a couple hundred milliseconds latency SLAs that you're dealing with. So it's the orchestration of a large number of very complex systems together that makes it such a difficult problem.

9:57Sam Charrington:In AI broadly, there's this very tight relationship between research and implementation. but an inference in particular like I've noted on several occasions like you know hear about the speculative coding work one day and like I hear about it all over in two weeks and everybody's already using it like do you find that that research to production timeline and inference is you know particularly rapid relative to other aspects of AI? I mean it might be the fastest timeline in the world. If you think about medicine, for example, it can take decades for research to reach a pharmacy. If you think about, you know, physics or engineering, it can take years to apply a new concept, material science.

10:50Even within AI, what moves faster? Training ones, for example, if you want to train a model off of a new technique, it can still take weeks or months to fine-tune the hyperparameters and find the exact right way to sort of express that technique. But with inference, the timeline is often hours. A new model architecture comes out, you have to figure out how to support it day zero. We had PoloQuant come out, that research paper, and an engineer on our model performance team had it implemented 31 hours later as a CUDA kernel. So the pace of applied research is very, very fast because it is a highly competitive industry and everyone is sort of searching for that next edge.

11:37Maybe the only industry where research goes into production faster than inference might still be like trading. To be clear, trading with a D, not training.

11:47Sam Charrington:Yeah, this might be a little bit too in the weeds, but with regards to PolarQuant, like the paper was a year old and then all of a sudden it got really popular last week or a couple of weeks ago? Like, what's the backstory there? Yeah, so it was sort of republished as a blog post and sort of caught everyone's attention. I think, like, as a little detour, you know, I grew up in Iowa, so I was very influenced by sort of the best Midwestern school of thought in economics, which is the Chicago School of Economics, famed creator of the Efficient Hypothesis. I think the most dangerous falsehood that you can expose an impressionable young person to is the efficient market hypothesis.

12:35And the idea that like everything is already figured out and everyone already knows everything because that's just not true. You know, with the volume of research techniques that are coming out, I think that it's easy to lose track of things. But it's also a case where there's a lot of low-hanging fruit in the inference world that has already been picked. And so now, as we try to, you know, figure out what's next to push the envelope, sometimes techniques that might not have seemed, you know, either feasible or worthwhile, even as recently as a year ago, all of a sudden become like, oh, wait, now we're ready for this.

13:15Now it's time to add this to the sort of industry standard and move the inference stack forward.

13:21Sam Charrington:It was just that, you know, the bump in the email that brought it back to everyone's attention and the timing was just right. It worked. You know, if the great thing about a launch that doesn't go viral is that no one knows about it if you want to launch it again. Exactly. You know, let's talk a little bit about who should care or who does care about inference engineering. Like for many, you know, that's somebody else's problem, you know, and it's, you know, the problem of engineers at like a handful of really large companies. But I suspect that there are a number of reasons why a broader set of folks should care about it.

14:03Sam Charrington:Like, what's your take on that? Yeah. Well, one reason to care about it is that it's a really interesting field that you might consider working in. There was, you know, three years ago, there were maybe a couple hundred inference engineers in the world. They probably didn't even call themselves inference engineers back then. And now there are thousands, tens of thousands, probably, depending on, you know, where you make the cutoff. and software engineering of course undergoing a massive change in the way it works and the best engineers are becoming substantially more productive there's a lot of questions around how many engineers we're going to need in the future and what our jobs are going to look like I believe that even with the advances in AI assisted code generation even with the increase in productivity on an individual engineering basis, there's still going to be demand for 10 to 100 times more inference engineers in a couple of years than we have today.

15:05I mean, I know at Base 10, we can't hire people who are knowledgeable about inference fast enough. And even if you're not wanting to work at an inference platform, the other reason to care about it is that every vertical AI application company is eventually going to have to figure out what their inference strategy is. And it's going to need knowledgeable members of technical staff who understand the actual technologies and trade-offs involved in inference to sort of develop and execute a competent strategy. Because great inference can be the difference between a really fast product that feels magical to users and a slow buggy experience that causes constant shown.

15:49Sam Charrington:You know, so let's talk about like when and why and how like knowing about inference impacts, you know, the decisions I'm making and the strategy or approach that I'm taking. Like a lot of people will be delegating inference to a provider, whether it's a, you know, first party model provider or a third party, you know, inference provider like Base 10 or Firefly or someone like, you know, what are the situations where I'm going to want to, you know, have deep knowledge about inference? Yeah. So I think the first thing is understanding what's possible, because in a world where most AI engineers are sort of trained on GPT and cloud models, you understand the world in a certain way.

16:39I have access to N models, and there's maybe three or five or ten. And each one of them has a sort of fixed level of intelligence, a fixed level of speed in terms of tokens per second, time to first token, maybe a range of that. A fixed cost in terms of input output tokens with maybe a couple variables around something like a cash token price or a discount for long running batch jobs. And then also a reliability. You know, you have an endpoint where you might get a couple nines of uptime and maybe need to engineer the rest of your product around that reality. And the first sort of lesson of inference engineering is going from taking all of these as sort of immutable discrete points and understanding them more as a spectrum.

17:31You can trade off between, for example, latency and throughput in a given inference engine by adjusting things as simple as batch size or things like, you know, adding or removing a speculation algorithm. When you do that, when you create sort of a spectrum of outcomes, an efficient frontier of high performance inference, then you start to understand like, wait, I can choose to change the way that I consume these systems. So, for example, you might start having priority queues in your product where paid user traffic gets prioritized over free user traffic. And maybe that's something that you couldn't have built previously.

18:18Or maybe you add a concept of, you know, maybe you understand, like, I have the ability to quantize models. And because I have my own very sophisticated and product-specific evals, I can do that with complete confidence that I'm not degrading the quality of the service my users are experiencing. And then when you quantize your own model and start to serve it faster and less expensive and with a great deal of confidence that you've calibrated it properly, then you are not subject to sort of random intelligence degradations from a third party provider who might not understand the exact nature of your workload.

19:01Instead, you've achieved the sort of price and performance goals that they're going for with that, but actually, you know, kept things appropriate for your task. So, you know, if you don't know a lot about inference, you might think like, oh, all quantized models are bad, when in fact you can calibrate them appropriately to your workload. There's all sorts of, you know, examples like this, even things as simple as, you know, understanding that the way KVCache reuse works means that you have to think about only your prefix and even one token difference early in a sequence can throw off the entire rest.

19:38And maybe that's how you structure your chat template or how you structure your inputs to your models. There's all kinds of ways where understanding what's possible with inference helps you build a better product and helps you deliver a better user experience.

19:55Sam Charrington:The last example you gave about the KV cache, KV caching, you know, is a clear example of it doesn't matter whether you're, you know, building out your own inference system, you know, you benefit from having that knowledge. A lot of the other examples, do they, you know, to take advantage of, you know, things like things that you mentioned, like, you know, changing batch sizes or, you know, different tiers of service? Like, do you have to host your own inference service and, you know, do inference on your own models in order to take advantage of those? or are there, you know, providers that like allow you, I imagine like everything, there's like a spectrum of the number of knobs you get when you're doing inference.

20:48Our philosophy at Base 10, not to, you know, cut a promo too much, is that like you as the end user should have access to every conceivable knob. And of course, along with like guidance on, you know, which ones are hot to touch and which ones aren't and what some sensible defaults are. But if you are making the transition from working with a closed model where you don't have access to the weights to working with either an open model or a model that you trained yourself and doing dedicated deployments, which basically everyone at a reasonable scale starts to need to do, where instead of paying on a per token basis, you're paying for the underlying GPU hardware, then you have access to these knobs.

21:35Whether or not you as the engineer have the sort of knowledge and confidence to tone them is why this discipline of inference engineering exists. But as companies of all kinds, especially like vertical AI native companies are achieving the massive scale that they are, everyone is thinking about taking control over their AI, owning their intelligence. and part of that means owning their influence outcomes, taking these knobs and figuring out the right positioning for their product and for their use case.

22:08Sam Charrington:Let's talk a little bit more deeply about the levels of control that one might pursue and it may or may not be a strict sequence, but what are the things that are happening in a business that, you know, drive the transition from, you know, one state to another. Like, so you talked about paying per token to paying per GPU time. You know, that spectrum goes to, you know, I guess I'm trying to get at like, you know, you could use a provider like you guys, you could, you know, stand up a Kubernetes cluster on like an infrastructure provider. You could, you know, host it all, you know, in your own building.

22:58Sam Charrington:Like, what are the dominant models or the dominant setups that you're seeing? And like, what drives a company to go from one to the next? The journey I sort of see is, in many cases, aligned with a product maturity cycle. And I want to specify this is really a product maturity cycle. This is not a company maturity cycle, because oftentimes you'll see relatively small companies further along in the cycle versus relatively large companies, just because they're sort of either more AGI pilled as a company, or just because the AI component of their product is more critical to what they're doing than the AI component of maybe the larger and more established product is.

23:50So as we think about the maturity cycle, there's a few options. Option number one is to rely entirely on pro-token closed model providers. And everyone starts here just about. Everyone should start here. It's really easy. You get frontier intelligence with an API key. And that's a hard thing to beat when you're starting out. I think the next level often looks like running into one of two problems, either a cost problem or a capacity problem and saying like either my token costs are just getting out of control and or I just can't get enough tokens to do the thing I want to do. And for that, oftentimes people start turning to hyperscalers, AWS, GCP, et cetera, and doing things like provision throughput purchases, doing things like spinning up models on like a bedrock or a vortex or an Azure AI Foundry or something like that.

24:51And this kicks the can down the road a little bit because now you have like a large scale commit with a provider who has, of course, massive scale and the ability to deliver you some capacity. And then the sort of next step that I see a lot of times is going on to more of a dedicated inference provider like a Base 10 where companies have, you know, the models with the weights that they own, and they start to set up, you know, a specialized deployment. I think the sort of other fork in the road there is to go for something that's a little bit more in-house, either, you know, building an in-house platform, staffing that team up and trying to deliver a sort of best-in-class inference platform internally.

25:38Or, you know, in some cases, like going really deep into edge inference, especially if you want to do a lot of sort of field work, you replace all of the distributed systems problems with Internet of Things problems, but you have a very similar challenge in scaling inference for sort of a large distributed edge network. I see a lot less of like companies going out and making enormous capital purchases of large numbers of GPUs, sticking them in the basement of their office and doing all that stuff internally. I do see that, especially in the medical field and medical research and healthcare companies do that.

26:17I've definitely seen that in the finance world as well and some of those other traditional enterprises. I guess the alternative to that is also like going out and making a very large spend commit. You saw like Jane Street, for example, announced today with CoreWeave. That's sort of a financialization difference between, you know, going out and buying a bunch of GPUs and sticking them in your basement. You know, I think that all of these approaches are valid in the right circumstance. I certainly wouldn't say that, like, all inference in the world needs to be run through a dedicated provider. But I will also say that, like, these technologies, like we talked about earlier, are, like, very challenging to do well and to keep up with as the industry continues to scale.

27:03So it's definitely becoming like more and more of a trend for folks to buy instead of build on the inference side, which is something that I'm like very grateful for. Yeah, yeah.

27:16Sam Charrington:One of the critiques that you hear, like about the businesses that are being set up to do inference for other companies as like the, you know, AI grows is that, you know, these GPUs are a rapidly depreciating asset and the companies are financing the purchases on the public markets and financing them through debt and all these other mechanisms. I'm just wondering if you have like a strong feeling about like you know the GPU lifespan as it relates to as it relates to inference and inference maturity and the way the the field is going so I've been through now three full GPU cycles ampere hopper blackwell and and I actually like haven't really been all the way through blackwell because blackwell isn't fully rolled out yet.

Read the full transcript

28:12The thing is, the time it takes for a GPU generation to first be manufactured and actually distributed to the point where everyone can get their hands on it is months to a year. And then there's another cycle of months to a year of porting all of the code in the industry to run on it. So actually, like, proper GPUs in particular still are very, very popular for inference. One big reason for that actually is because so much open source work comes out of Chinese labs who, due to export controls, generally work on Hopper GPUs and not Blackwell GPUs. So you generally get like either, you know, FPA to int4 kernels.

28:52You get things that are built for Hopper's asynchronous programming paradigm instead of the slightly different paradigm of Blackwell kernels you get things that are models that are built for the size and restriction of say like an 8x h200 node instead of say a gb300 nbl72 system so there's a lot that still is built for blackwell oh sorry for hopper i think that like ampere has finally fallen out of favor in inference for the most part, mostly because it doesn't have support for FBA quantization. So you have to run models at full precision or suffer sort of catastrophic quality loss. And that makes Hopper GPUs like much more cost effective for inference versus Ampere.

29:45But due to both like Blackwell shortages, as well as the large number of sort of Hopper optimized workloads in the industry, like Hopper has had a ton of staying power you know an h100 is more expensive now than it was a year ago uh in terms of a rental basis and i think that like as we continue to radically underestimate the demand for inference um even though we're able to you know scale models bigger make them cheaper make them faster there's going to be a very strong sustained demand for hopper inference i think there's going to be a very very strong sustained demand for blackroll inference even as ruben rolls out and yeah the other thing about hopper is that they are really good for like smaller models too um as as we're getting bigger and bigger gpus we're relying more on like mig multi-instance gpus to split them up and right size them for small models stuff like embedding models voice in voice out models like these can be one to three billion parameter models maybe as big as like seven or eight billion in many cases and so you don't need like the massive brand new 72 blackwell gpus in a rack to run these things you need like a small slice of a hopper gpu to do them in a very efficient and cost-effective way.

31:16So long-winded answer, but I am very bullish on the sort of valuable lifespan of GPU generations, especially starting with Hopper. I do think that, again, like the older stuff, actually, I mean, Lovelace, like we still run a lot of L4 and L40 workloads. Again, a lot of times for these smaller models or like more cost-sensitive batch workloads. But yeah, before that, Ampere, you know, the T4s, like all that sort of stuff, you know, used to have like Kepler GPUs, I remember back in the day. We don't really see demand for that anymore. But starting with the Lovelace Hopper like 2018 GPUs, so that's like 2018, 2020.

32:06And then we're still using them like five, six years later. So it's got staying power. Okay.

32:14Sam Charrington:Okay. We talked a little bit about inference engineering as kind of a career direction. And you kind of alluded to the rise of like AI-assisted coding. I'm curious. Do you think that inference engineering is particularly more or less resistant to being fully automated via LLMs in Gen AI? Does the systems nature of it make it more robust or can LLM figure out how to write the CUDA and the the cube control commands and like all the the random stuff just as easily as or maybe even easier than other aspects of software engineering well there's definitely certain pieces of the stack that are getting ai accelerated right now i think one great example is CUDA kernel writing which is a a very sophisticated several months like there's a whole bunch of you know open source and startups that are like doing custom cruder kernels on demand you know exactly exactly so that's big the thing is there's first off like there's there's a lot of there's a lot of demand for more specialized inference inference systems are still very generic relative to the models you're taking a sort of specific architecture and running it with a pretty general set of kernels, a pretty general set of frameworks and engines on top of that on very, very general hardware.

34:03So there's a lot of unlocks to be made in terms of making the inference system more specific to the exact workload it's running, as well as adjusting that inference system on the fly. So I could imagine a world where the inference systems we use are 10 times more complex than the ones we have today.

34:26Sam Charrington:You made a really interesting point in the book about how inference unlocks, essentially, I'm paraphrasing here, but inference unlocks are essentially adding constraints to the way you think about inference. Yeah, yeah. So understanding the nature of the workload and, you know, giving, making, making something more specific to exactly what you want to do. It's like it's not about being able to do everything reasonably well, which is where our systems are today as an industry. It's about being able to do the exact thing you want as well as possible. And so the, you know, the first thing that LLM-assisted inference, I hope, unlocks is more of that.

35:12The ability to, you know, create much more workload-specific systems. the you know the the other pieces that are that are getting assisted by by ai programming are certainly you know the the developer experience of deploying models and influencing them um the creation of for example like if you want to create a spec deck algorithm for a model like you need a large deck algorithm sorry speculative decoding um got it yeah if you want a speculation algorithm like a like say like an eagle three model you need to create a large amount of samples you know that sort of synthetic data generation can be ai assisted there's a lot of work there's a lot of sort of grunt work in this um influence industry for for lms to take over for us and you know Eventually, I do see them owning certain systems end-to-end.

36:12But one of the things we like to say at Base Dunes League, you can't vibe code uptime. And ultimately, there is still going to need to be, for these mission-critical systems that have hundreds of millions or billions of dollars of economic value relying on them, there's going to need to be human owners who can be accountable for the results of this system. And, you know, of course, like we're going to continue to develop better and better tooling to amplify the impact of these people, but they're not going to go away entirely.

36:46Sam Charrington:We talked a little bit about this idea of like adding constraints as a way to get to higher performing systems or more efficient systems. At the same time, like, we're also moving as an industry, I think, to, you know, these multimodal models that can do everything, agents that can do everything. Like, are these ideas at tension or, you know, does it just kind of point to a particular direction for inference? I think that the demand for agents and multi-step inference is, in fact, what is driving the need for this very specialized inference optimization. In a chat world, you are making, every time a user takes an action, you are making one request to one model.

37:43In an agent world, when a user takes an action, you are making dozens, hundreds, maybe even thousands of requests, oftentimes across different models. And if you want your agent to be fast and reliable, then every single one of these models, as well as every connection between them, needs to be fast and reliable. And so that makes inference a more critical challenge. And it also means that having a system where you are able to very quickly optimize individual models and very quickly optimize individual workloads, perhaps with a sort of AI accelerated programming paradigm, means that you can actually build and scale these agents much more reliably.

38:28Sam Charrington:Talk about the, the multimodality aspect of that and the kind of disparate workload aspect of that. Like, I think I'm asking two distinct questions here. One is like, you know, we've got, you know, multimodal models, like, what is different and interesting about them from an inference perspective, if anything? The other is more trying to get at in the agentic world. You kind of spoke to calling, you know, this fan out of like requests to like different types of models. I don't know. This question is just kind of occurring to me. You also have like tools like, you know, have you thought about or to what degree are people like really focusing on like tool engineering or tool inference engineering.

39:22Sam Charrington:I don't know if this makes sense. But I'm curious, like, it's probably a separate question of like tool at scale system optimization and engineering that maybe is out of scope of inference. But, you know, if not, I want to hear your thoughts on that. I like to think that most things are in scope. Like we want to take as much responsibility for the outcomes of the system as possible. And the provisioning, sort of provisioning of tools and figuring out how to supply tools to the model, that is somewhat more of a context engineering problem. But the actual execution of those tools, you know, is an important part of an inferencing system in an agentic world.

40:15I will say that like the model modalities like you were talking about is an interesting place to focus just because there is another example here of how understanding inference helps you build better products. Let's take classification as an example of a task. if I want to classify some data, I could throw that data into Opus 4.6 with max thinking high. And I'm pretty sure I would get a really, really good classification out of that. And then I would pay a whole lot of money for it as well. And if I want to use a traditional ML-based classifier like I was playing with in early 2022, If I have the right one, maybe one that I wrote with the agent coding loop, I can also get a really, really good classification and I can get it effectively free.

41:24So understanding like the different tasks that different models are capable of means that, of course, you can save cost by moving it over to a smaller model. But you can also save time. And that's maybe even more important. So like, let's say like named entity recognition as a, you know, over the top example. No one is truly doing named entity recognition, which is sort of extracting keywords from sentences. No one is truly using frontier LLMs to do that. Or at least I sure hope they're not. But even if you're using, you know, like a flash model type of thing for that versus as a specialized model, maybe a highly optimized small LLN can do that task in 500 milliseconds.

42:13We just released a named activity recognition runtime that does it in one millisecond. One, not 500. And if you have an agent that is doing this a hundred times for user requests, all of a sudden this has gone from something where you have to look at the spinner to something that happens instantly. And so having optimized run times for every different thing you want to do, it might seem like overkill at first. But when you look at the scale that many of these sort of frontier AI companies are operating at, like it actually very much isn't. There are tons of workloads sitting out there with, you know, seven, eight figure spend behind them in the inference world that are just running on these very generic systems and need something that's very specialized.

43:05And that specialization can often look like, you know, a modality specific thing. you know a lot of a lot of models are kind of congregating around two ways of of being one being a sort of auto aggressive token look of the world and even like voice in voice out models look like that embedding models look like that the other being a diffusion look at the world which is more often seen in in video generation and image generation but also like there are diffusion llms there's they're always sneaking diffusion into all kinds of different things and there's a lot of sort of promising research there.

43:43But within each of those, you still have different components to consider with a vision language model. Vision language models can take images as input as well. They have this tiny little 1 billion parameter vision encoder in front of your trillion parameter LLM. And honestly, that little vision encoder often causes a lot more runtime problems than the LLM itself does just because the vision encoders right now, they're so heterogeneous across the industry and there's so much less software support for all their different quirks. So yeah, so that's just to say the specialization of the inference systems to the workload is still in very, very early days and there's a lot more to build across every single thing other than just take a giant LLM and run it as fast as possible.

44:36Although there's still a lot to build there too.

44:38Sam Charrington:Yeah, I thought as you were talking that you were, you know, when you can break everything down to, you know, sequential autoregressive models or diffusion models, then like, you know, all of the problems get abstracted away into these two types. But if the VLM is any example, then, you know, each of these modalities has its own quirks that you need to address. It does. I think like one place, like sometimes you do get sort of unexpected carryover benefits. One place for that actually is like text to speech models. Many modern text to speech models are basically extremely fine tuned LLMs. what you do is you extend the vocabulary of the llm the the different things it's able to represent in a token with yeah we you you send it to actually like audio waveforms you add like an encoded representation of a couple tens of thousands of audio waveforms and then you fine tune the model to to generate those a token at the time and stream them out and decode them and for that's why for those sort of models we were able to just like grab a custom version of tensor rtlm stick the tts model in it and let it rip and all of a sudden get really really excellent performance but if you try to do the same thing with again like you know vlms is an example i talked about or you know if you try and do the exact same thing with whisper for audio in which does have the encoder and the decoder component then then then then it doesn't work so you do get a free lunch here and there, but a lot of times you do have to do this very specific work.

46:18And even when you don't, even in a case like TTS where you are able to sort of reuse an existing system, there's still modality-specific considerations. With LLMs, you just want to make as many tokens as possible. You want your tokens per second to be as high as possible. With text-to-speech, you only need a certain number of tokens per second for real-time speech. Oftentimes, that's about like 80 to 100 depending on the model so if you get your system to like 100 tokens per second all of a sudden what you want to do is increase the batch size so that you can run you know more concurrent streams or you know maybe you want to leave room for like retries and checks and that kind of stuff so there's or you know figure out how to get that 100 tokens per second now on cheaper hardware there's you know a cap on on your your tps speed that you care about and that's novel.

47:10So, so yeah, so even, even in a similar architecture across modality, you might have a different performance goal and having a different performance goal completely changes the way you architect your system.

47:21Sam Charrington:So you mentioned, uh, TensorRT and that made me think about like the software side of this, like VLLM, TensorRT, like where do you think we are in the cycle? Like, Are we like in the explosion part where we're getting a lot more options? Or are we at the narrowing part? It kind of feels like a little bit of both to me sometimes. You know, the market here is remarkably concentrated in open source. If you look at, for example, JavaScript frameworks, at the height, there were dozens of very viable frameworks you could build an application on. If you look at programming languages, there are still like maybe eight languages that I think most people would say are a reasonable pick to build a website with.

48:14There are, of course, a large number of really promising open source projects across the entire AI community. I think there's a ton of really interesting work in agent orchestration and context work and that kind of stuff. Within inference, there's really been a concentration around like three major open source runtimes, the VLLM, SGLang, and TensorRT LLM. And I think part of that is just the complexity of standing up a new runtime from zero is very, very high. And so most people find it more useful to sort of contribute to an existing one. there's a lot of really good open source work around inference optimization outside of that like good open source kernel libraries open source quantization tools kv cache reuse tools new speculation stuff that kind of thing but i think it's it's just uh due to the complexity of building a complete inference engine from scratch that like i mean even even the influence providers tend to take a open source engine and build a bunch of customization on top of it versus like completely reinventing everything from scratch unless they have to due to either like being on specialized hardware that they design themselves or like you know creating their programming language or whatever the sort of practice in the industry is to take the best of open source and the best of what you can build yourself and put them together

49:50Sam Charrington:and you mentioned hardware what do you see happening on the hardware side of the equation i mean i'm team green all the way like the book cover is green for a reason but i think that there's a argument to be made for the specialization at the hardware layer as well i think one of the coolest demos I saw recently was Talos, which was that ASIC, they burned the actual LAMA 3.18B onto a chip and got to some ridiculous like 16 ,000 tokens per second number, which was really cool and sort of an interesting research project in the direction of like where this industry could go. I think that what we're seeing, you know, maybe 2026 is the year of a lot of things.

50:37One thing 2026 could be as the year of disaggregation with, for example, NVIDIA buying Glock and having more specialized pre-fill compute versus decode compute. We saw the same thing on AWS with combining their Tranium chips with Cerebris for decode. I think that there's definitely something to be said for these systems, especially for like very, very large scale workloads where you have one model that you are sending, you know, millions of concurrent requests to. And so, yeah, the increasing specialization in hardware is really just beginning and I think is going to offer a lot of unlocks for the industry, but it's certainly not going to entirely replace the need for very sophisticated software as well.

51:24Sam Charrington:You started going down this path of 2026 is the year of like, what else are you seeing? What are you excited about? What's your crystal ball saying? I mean, 2026 is the year of influence. Like 2025 was too, 2024 was, 2027 will be. But, well, it's the year of influence. And Linux on the desktop. I think that this year we're going to see a real increase in ownership of intelligence. You've seen like you saw with Shopify, them moving to a Quen model and saving millions, tens of millions on their workloads. You see companies like Coso coming out with very sophisticated models like Composer that are allowing them to really create a novel experience for the users who are depending on their platform.

52:20I am a daily user, for example. So I think that's the trend that I'm most excited about is companies understanding that if they want to build a really differentiated product, they need to be differentiated at every level. And that's starting to include the model level.

52:35Sam Charrington:And for folks that want to dig in more is, you know, where can they get the book? Is it easily accessible? Yeah. So Inference Engineering is available at Base10.com slash Inference Engineering. There's a free PDF and free e-book, e-pub. I'm working on an audiobook with actually with a customer of ours, which is very exciting. So that'll be out hopefully in a few weeks. Physical copies are a little harder to come by. I massively underestimated the demand for this book. I've already done 20 ,000 digital copies and given away nearly my entire first print run. I should have a Shopify store up next week, which will allow people to self-serve stuff.

53:16But there's always copies at any base 10 booth at a conference or trade show. What's the next conference you're going to? I'm going to AIE Miami and GTC next week. We're going to be at AI Engineer World's Fair this summer. We do a bunch of events around San Francisco and New York where we hand these out. So, yeah, it's a little bit of a scavenger hunt to get your hands on one. But inference is here and it's time to learn about it. So come through. Awesome.

53:48Sam Charrington:Awesome. Well, Philip, thanks so much for jumping on and talking with us about inference engineering. Hey, Sam, thanks so much for having me and appreciate the conversation. Same. Thank you.

From the publisher

In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today’s runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most.

The complete show notes for this episode can be found at https://twimlai.com/go/766.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
How to Engineer AI Inference Systems with Philip Kiely - #766The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 55 min
Listen in VO