In short
The episode argues that “agentic AI” will increasingly rely on small language models (SLMs) rather than large LLMs, driven by capability-per-parameter, better operational fit for agents, and major cost/energy advantages. It cites agent adoption (over half of large IT enterprises using agents; 21% adopting in the last year) and an economic mismatch (about $5.2–$5.6B agent LLM API serving vs ~$57B cloud infrastructure investment in 2024).
Key claims
SLMs under 10B parameters can match larger models on many tasks; SLMs are 10–30x cheaper to serve; PEFT fine-tuning (LoRA/Q-LoRA) can specialize models in hours; local/edge deployment improves latency and privacy; agents can route among heterogeneous models and learn via feedback loops.
Notable examples
Microsoft Phi-2 (2.7B) and Phi-3 (7B); NVIDIA “Pneumotron H” (2–9B); DeepSeek R1 Distill Qwen 7B outperforming Claude 3.5 Sonnet/GPT-4o on some benchmarks; Toolformer-style tool augmentation; NVIDIA “Dynamo” for inference scheduling.
Guests
No guest names or backgrounds are provided in the transcript.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOMarket Dynamics of AI Agents
0:45 to 2:15
Discussion on the current market valuation and rapid adoption of AI agents.
“So when you think AI agents, I guess like me, you probably think large language models, right?”
Challenges with LLMs
2:15 to 4:00
Exploring the financial discrepancies in the LLM infrastructure versus market value.
“Drawing on some really interesting new research, we're going to make the case that small language models SLMs are actually the future for agentic AI.”
Introducing Small Language Models
4:00 to 5:55
Argument for the shift from LLMs to small language models (SLMs) for agentic AI.
“There's a real need to think about the rising costs, sure, but also the environmental impact of these huge models.”
Defining Small Language Models
5:55 to 7:55
Clarification of what constitutes small language models and their operational capabilities.
“It's actually been shown to outperform big names like Claude 3.5 Sonnet and even GPT-40 on some common sense reasoning benchmarks.”
Capabilities of SLMs
7:55 to 10:15
Exploration of the effectiveness and performance of small language models compared to larger counterparts.
“That speed sounds incredibly valuable for businesses.”
Economics of Small Language Models
10:15 to 12:30
Discussion on the cost-effectiveness and operational efficiency of SLMs.
“Ah, so even if the LLM could write a sonnet, the agent only asks it to extract a date from an email.”
Flexibility and Modularity of SLMs
12:30 to 14:03
How SLMs support modular design and enhance adaptability in AI systems.
“Well, the big one, challenging V1, that SLMs are powerful enough, is the classic.”
Economic Considerations of SLMs vs LLMs
14:03 to 15:10
Explore the economic arguments for centralized LLMs and how SLMs might challenge that status quo.
“It's the argument that centralized LLM inference will still be cheaper overall due to economies of scale.”
Practical Roadblocks for SLM Adoption
15:11 to 17:23
Discuss the hurdles SLMs face in gaining traction compared to LLMs.
“Yeah, it's basically, OK, maybe SLMs could win, but LLMs just have too much momentum, too much of a head start.”
Roadmap for Transitioning to SLMs
17:24 to 20:19
Learn about the six-step algorithm for converting agents from LLMs to SLMs.
“There's growing recognition in the AI field itself about SLM utility.”
Show all 11 chapters
Future Implications of SLM Adoption
20:20 to 21:24
Reflect on how a shift towards SLMs could reshape the AI landscape.
“We've unpacked a seriously compelling case challenging the LLM dominance narrative for AI agents.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. This is where we take a whole stack of articles, research notes, and really try to cut through all the noise to get you the most important, maybe even surprising insights. And today we're definitely diving into something that's not just, you know, growing fast, it's exploding. AI agents. Yeah, exploding is the word. I mean, look at the number. It's pretty staggering. The AI agent sector valued around, what, 5.2 billion U.S. dollars late 2024. Okay. But the projection, nearly 200 billion by 2034, that's, well, that's massive growth. And it's not just pie in the sky, right?
0:35People are using them now. Oh, absolutely. Over half of large IT enterprises are already using AI agents actively. And get this, 21 % adopted them just in the last year alone. Wow, that's a really fast uptake. So when you think AI agents, I guess like me, you probably think large language models, right? The LLMs we hear so much about. That's been the standard thinking, yeah. They're seen as the core intelligence. The brains of the operation, handling the big decisions, breaking down tasks. Precisely. LLMs have been that foundational piece. And the typical setup involves agents talking to, you know, LLM API endpoint.
1:11Hosting the cloud, big centralized system. Exactly. That's become the sort of default way the industry operates. Okay, but here's where, for me anyway, it gets a bit weird. Confusing even, when you dig into the financials. So the market for actually serving those LLM APIs for agents was about$5.6 billion in 2024. Seems reasonable. OK, yeah. But the investment in the cloud infrastructure to host those models, it jumped to a massive$57 billion. Same year. $57 million. Yeah, that's like a tenfold difference. It just feels off, doesn't it? Like maybe the system we've built is, I don't know, not quite sustainable.
1:51It absolutely raises flags about the economics long term. You see this huge upfront spend, this bet on LLMs remaining dominant, assuming the ROI will eventually look like normal software. But that gap. That huge gap between the serving market value and the infra cost. It points to, well, a potential imbalance, a big one. So that brings us to what we're doing in this deep dive. We're basically here to challenge that status quo, that LLM-centric view. Right. Drawing on some really interesting new research, we're going to make the case that small language models SLMs are actually the future for agentic AI.
2:25Not the big LOMs. Not primarily, no. We think SLMs are already powerful enough, they're a better operational fit, and crucially, they're just more economical for most agent tasks. Okay, powerful enough, better fit, more economical. Let's dive in. But first, maybe we should quickly define our terms. What exactly is small here? Good point. It's not arbitrary. For this discussion, think of an SLM as a language model that can, you know, actually fit and run effectively on a standard consumer device. Like a laptop or even a powerful phone. Potentially, yeah. The key is inference latency low enough for a single user's requests.
2:59As of now, 2025, we're generally talking models under 10 billion parameters. Okay, under 10 billion. And an LLM is just anything bigger. Exactly. Anything that's not an SLM by that definition. And we'll use agent and agentic system pretty interchangeably, meaning the software itself or all its parts working together. Got it. Practicality is key. So the core idea, the main argument from our sources is that SLMs are the future here. And it rests on three pillars. That's right. Three key reasons. One, they are in principle sufficiently powerful for what agents need to do. Let's call that V1. Okay, V1.
3:32Powerful enough. Two, they are inherently more suitable operationally for how agentic systems work. That's V2. V2, better operational fit. And three, they are necessarily more economical for the vast majority of LM uses within agents. That's V3. V3. Just plain cheaper. I like that phrasing. Yeah. And our sources seem to suggest this isn't just a tech preference. They frame it almost as a moral imperative. What's that about? Yeah, it sounds strong, but it points to a bigger picture. There's a real need to think about the rising costs, sure, but also the environmental impact of these huge models.
4:07The energy consumption for training and running them. Exactly. So moving towards SLMs where appropriate isn't just efficient, it's arguably more responsible. It's about sustainable AI development, using resources wisely. Right. That makes sense. A push towards greener, more accessible AI. OK, let's dig into those arguments. Argument one, SLMs are already powerful enough. How is that possible when we keep hearing bigger is better? Well, the scaling isn't as straightforward as just more parameters equals more smarts anymore, especially for SLMs. The capability curve, how much better they get as they grow, seems to be getting much steeper for these smaller models.
4:46Steeper, meaning you get more bang for your buck, parameter-wise. Precisely. Newer SLMs are punching way above their weight, getting surprisingly close to what much bigger models could do previously. It's becoming more about capability per parameter, not just the raw count. Okay, that's a really interesting chip, but Doc is cheap, right? Do we have solid examples, like specific models showing this? Oh, definitely. Let's look at Microsoft's PHY series, PHY-2. It's tiny. 2.7 billion parameters. Yet it achieves common sense reasoning and code generation on par with models around 30 billion parameters.
5:20And it runs about 15 times faster. 15 times faster. Doing the same job as something 10 times its size. Yeah. And then there's PHY-3 small. That's 7 billion parameters. It matches 70 billion parameter models in language understanding and coding tasks. Wow. Okay, that is genuinely surprising. So Microsoft's seeing huge efficiency gains. Yeah. Anyone else? Absolutely. NVIDIA has the Pneumotron H family, models in the 2 to 9 billion range. They're getting instruction following encoding accuracy similar to 30 billion parameter LLMs, but using way fewer computational resources, FLOPs, basically. An order of magnitude fewer, you said.
5:57An order of magnitude fraction, yeah. And check this one out. DeepSeq R1 Distilquin 7B. It's a 7 billion parameter SLM. Okay. It's actually been shown to outperform big names like Claude 3.5 Sonnet and even GPT-40 on some common sense reasoning benchmarks. Outperform GPT-40, a 7 billion parameter model. On specific tasks, yes. It completely flips the script that you need those giant models for top performance. That really does challenge the conventional wisdom. And there's more. Techniques like Toolformer show that even a 6.7 billion parameter model, if it's augmented to use external tools or APIs smartly, can beat something massive like GPT-3, which is 175 billion parameters.
6:38So it's not just the model, it's how you use it, how you augment it. Exactly. It underlines that capability enhanced by design and tools is the real bottleneck, not just parameter count. SLMs clearly offer enough reasoning power for a huge chunk of agentic work. Okay, point taken. They're surprisingly capable. Let's move to V3. They're more economical. The just plain cheaper argument. How much cheaper are we talking? The difference is pretty stark. A typical 7 billion parameter SLM can be anywhere from, say, 10 to 30 times cheaper to actually run to serve than a big 70 or 175 billion parameter LMM.
7:1410 to 30 times across the board. Yeah. looking at things like latency, energy use, the raw compute needed, FLOPs, this lower cost makes real-time responses possible at a scale you just can't easily achieve with the huge models. And systems like NVIDIA Dynamo are now emerging to make SLM inference super efficient. And beyond just running the model, what about developing and updating them? Is that cheaper too? Massively. That's the fine tuning agility, using techniques called parameter efficient fine tuning or PFT things like LoRa or QLaura. Right. Heard of those? You can fine tune an SLM with just a few GPU hours.
7:49That means you can add new skills, fix problems, or specialize a model literally overnight, not weeks or months like with giant models. That speed sounds incredibly valuable for businesses. Faster iteration, lower dev costs. Huge advantage. And you mentioned running them locally. Yeah. On the edge. Yes. Edge deployment. Think systems like NVIDIA chat RTX running SLMs right on your local consumer GPU. So the agent runs entirely on my machine. Exactly. That gives you much lower latency, real-time responses, and crucially, stronger data control. Your data doesn't need to leave your device. Big plus for privacy and security.
8:25Okay, so we're talking cheaper to run, faster to update, potential for local execution. It sounds like you end up building systems differently, like with building blocks. Precisely. It encourages this modular Lego-like design. Instead of one giant monolithic LLM trying to do everything, you scale out by adding small, specialized, expert SLMs. Like having little specials for summarization, another for coding help, another for translation. Exactly. These systems end up being cheaper overall. They're faster to debug because the pieces are smaller, easier to deploy, and they just align better with the varied specific tasks agents actually do.
9:02Which leads perfectly into the flexibility argument V2 and this idea of democratization. Democratization. Right. Because training and fine-tuning SLMs is so much cheaper, it becomes practical to have lots of specialized models for different agent routines. You can adapt quickly as user needs change or new regulations come in. That adaptability sounds key. And the democratization part. Well, if it's cheaper and easier to build and deploy these models, then more people and smaller organizations can actually participate. It's not just the tech giants. Whoa. More players in the game. More players, potentially more diverse perspectives baked into the AI, less risk of bias from just a few huge models, and likely faster innovation across the board because more people are experimenting.
9:44Okay, so SLMs. Powerful enough, cheaper, more flexible, better for modular systems, potentially more democratic. So why are agentic systems in particular such a good home for them? Our sources had specific reasons for this too. They did. The first one is really insightful. Agents actually expose only a narrow slice of an LM's functionality. What do you mean by expose? Well, an agent uses careful prompting, context management. It basically forces even a generalist LLM to focus on a very specific task within a narrow range. Ah, so even if the LLM could write a sonnet, the agent only asks it to extract a date from an email.
10:21Exactly. And for that narrow, specific task, a fine-tuned SLM can often do the job just as well, maybe even better because it's specialized, but way more efficiently. The LLM's general knowledge is often just overkill. Makes sense. And this links to the need for strict behavioral alignment in agents. It does. Agents often need outputs in very specific formats, right? JSON, XML, Python code for calling tools, maybe markdown for display. Yeah, predictable structure is important for the code used in the LLM. Crucial. And you can train an SLM to always stick to one format. It's much less likely to hallucinate or randomly deviate compared to a general purpose LLM that might get creative when you don't want it to.
11:02That reliability is a big operational plus. And agents themselves aren't just one monolithic thing. They can be heterogeneous using different models. Yes, exactly. Think of the agents controlling code. It can, in theory, choose which LM to call for which task, like a smart dispatcher. So the main agent logic could use a slightly larger model for planning. Right, for the high-level stuff. But then for a quick sentiment check or a data extraction, it could route that query to a super-efficient, specialized SLM. Like having different tools in a toolbox. Perfect analogy. This modularity, this ability to mix and match LMs of different sizes and specializations, is a natural fit for SLMs.
11:41A system might have one LLM and several SLMs working together. And the final point here is that these agent interactions actually help improve the models over time, a feedback loop. Absolutely. Every time an agent calls an LM or a tool, that interaction, the prompt, the response, the tool use is valuable data. If you capture it. If you capture it, yeah. Imagine a logger within the agent. You collect this data, clean it up, filter it, and then you can use it to further fine-tune your specialized SLMs. Making them even better and more efficient for those specific tasks. Exactly. It creates this cycle of continuous improvement, making the whole system more effective and cost-efficient over time.
12:20Okay, the case for SLMs and agents seems really strong, but no idea is without its critics. What are the main counter-arguments people pushing back against this? Well, the big one, challenging V1, that SLMs are powerful enough, is the classic. LLMs will always have superior general language understanding. The bigger is fundamentally better argument. Right. They point to spaling laws, empirical evidence showing bigger models generally do better, and this idea of a semantic hub in LLM. A semantic hub, like a central understanding core. Sort of, yeah. Hypothesized ability in large models to integrate and generalize meaning across vast domains better than smaller models can.
13:00Okay, so size grants better generalization. How do our sources rebut that? They hit it from a few angles. First, scaling law studies often assume the architecture of the model stays the same, just bigger. But recent SLM research shows big gains from architectural innovations, which those laws don't really capture. So smarter design, not just bigger size. Exactly. Second, the ease of fine-tuning SLMs lets you achieve high reliability on specific tasks, which general scaling laws don't account for. Third, SLMs are cheaper to run, which means you can afford more computation at inference time, like having the SLM think longer or check its work enhancing its reasoning.
13:37More thinking time for the SLM. Right. And finally, remember how agents break problems down. Good agent design decomposes complex tasks into simple subtasks. For those simple steps, you often don't need the super broad abstract understanding of an LLM semantic hub. An SLM focused on the subtask is sufficient. Okay, that decomposition argument feels powerful. Address the specific small task, don't worry about the whole universe. What's another counter? Another one targets the economics, V3. It's the argument that centralized LLM inference will still be cheaper overall due to economies of scale.
14:14Ah, the big cloud provider argument. Cheaper to run one massive data center efficiently than lots of little ones. That's the gist. Plus, they argue it's harder to keep lots of specialized SLM endpoints fully utilized, and managing distributed infrastructure and talent is expensive. That sounds plausible, at least. What's the countercounter? We have to acknowledge it's a valid point, and the exact math depends heavily on the specific use case. It's complex. However, the sources highlight big improvements in inference scheduling, efficiently managing requests across many models and systems like NVIDIA Dynamo, are specifically designed to tackle these load balancing challenges for distributed systems.
14:53So technology is catching up to make distributed systems more efficient. It seems so. And importantly, the actual costs for setting up infrastructure are generally trending downwards thanks to tech advances. So while it's a real consideration, the advantage of centralization might be shrinking. OK, interesting. And the last pushback, kind of a pragmatic one. Yeah, it's basically, OK, maybe SLMs could win, but LLMs just have too much momentum, too much of a head start. Industry inertia. The path of least resistance is sticking with the big models everyone's already invested in. Exactly. And we have to agree that's a real possibility.
15:29Inertia is powerful. But we'd argue the advantages we've laid out, the power, the economics, the flexibility, the operational fit, they're so significant that they provide a very plausible pathway to overcoming that inertia. The benefits might just become too compelling to ignore. So if SLMs are so great, why isn't everyone using them for agents already? What are the actual practical roadblocks holding things back? It really boils down to practical hurdles, not fundamental flaws. The biggest one is probably the massive existing investment in centralized LLM infrastructure. All that money already spent on building systems for LLMs.
16:04Right. So naturally, the tools, the workflows, the industry focus has been on optimizing that paradigm. Decentralized SLM approaches have been relatively overlooked. It's like, we built a superhighway, so we're optimizing for semi-trucks, even if delivery vans on local roads might be better for some things. That's a perfect analogy. The infrastructure dictates the focus. Second, there's the issue of benchmarks. How we measure if a model is any good. Exactly. SLM development often still follows LLM trends, aiming for high scores on generalist benchmarks, like understanding literature or broad knowledge.
16:41But that's not necessarily what an agent needs. Often not. Research shows that when you evaluate SLMs on agentic utility benchmarks, how well they actually perform specific agent tasks like tool use or following instructions, they frequently outperform larger models. They're being judged by the wrong metrics sometimes. Right. Measuring apples by orange standards. Oh. And the third barrier. Just a general lack of popular awareness, maybe. LLMs get all the headlines, all the big marketing pushes. The shiny new objects. Yeah. SLMs, even though they might be perfect for many business or industrial scenarios, just don't get the same level of press.
17:16They're sort of the quiet workhorses. But you sound optimistic these can be overcome. I am. Tech advancements, like those inference scheduling systems, are chipping away at the infrastructure inertia. There's growing recognition in the AI field itself about SLM utility. And honestly, the economic benefits are eventually going to force businesses to pay attention. It just makes financial sense in so many cases. Okay, so this isn't just theory. Our sources actually laid out a practical roadmap, didn't they? A kind of algorithm for converting an agent from using LOMs to SLMs. They did. It's a concrete six-step process.
17:52Step one is crucial. Secure usage data collection. Logging everything the agent does internally. Pretty much. Every time the agent calls an LM or a tool, capture the prompt, the response, what tool was called, how long it took. But critically, do it securely and anonymize the data. Think back to that logger function we mentioned. Right. Capture the raw materials securely. Step two. Data curation and filtering. You gather enough data, maybe 10K to 100K examples for fine-tuning an SLM, and then you meticulously clean it. Remove any sensitive stuff, PII, PHI, using automated tools or careful paraphrasing.
18:27Get clean, safe data. Makes sense. Then you need to figure out what patterns are in that data. Exactly. Step three. Task clustering. You use unsupervised machine learning, clustering techniques on all those prompts and agent actions. The goal is to automatically find recurring types of tasks the agent is performing. Like 20 % of the time it's summarizing emails or 15 % it's extracting product codes. Precisely. You identify these natural clusters, defining specialized tasks like intent recognition, data extraction, summarization, maybe specific kinds of code generation. So you've identified the jobs SLMs could do.
19:02Then you pick the right SLM for each job. Correct. Step four is SLM selection. For each task cluster you found, you pick the best candidate SLMs. You look at their capabilities, benchmark scores, the relevant ones, licensing, how easy they are to deploy, their resource footprint. Find the right tool for the job. And then comes the training, the specialization. Step five. Step five. Specialized SLM fine-tuning. This is where you take the data specific to each task cluster and fine-tune your chosen SLMs on it. Using those PEFT techniques like LoRa. Yes, things like LoRa or QLaura make it really efficient in terms of cost and memory.
19:37You could also use knowledge distillation here. Knowledge distillation. That's where you train the SLM to mimic the outputs of a larger, more capable LLM on that specific task. So the SLM learns the nuances without needing the LLM's full size. Clever. Transferring the smarts. And it's not a one-and-done process. Definitely not. Step six is iteration and refinement. This is a continuous loop. You periodically retrain the specialized SLMs and maybe any router model that directs traffic using new data collected from the ongoing agent operations. So the system keeps learning and adapting. Exactly.
20:12It adapts to changing usage patterns, gets more efficient, stays performant. It's about continuous optimization. Wow. Okay, this has been, well, a really deep dive. We've unpacked a seriously compelling case challenging the LLM dominance narrative for AI agents. Yeah, we've seen SLMs aren't just small. They're surprisingly powerful. They're way more economical. They offer incredible flexibility. And they just seem inherently suited to how agents actually work those narrow tasks. The need for structure, the modularity. It really paints a picture of a future with significant cost savings, much more sustainable AI practices, and maybe a more democratized landscape where more people can build and innovate.
20:53It boils down to making powerful AI more accessible, more efficient, and maybe even more tailored to specific needs. So the final thought for you, our listener, to maybe chew on. If this shift from giant LLMs to specialized SLMs really takes hold, think about the ripple effects. How could that fundamentally change the competitive dynamics in the AI world? And maybe more excitingly, what entirely new kinds of custom, hyper-specialized AI solutions could this unlock across all sorts of industries? Something to ponder. Thanks for joining us on The Deep Dive.
From the publisher
This research paper proposes that small language models (SLMs) are the future of agentic AI, challenging the current reliance on large language models (LLMs). The authors argue that SLMs are sufficiently powerful, more operationally suitable, and more economical for the repetitive, specialized tasks common in AI agents. While acknowledging the current dominance and investment in LLMs, the paper provides an algorithm for converting LLM-centric agents to SLM-first architectures, highlighting the significant economic and operational benefits of this shift. It also addresses counterarguments regarding LLM general understanding and the economics of centralized inference, advocating for heterogeneous agentic systems where SLMs handle most tasks, with LLMs used sparingly for complex reasoning.




