Small Language Models are the Future of Agentic AI

7 Oct 2025 · 19 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that agentic AI should shift from relying on centralized large language models (LLMs) to an “SLM-first” architecture using small language models (under ~10B parameters) for most agent subtasks. It frames the current setup as an unsustainable infrastructure paradox: agents depend on big cloud LLMs (citing ~$57B capital investment) even as agent markets grow (projected ~$200B in a decade).

Guest backgrounds

No guests are mentioned in the transcript; it’s a solo “Deep Dive” discussion based on an NVIDIA Research position paper.

Key claims

SLMs are 10–30x cheaper and faster for inference; fine-tuning specialized SLMs takes “a few GPU hours” vs “weeks” for large models. Agents need behavioral alignment (strict JSON/YAML/XML outputs) more than broad conversational ability, so small fine-tuned models can be more reliable. Use heterogeneous systems: a big model can handle planning while SLMs handle parsing, extraction, and tool calls.

Notable examples

Microsoft 2.7B model (common-sense reasoning/code comparable to 30B; ~15x faster); NVIDIA Hamba 1.5B (outperforms 13B on instruction following); DeepSeek R1 (7B) beating Claude 3.5/Sonnet/GPT-4o on some reasoning; Salesforce XLM-2AB for tool calling surpassing GPT-4o; Toolformer (6.7B) outperforming GPT-3 (175B) on tool use. Case estimates: MetaGPT could shift ~60% of LLM queries to specialized SLMs; “Cradle” computer-control agent ~70%. Roadblocks: inertia, wrong evaluation benchmarks, and lack of awareness; proposed roadmap is data logging, task clustering, and specialized fine-tuning (e.g., LoRa/knowledge distillation), with continuous retraining.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Agentic Systems and Market Growth

0:45 to 2:30

Exploration of the rapid growth of AI agents and the paradox of reliance on large models.

“And these are almost entirely served by centralized cloud infrastructure.”

Defining Small vs. Large Language Models

2:30 to 5:30

Clarification of what constitutes small language models and their emerging role.

“and also look at, well, what's stopping everyone from just switching over right now?”

Economic Advantages of Small Language Models

5:30 to 9:30

Discussion on cost efficiency and operational advantages of SLMs over LLMs.

“We definitely need to challenge that assumption that bigger size is always necessary.”

Capabilities of Small Language Models

9:30 to 11:30

Evidence and examples demonstrating the effectiveness of SLMs in various tasks.

“Maximize capability while slashing the overall cost, radically minimizing it potentially.”

Counterarguments and Barriers to Adoption

11:30 to 14:00

Analysis of the skepticism around SLMs and the barriers preventing their widespread adoption.

“First, that load balancing dozens, maybe hundreds of different specialized SLM endpoints is just too complex and inefficient compared to hitting one big endpoint.”

Introduction to SLMs and Industry Shift

14:00 to 14:35

Learn about the current landscape of small language models and the industry's need for a shift.

“SLMs simply don't get the same marketing buzz or intense press coverage as the big proprietary LLMs.”

Phases of Transitioning to SLMs

14:35 to 16:18

Explore the three critical phases for transitioning from LLMs to SLMs.

“There are basically three critical phases outlined.”

Impacts of SLM Adoption

16:18 to 17:40

Understand the potential impacts and cost savings from adopting specialized SLMs.

“What kind of impacts can this actually have?”

The Future of AI with SLMs

17:40 to 18:59

Discuss the implications of an SLM-first approach for responsible and sustainable AI.

“And if we connect this to the bigger picture, this move towards SLM first isn't just about, you know, optimizing an IT budget.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. We are jumping straight into the engine room of the AI revolution today. We're focusing specifically on agentic systems. Right. If you've been watching the field, you know that AI agents, these are bits of software designed to plan, reason, use tools, all pretty autonomously. They're really expanding fast. Oh, absolutely. The market for these agents, I mean, it was valued at something like$5.2 billion just last year. Wow. And the projections, they're talking potentially$200 billion within the next decade. It's exponential. It really is. But what's crucial to grasp here is the massive infrastructure paradox that's kind of powering it all right now.

0:37Paradox. How so? Well, these agents, they currently rely really heavily on the biggest, most generalist models available, large language models, LLMs. Okay. And these are almost entirely served by centralized cloud infrastructure. We're talking an estimated$57 billion capital investment backing that whole centralized approach. 57 billion. And that dependency, that's exactly where things start to look a bit misaligned, isn't it? Our source material today, it's a key position paper, actually, from NVIDIA Research. It suggests this dependency is, well, fundamentally unsustainable for where agents are headed.

1:16Right. So the central position of this deep dive is this. Small language models, SLMs, are really poised to become the standard architecture for agentic AI. And there are solid reasons why. Exactly. Why? Because they offer the necessary power. They're way more economical. And operationally, they're just better suited for specialized tasks. Before we really unpack those three arguments, power, economy, operations, we should probably quickly nail down the definitions. Because small, well, it's relative, isn't it? Definitely need to do that. So when the researchers talk about an SLM small language model, they're defining it mainly by its utility.

1:51Can it do the job locally? Is it compact enough to fit onto a common consumer device or maybe small local server? And can it run inference fast enough low latency to be practical for just one user? Makes sense. As of, say, 2025, that generally means models under about 10 billion parameters. Okay, under 10 billion. And then for this conversation, an LLM, large language model, is just anything bigger than that. Pretty much. Anything that doesn't fit that SLM criteria. It's the big centralized brain that currently does most of the heavy lifting. Got it. So our mission today is to dive into these arguments for shifting towards an SLM-first setup and also look at, well, what's stopping everyone from just switching over right now?

2:34The Ertls, yeah. Let's kick off with maybe the most persuasive point. The sheer economics money talks. It really does. And the economy argument here is, frankly, overwhelming. It's probably the main driver for this potential shift. How big a difference are we talking? OK, think about the resources needed. Serving a 7 billion parameter SLM. It's estimated to be somewhere between 10 and 30 times cheaper. 10 to 30 times. Yeah, across the board. Latency, energy use, even the raw computation, the FLOPs. Wow. Now, multiply those savings by the billions, maybe trillions of agent calls we expect in the next few years.

3:12Centralized LLM inference just starts to look completely cost prohibitive at scale. And it's not just the cost per answer, is it? It's also about how quickly you can develop and change things. Agility. Exactly. And that connects straight to the operational advantages of SLMs. Okay. The sources really highlight this. Fine tuning an SLM, maybe using techniques like LoRa or Q LoRa. It takes just a few GPU hours, literally. A few hours. Compared to what for the big ones? Weeks. Weeks of dedicated, expensive resources if you want to re-specialize one of the giants. So with SLMs, you can add new features, specialize an agent for a niche task, maybe adapt to a new local regulation basically overnight.

3:54That's kind of rapid iteration. That's just impossible with the current big model setup. Pretty much. and that cost efficiency, that speed, it directly enables operational flexibility. How so? Well, because it's so fast and cheap to train and adapt these smaller models, it suddenly becomes really practical to deploy a whole system of multiple specialized expert models. Ah, so instead of one giant model trying to do everything. Right. Instead of that one giant Swiss Army knife, the monolithic LLM, the paper, suggests we should think more like a specialized toolbox. Yeah. You've got multiple SLMs, each one perfectly tuned for its specific job within the agent.

4:30That makes a lot of sense, like different tools for different tasks. Precise. And this structural change, it has a really interesting side effect, doesn't it? Democratization. Yeah, that's a fantastic byproduct. Yeah. When developing and maintaining these agents becomes cheaper, the barrier to entry just plummets. Well, more people can build them. Exactly. More organizations, independent developers, they can start building their own LMs. That obviously spurs innovation, but maybe more importantly, it allows for a much more diverse range of perspectives to actually get embedded into these agents.

5:01Which could help reduce some of the systemic bias issues we see with centralized models. Potentially, yes. Avoiding that centralization could definitely help foster more diversity. Okay, but this brings us right to the big question, the main point of skepticism, I imagine. Power. Mm-hmm. The capability question. For years, the message has been bigger is better. More parameters mean smarter AI. Are these small models really powerful enough, or are they just, you know, good at narrow tests? That is the absolute key question. We definitely need to challenge that assumption that bigger size is always necessary.

5:35So what's the evidence? Well, the research provides some pretty specific examples showing SLM capabilities have advanced dramatically. The scaling curve, how capability increases with size, it seems to be getting much steeper at the smaller end. Meaning small models are punching above their weight now. You could say that. They're hitting performance levels that used to require models maybe 10 times their size. Okay, let's hear some examples. The aha facts, as the paper calls them. Right, the aha facts. So Microsoft's fight to. It's only 2.7 billion parameters. Tiny by today's standards. Tiny, but it achieves common sense reasoning and code generation scores comparable to older 30 billion parameter models.

6:13And it runs about 15 times faster. 15 times faster. Okay, what else? It's not just Microsoft. NVIDIA has Himba 1.5b, again, very small. It actually outperforms 13 billion models on instruction following tasks. Instruction following is crucial for agents. Absolutely. Then there's DeepSeek R1 to still Quinn 7b. That one showed really strong common sense reasoning, sometimes even beating big proprietary models like Claude 3.5, Sonnet, and yes, even GPT-40 on certain benchmarks. Beating GPT-40, seriously. On specific reasoning tasks, yes. And maybe most critical for agents which constantly need to interact with other software.

6:52Tool use. Exactly. Salesforce has a model, XLM-2AB. It hits state-of-the-art performance on tool calling. Actually surpassed GPT-40 there too. Okay, that's compelling. I remember Toolformer. It was a 6.7 billion parameter model. It outperformed the massive 175 billion parameter GPT-3 back in the day simply because it was fine-tuned specifically for effective tool use. So augmentation, like specialized training, can make a smaller model outperform a larger generalist one. Precisely. Augmentation works and it reduces the need for just raw brute force scale. So why? Why is a small model suddenly good enough for these agent tasks?

7:27It boils down to how agents actually use these language models. Okay. Agentic applications, they deliberately constrain the big generalist LLMs. They use really rigid prompts, focusing them on narrow, often repetitive tasks. So the agent isn't asking for a poem one minute and complex code the next. Not usually within the same step, no. It doesn't need the LLM's entire library of knowledge or its conversational skills for most subtasks. What it really needs is extremely close behavioral alignment. Behavioral alignment, meaning? That's the key insight the paper stresses. It means the output has to conform to a very strict, predictable format.

8:06Like it must return this specific type of JSON or YAML or XML every single time. Why? So the code can read it. Exactly. So the surrounding code can parse it reliably and the agent's workflow doesn't break. Ah, and a big generalist LLM, ironically, might actually mess that up sometimes. That's the issue. A generalist LLM is more prone to, well, hallucinating slightly. or just deciding to respond in a slightly different format than expected. That breaks the automation. Whereas a small model, fine-tuned just for that one output format, is much more reliable. It's trained explicitly to output that format and only that format.

8:44It's predictable. It's robust for that specific task. Okay, that makes a lot of sense. Predictability over sheer power for automated tasks. Which leads perfectly into this idea of heterogeneous systems. Meaning mixing and matching models. Exactly. Agent systems are naturally modular. You can have LMs calling other LMs, so you don't have to throw out the big LOM entirely. How would that work? You could have the centralized LLM handle the really high-level stuff, the overall planning, maybe complex reasoning about the user's ultimate goal, the root agency. Okay. But then it delegates all the subordinate repetitive tasks like parsing a specific user request, extracting details, or calling a particular tool to these highly specialized, super cost-effective SLMs.

9:27So you get the best of both worlds. High-level smarts from the LLM, efficient execution from the SLMs. That's the idea. Maximize capability while slashing the overall cost, radically minimizing it potentially. Right. Now, we have to address the counter-arguments. The folks backing the big LLMs, they have some strong points based on, you know, years of research into scaling. Absolutely. We need to look at the dominant beliefs. The main one is probably this idea that LLMs will always have a fundamental edge because of their generalized language understanding. A scaling laws argument. Right. Scaling laws support it.

10:03And there's this hypothesized mechanism, the semantic hub. Semantic hub. The theory is that large models integrate information across hugely diverse domains much better than small models can. This lets them handle unexpected situations or novel complexities more effectively. And the argument is you just can't distill that broad understanding into a tiny model without losing something critical. Exactly. That you'll lose crucial performance on the edge cases or the really novel tasks. Okay, that sounds like a reasonable concern, especially for things we haven't seen before. What's the rebuttal? The rebuttal presented in the paper is that while semantic hub might be real, it's actually less relevant for how advanced agents work.

10:42How so? Because good agent design is all about decomposing complex problems into simpler, manageable subtasks before you even call a language model. Ah, so by the time the model gets the request, the complexity has already been broken down. Precisely. The task handed to the LM, whether it's large or small, is usually simple, tightly scoped, and often highly repeatable. At that execution layer, that amazing generalized understanding becomes, well, a bit of a high-cost luxury you might not need. Interesting. Okay, what's the second main counterargument against SLM? This one's purely economic, focused on the operations side.

11:18It's the economy of scale argument for centralization. Meaning the big cloud providers running huge LLMs are just so efficient, it outweighs the SLM cost savings. Kind of. The argument has two parts. First, that load balancing dozens, maybe hundreds of different specialized SLM endpoints is just too complex and inefficient compared to hitting one big endpoint. Managing all those little models sounds tricky, yeah. And second, the setup costs. You need specialized talent. You need potentially complex infrastructure management. The argument is these initial and ongoing costs might actually eat up any savings you get from cheaper inference.

11:55Centralization is just simpler to manage. That feels like a very practical, real-world challenge. Hasn't that historically been true? It has been a significant hurdle, yes. But the sources counter that modern advances are actually solving these operational problems pretty quickly. Like what? Things like breakthroughs in inference scheduling, better system modularization tools. They say systems like NVIDIA Dynamo as examples. These are specifically designed to tackle the load balancing issues that used to make heterogeneous systems difficult. So the tech for managing multiple small models is catching up.

12:29It seems so. And analyses also show that the setup, costs-finding talent, managing the infrastructure, those are consistently falling as the tooling gets better and more people learn the skills. The tide seems to be turning against that centralized cost argument. Okay. So if the arguments for SLMs are stacking up economics, operations, even capability for many tasks, why hasn't this shift just happened already? What are the roadblocks? Yeah, good question. The researchers point to three main barriers. They're mostly about inertia and maybe focusing on the wrong things. Inertia like existing investments.

13:04Exactly. Barrier number one is just massive inertia. We mentioned the tens of billions already poured into centralized LLM infrastructure. Right, the$57 billion. No major company wants to just walk away from that kind of investment overnight, even if a cheaper path is emerging. That momentum is a powerful force. Understandable. What's barrier two? The second one is about how SLMs are currently trained and evaluated. The focus in development is still often on generalist benchmarks. Trying to make them chat like big models. Kind of, yeah. Trying to compete on conversational ability or broad academic tests instead of focusing laser sharp on metrics that actually matter for agentic utility.

13:42Like? Like reliable tool calling, precise format following, speed on specific tasks. The market, you could argue, is still measuring the wrong things for this particular application. Okay, measuring general smarts instead of job-specific skills. And the third barrier. Just sheer lack of awareness, really. SLMs simply don't get the same marketing buzz or intense press coverage as the big proprietary LLMs. The big names get all the headlines. Pretty much. But notice these are all practical hurdles, right? They're not fundamental flaws in the SLO approach itself. Right, which suggests they can be overcome.

14:16So the industry needs a plan, a blueprint for making the switch. Recognizing that, yeah, the sources actually lay out a pretty practical step-by-step conversion roadmap. It's for organizations that are ready to move their agents away from relying solely on expensive LLMs towards a cheaper SLM-first design. Okay, let's hear it. What are the steps? There are basically three critical phases outlined. Phase one is all about secure usage data collection. Okay, data first. What kind of data? You need to instrument your agent system to log everything that happens when the agent calls out to a language model, assuming it's not direct human interaction.

14:54So log the exact prompts sent to the LLM, the responses that come back, any tool calls involved. Basically create a detailed record of how your agent is actually using the big LLM in the real world. Exactly. Build a high quality task specific data set from production usage. Makes sense. Phase two. Phase two is task clustering. You take all that log data and you use unsupervised machine learning techniques to find patterns. Patterns like common types of requests. Precisely. You look for recurring types of interactions. Maybe you find that, say, 30 % of all the LL on calls your agent makes are always about recognizing the user's intent from their initial query.

15:30Okay. And maybe another 20 % are always summarizing, I don't know, quarterly financial reports in a specific way. You identify these recurring, relatively simple tasks. And those become the targets for specialized SLM. They're your prime candidates, yes. Which leads to phase three, specialized SLM fine-tuning. Taking an existing SLM and teaching it the specific task. Right. You use that clean cluster data from phase two, and you fine-tune your chosen base SLM. You might use techniques like LoRa or maybe knowledge distillation. Knowledge distillation, like teaching the SLM to copy the big LLM. Essentially, yes.

16:08Training the SLM to mimic the high-quality, reliable outputs that the generalist LLM produced for those specific clustered tasks identified in Phase 2. Okay, so collect data, find the repetitive tasks, train small models to do just those tasks. Seems logical. What kind of impacts can this actually have? Do they give examples? They do provide some case study estimates. For instance, they looked at MetaGPT, which is a pretty complex agent framework for software development. Their analysis suggested it could probably handle about 60 % of its current LLM queries using specialized SLMs instead. 60 % replacement.

16:42That's substantial. It is. And for agents that are maybe more focused on repetitive actions, like one called Cradle, which does general computer control tasks, the viability estimate jumped even higher, maybe around 70%. 70%. That translates directly into potentially huge cost savings. Massive cost savings. And potentially faster response times, too, for those tasks. So let's try and wrap this up. What's the core takeaway from this deep dive? I think the main message is that this potential shift towards an SLM-first architecture for agents, it's not really driven by, you know, chasing the latest tech trend or hype.

17:15It's not just novelty. No, it seems driven by pure pragmatism. The argument is that SLMs now offer sufficient power for a large chunk of agentic tasks. And they combine that with just dramatic economic and operational advantages, especially for those repetitive, narrowly scoped tasks that seem to make up the bulk of what agents actually do moment to moment. Sufficient power, much cheaper, much faster to adapt. That's the combo. That's the core argument. And if we connect this to the bigger picture, this move towards SLM first isn't just about, you know, optimizing an IT budget. There's more to it.

17:50The paper suggests it's maybe an imperative for promoting more responsible AI deployment and critically sustainable AI deployment. Sustainable because of the lower energy costs. Exactly. In an age where infrastructure costs and the huge energy footprint of AI are becoming major concerns, both commercially and societally, this shift offers a path towards more sustainable growth. That is a powerful final thought. Responsible and sustainable AI. Okay, and to leave our listeners with something to chew on. Ah, yes, the provocative thought. The source material mentions that this conversion process isn't a one-off.

18:24It includes a step for continuous iteration. Step S6, yeah. Retraining the SLMs. Where those specialized SLMs that you've deployed are periodically retrained using new usage data logged from your system. So if agents become these expert systems that are constantly refining themselves based on your specific logged workflow data. Hmm, a feedback loop. What does that closed loop of self-improvement mean for the future of truly localized, personalized, customized intelligence? An agent that learns your specific ways of working better and better over time. That's a fascinating question to ponder. What does that level of personalization look like?

19:01We'll leave you to think on that. Until next time, keep digging deeper.

From the publisher

This paper presents a strong **position statement** arguing that **Small Language Models (SLMs)** are the **future of agentic AI**, despite the current dominance of **Large Language Models (LLMs)**. The authors contend that SLMs are **sufficiently powerful**, **more economical**, and **operationally more suitable** for the specialized and repetitive tasks common in AI agents. They provide **arguments grounded in modern SLM capabilities** and **inference efficiency**, advocating for a shift to **SLM-first architectures** or **heterogeneous systems** that use LLMs only when necessary. Furthermore, the paper outlines a **conversion algorithm** to help developers migrate existing LLM-based agents to more efficient SLM solutions and discusses **barriers to adoption** such as industry inertia and infrastructure investment.

More from Best AI papers explained

All 475 episodes
Small Language Models are the Future of Agentic AIBest AI papers explained · 19 min
Listen in VO