In short
How reinforcement learning (RL) can train LLM-based agents for multi-step, tool-using problem solving, and why training recipe choices beat raw model scale.
Guest backgrounds
No guests mentioned in the transcript.
Key claims
A 4B model (DemiAgent 4B) can match or beat much larger (14B/32B) tool-using agents by optimizing three pillars: real end-to-end multi-turn trajectories, model-aware diverse data, and an RL algorithm (GRPOTCR) with clip-hire and overlong reward shaping; plus deliberative reasoning mode.
Notable examples
Switching from “stitch-style” synthetic tool traces to real trajectories yields 40%+ gains on hard tests (e.g., AM 2024). Diverse math/science/coding data improves training stability and reaches accuracy ~40% faster. Deliberative agents achieve 70%+ tool-call success, while reactive agents call tools rapidly with low effectiveness. Long-reasoning instruction models can avoid tool use during agent training.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Shift to Agentic Models
0:45 to 1:53
Exploration of how LLMs are evolving into problem-solving agents.
“This deep dive is digging into a huge study that tried to cut through all that confusion.”
Surprising Findings on Model Size
1:53 to 2:18
Smaller models can outperform larger ones when optimized correctly.
“It feels like we should focus less on just building bigger brains and more on, I guess, building better teachers for them.”
Data Curation Techniques
2:18 to 3:25
Discussion on the importance of real data for training models effectively.
“Ah, like patching it together after the fact.”
Diversity in Data
3:25 to 5:25
How diverse training data enhances policy entropy and model adaptability.
“But this study also mentioned diversity in that data, right?”
Model-Aware Data Selection
5:25 to 7:05
Strategies to curate data for specific model capabilities to enhance learning.
“What changes did they find there that made such a dramatic difference?”
Algorithmic Enhancements
7:05 to 8:43
Introduction to effective algorithms and techniques that improve training.
“That's like a 75 % reduction in training time.”
Exploration vs. Exploitation
8:43 to 9:00
The trade-off between exploration and exploitation is altered by external tools.
“The main bottleneck for training efficiency then becomes the gap between those two.”
Internal Reasoning Modes
9:00 to 12:10
Exploration of different reasoning strategies agents use during problem-solving.
“So the goal isn't just balancing exploration and exploitation anymore.”
Quality Over Quantity in Tool Use
12:10 to 12:39
Deliberative strategies lead to better tool usage outcomes in agents.
“They were caught between two opposing goals.”
Future Directions for AI Development
12:39 to 14:00
Insights into how AI development may shift towards strategic planning and tool orchestration.
“Okay, let's try and bring all this together then.”
Show all 11 chapters
The Future of AI Development
14:00 to 15:14
Explore how careful techniques and optimization impact AI tool use and development.
“It really validates that careful technique, not just raw model scale, is driving the frontier of intelligent tool use right now.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. If you've been following AI, you know this story isn't just about large language models generating convincing paragraphs anymore. The real cutting edge now. It's LLMs becoming actual agents, problem solvers that can, you know, use external tools like code interpreters. To tackle really complex challenges, math, science, coding, things that go way beyond just text. Exactly. But turning a standard LLM into one of these brilliant tool using agents, it's surprisingly tricky. It takes some pretty sophisticated training, often using reinforcement learning, RL, to get them to learn how to handle multi-step tasks, use tools, get feedback.
0:38Right. And there's a lot of debate about the best way to actually do that RL part effectively for agents. OK, let's unpack this. This deep dive is digging into a huge study that tried to cut through all that confusion. We're looking at the sort of three key pillars for getting this right. The data, the algorithm and what they call the reasoning mode, finding the practical steps that really work. And we probably need to start with the big surprise finding because it really challenges that whole bigger is always better idea in AI, this study. They found that by really optimizing these three areas, data, algorithm, reasoning, they could train a much smaller model, just 4 billion parameters.
1:18Wait, hold on. A 4 billion parameter model? Yeah, a 4B model and get it to achieve superior agentic reasoning performance. You're saying a model that's like a quarter, maybe even an eighth the size of some of the big ones out there is actually better at using tools. That's what they found. Technique really trumps sheer scale here. The specific recipe they developed let these, you know, compact 4B models match or even beat giants like 32 billion parameter models on tough reasoning tests. Wow. It really proves that optimizing the learning process is key, maybe more key than just scaling up the model size right now.
1:53That's fascinating. It feels like we should focus less on just building bigger brains and more on, I guess, building better teachers for them. So let's jump into that first pillar then, data curation. What kind of data do you need to even start building these agents? Well, the study points out a pretty big flaw in how a lot of this data gets made currently. Many training setups use what they call stitch style or synthetic data. Basically, they take some internal thinking steps from a model and then just kind of artificially glue tool outputs into the chat history. Ah, like patching it together after the fact.
2:26Exactly. It looks like an interaction, but it's fundamentally missing the crucial bits. The agent never actually learns when it should use the tool or why now is the right moment or even what to do next after the tool gives a result. It lacks the decision-making context. Precisely. And they found that models started on this synthetic data were, well, pretty unstable. The performance wasn't reliable. Okay, so if stitching is the wrong way, what's the right way? What's the fix? The fix is using real end-to-end multi-turn trajectories, actual conversations, if you like, where the agent reasons, calls a tool, gets real feedback, and then decides the next step.
3:01The whole loop. And the difference is huge. When they switched the QIN3-4B model from synthetic initialization to this real trajectory data, its performance jumped massively. Massively. Over 40 % improvement on a key metric pass at 32 on hard tests like AM 2024. It's just real data seems indispensable for teaching the actual behavior of tool use. Okay, so real data, not faked interactions. But this study also mentioned diversity in that data, right? Mixing math, science, coding tasks. Why is that mix so important for the RL training itself? Yeah, that's another key finding. It really boils down to something called policy entropy.
3:40Okay. Think of policy entropy as like the agent's flexibility in problem solving. It's range of strategies. It's exploration breadth. If you only feed it math problems, its strategy set becomes very narrow, very deterministic, just focused on math tricks. Right, it gets tunnel vision. Exactly. But if you give it diverse tasks, math, science, code, you force its policy to stay broader, more adaptable, like making a student study multiple subjects, not just cramming for one test. And this higher entropy, this flexibility, leads to faster and more stable training. They found agents reached high accuracy levels about 40 % faster compared to training on just a narrow math-only data set.
4:17That makes a lot of sense. Keep the learning broad. Now, what about those cases where you have a weaker base model? Sometimes a smaller model seems to hit a performance wall, right, even with good diverse data. That's a really important point. And the solution they found is what they call model-aware data selection. See, if you give a model problems, it either gets right every single time, 100 % accuracy, or fails at every single time, 0 % accuracy. Those problems are actually useless for learning. Because there's no useful feedback signal. Exactly. There's no gradient, no signal telling the model how to improve.
4:51It's either perfect or impossible. So the trick is to curate the data set for that specific model. You actively toss out the problems that are clearly too easy or way too hard for that model's current ability. So you focus the training effort where it counts. Precisely. You focus the training time only on the problems where the agent has, you know, a non-trivial chance of making progress. This amplifies the critical learning signals and is especially crucial for smaller models to help them overcome stagnation. Okay, so we've got the foundation. Real, diverse, model-aware data. Now, let's talk about tuning the engine.
5:27The algorithms. What changes did they find there that made such a dramatic difference? Right. The second perspective. They looked at GRPO-based algorithms, it's a type of RL optimization, and found a specific recipe, they call it GRPOTCR, that was highly effective. And it relies on two techniques that sound technical but are actually quite straightforward in concept. A clip hire strategy and overlong reward shaping. Okay, can you break those down? Because clip hire and overlong reward shaping sound a bit like jargon, but you said they led to huge gains. They did. Massive gains. So clip hire, think of it like giving the agent a slightly longer leash during training.
6:06RL algorithms often have a safety mechanism, a clip, that stops the policy from changing too drastically in one step. The clip hire strategy just relaxes that constraint a bit if the agent finds a really promising you behavior that gets a high reward. So you trust the agent more when it finds something good. Pretty much. You let it run further down that promising path, even if it's a bigger change from its old strategy. Then there's overlong reward shaping. This tackles the issue where agents might generate really long rambling outputs. Instead of just failing the agent hard when it hits a length limit, this gives it a gentle penalty signal as it gets close to the limit.
6:42Ah, so it nudges it to be concise without just shutting it down. Exactly. It guides it away from unnecessary rambling without completely stopping useful exploration that might just need a bit more length. And the combination. The GRPOTCR recipe using these two tricks achieved results in just 100 training steps that took the standard, more conservative baseline about 450 steps. Whoa, 100 steps versus 450? That's like a 75 % reduction in training time. It's a massive efficiency game. Okay, that reduction is staggering. Now, this makes me think about the classic RL tradeoff, exploration versus exploitation.
7:18Is that still the main balancing act here, or do these external tools change things? They absolutely change the dynamic. In typical RL, yeah, you often have to sacrifice exploring new options to exploit the ones you know work, or vice versa. But with agentic RL, because the agent interacts with external tools over multiple turns, those tools actually bring in new information. They expand the possible solution space. So the tools themselves help with exploration. In a way, yes. And what's really fascinating is that the effective recipes, like this GRPOTCR, seem to improve both exploration and exploitation at the same time.
7:55Okay, wait. How does that work? Usually it's a trade-off. To get that, we need those metrics you mentioned earlier, PASIC-K and AVERAGE-AT-K. Can we quickly define those? Yes, please. Good idea. For listeners maybe not deep in RL, what are PASIC-K and AVERAGE-AT-K measuring? Sure. So PASA-K is basically a measure of exploration or potential. It asks, could the model solve this problem if it had K chances, like K attempts? It tells you the agent's ultimate ability boundary. Okay. It's best possible performance given a few tries. Right. Then average at K measures exploitation or reliable performance.
8:28It's the average success rate across those K attempts. How well does the model do typically? Got it. Potential versus typical success rate. Exactly. And in this agentic RL setting, with good recipes, both metrics go up together. The agent's potential, pass at K, increases, and its reliable performance, average at K, also gets better simultaneously. The main bottleneck for training efficiency then becomes the gap between those two. How quickly can the agent turn its potential ability into reliable, average performance? Interesting. So the goal isn't just balancing exploration and exploitation anymore.
9:04It's more about closing that gap between potential and reliable execution. Precisely. And this links back to that policy entropy idea we discussed. There's an entropy sweet spot. Weaker models often get stuck. They need a larger exploration budget, a higher value, for instance, to break out of their bad habits and find better strategies. They need more leash. Right. But stronger models, which might already be quite exploratory, actually need tighter bounds. If you give them too much freedom, their exploration becomes chaotic, unstable, and the training can even collapse. So you need to tailor the exploration budget, that clipping value, to the model's starting capability.
9:39Yes. Hitting that sweet spot of balanced entropy is key. Not too low, not too high. Okay. Data optimized, algorithm tuned. Let's hit our final pillar. The agent's internal reasoning mode. How does the model decide between thinking more internally versus calling one of those external tools? This was another really interesting observation. The study identified two distinct strategies agents tend to adopt. First, there's what they call the reactive mode. This is characterized by short bursts of internal thinking, often quite shallow, followed by frequent, almost rapid-fire tool calls. Like it's constantly asking the tool for help without much planning.
10:17Kind of. And this mode is usually adopted by the weaker agents, and maybe unsurprisingly, it leads to pretty low success rates when using the tools. Lots of calls, but many are ineffective. Then there's the second strategy, the deliberative mode. Here, the agent spends more time on internal reasoning. It seems to meticulously plan its next step before making fewer, but much more targeted and effective tool calls. Ah, so it's moving from being like a frantic micromanager constantly checking in to more of a high-level strategist who delegates specific tasks carefully. That's a perfect analogy. And the key finding is that this deliberative mode leads to far superior performance.
10:57Agents acting deliberately achieved over 70 % success rates with their tool calls because they actually invested the internal thought time to figure out the right call to make rather than just reactively hitting the tool API. It really underscores a quality over quantity principle for tool use. That strongly suggests we should just take models already graded internal thinking, like those long chain of VOD or long COP models, and just use them as agents. Yeah. Right. That give them an advantage. You'd think so, but ironically, no. Those models actually struggle quite a bit when turned into agents.
11:28When researchers tried using models that were pre-optimized for very long internal reasoning chains, they showed this strong bias against using tools. Really? They avoided the tools? Yeah. When faced with a complex task where a tool would be helpful, they tended to over-rely on their internal reasoning and actively avoided making tool calls. Their average number of tool calls actually dropped towards zero during agent training. So their strength, their deep internal thinking, became a weakness. Their ingrained habit to think longer actually stopped them from acting externally when needed. Exactly that.
12:04It created a conflict. Even when they tried to kickstart them with supervised fine-tuning, specifically showing them how to use tools, the model struggled. They were caught between two opposing goals. Learn to use tools, but also suppress this deep-seated bias towards just thinking internally. That's counterintuitive. It is, and the conclusion from the research is pretty clear. Standard instruction-following models, when trained on agentic reasoning tasks from the ground up using good data and RL recipes, ultimately scale better as agents. They don't carry that baggage, that conflicting prior about internal versus external actions.
12:39Okay, let's try and bring all this together then. Summarize the big takeaways from this deep dive into making RL work for AI agents. It really sounds like we have a practical, maybe even revolutionary recipe here. I think we do. It breaks down nicely. First, data. You absolutely need real interaction trajectories. They need to be diverse across different tasks. And they should be model aware, tailored to the model's learning zone. This builds a strong foundation with high policy entropy. Right. Real diverse tailored data. Got it. Second, algorithms. Simple, targeted tweaks within existing RL frameworks like that clip hire strategy and the overlong reward shaping in GRPOTCR can give you exponential gains in training efficiency.
13:18We're talking turning hundreds of steps into maybe dozens. Huge speed ups from simple fixes. Okay. And third. Third, reasoning mode. Quality over quantity wins. Agents that adopt a deliberative approach. thinking carefully, planning, then making fewer highly targeted tool calls. Vastly outperform reactive agents that just call tools frequently without deep planning. And the proof is in the pudding, right? This DemiAgent 4B model used exactly these techniques. That's the practical outcome, yes. DemiAgent 4B, using this systematic recipe, achieved state-of-the-art performance for its size on major agent benchmarks, matching or beating models many times larger, like 14B or even 32B parameter models.
14:01It really validates that careful technique, not just raw model scale, is driving the frontier of intelligent tool use right now. It does. It shows how much performance can be unlocked through optimization. So what does this all mean then? Where does this point for the next generation of AI development? Well, perhaps the biggest remaining challenge is how to reliably scale up that deliberate reasoning we see in the best agents. We've shown smaller models can do it exceptionally well with the right training. So the future direction this research hints at, it might involve a shift in focus, maybe less emphasis purely on making the LLM's internal reasoning engine bigger or better at calculation, and more towards improving its ability for high-level strategic planning and efficient tool orchestration.
14:44You mean the agent's core intelligence becomes less about doing the work itself and more about planning the workflow, knowing when and how to effectively delegate tasks to external tools. Exactly. The intelligence lies in planning the sequence, choosing the right tool for the right subproblem, interpreting the results, essentially managing the whole process rather than just crunching numbers internally. A shift from pure computation to strategic workflow management. A shift from internal calculation to strategic delegation. That's definitely a provocative thought for you, our listeners, to explore as you think about where AI agents are headed.
15:19Until next time on The Deep Dive.
From the publisher
The research paper systematically investigates how reinforcement learning (RL) can enhance the agentic reasoning capabilities of Large Language Models (LLMs), particularly in tool-integrated environments. The authors conduct a comprehensive empirical study across three main dimensions: data curation, algorithmic design, and reasoning mode to demystify optimal practices for agentic RL. Key findings include that real end-to-end trajectories are crucial for strong Supervised Fine-Tuning (SFT) initialization, while high-diversity and model-aware datasets improve training efficiency and exploration; algorithmically, techniques like clip higher and overlong reward shaping are effective for performance gains. Furthermore, the study identifies that a "deliberative mode" characterized by fewer but more successful tool calls outperforms frequent, reactive tool usage, and the authors introduce a new model, DemyAgent-4B, which achieves strong performance on challenging benchmarks compared to significantly larger models.




