Emergent hierarchical reasoning in LLMs through reinforcement learning

14 Dec 2025 · 13 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How reinforcement learning reveals emergent hierarchical reasoning in LLMs, explaining “aha moments” and length scaling, and proposing Hierarchy Aware Credit Assignment (HICRA/IICRA) to target strategic reasoning tokens rather than treating all tokens equally.

Guest backgrounds

No guests are named in the transcript; it’s a host-led discussion referencing “the sources” and “research” without identifying speakers.

Key claims

RL performance jumps come from learning a high-level planning layer (strategic grams) after reliable low-level execution is forged; strategy needs exploration (higher semantic entropy) while execution needs certainty. Prior methods like GRPO waste credit on execution tokens.

Notable examples

Strategic grams such as “first, let’s define our variables” and “let’s deduce” appear once or twice per solution but recur across problems; phase 1 lowers perplexity/entropy for execution tokens, phase 2 raises semantic entropy for strategic tokens. HICRA rewards planning tokens more on correct answers and dampens penalties for strategic failures, improving text-only and multimodal (image) reasoning.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI's Aha Moments

0:45 to 2:00

Explores the phenomenon of AI models making unexpected leaps in performance.

“We were treating every single word, every token the model generated as if it were the same.”

Internal Structure of AI Reasoning

2:00 to 3:18

Discusses the two types of reasoning tokens in AI: low-level execution and high-level planning.

“You've got your high-level planning, the strategy, and then you have the low-level, almost mechanical execution.”

Strategic Grams in AI

3:18 to 4:21

Explains how strategic grams help AI identify high-level reasoning patterns.

“It's the model deciding what to do next, not just how to do the math.”

Phases of AI Learning

4:21 to 5:39

Describes the two phases of AI training: mastering execution and steering with strategy.

“The research suggests that the whole RL training process isn't this one continuous slog.”

The Shift to Strategic Planning

5:39 to 7:40

Examines how AI transitions from executing tasks to developing strategic thinking.

“which really supports this idea that these low-level skills are just the ticket to the game.”

Introducing Hierarchy Aware Credit Assignment

7:40 to 9:20

Discusses the new RL method IICRA that optimizes AI training for strategic reasoning.

“And that perfectly explains the length scaling thing, too.”

Evaluating AI Learning Effectiveness

9:20 to 11:09

Explores the need for new metrics to measure AI learning and reasoning performance.

“It encourages what they call anisotropic exploration.”

Rethinking AI Training Strategies

11:09 to 12:54

Encourages a shift in AI training to focus on strategic ideas rather than individual tokens.

“The sources also made a point about not just targeting any high entropy token, but specifically the planning tokens.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you've been following the world of advanced AI, you've probably heard these strange stories from the training labs. You know, these moments where a model is struggling with, say, advanced math problems and then suddenly, boom. A aha moment. Exactly. A massive, totally unpredicted leap in performance. And what's even weirder is this other effect they call length staling. Right. The idea that sometimes for an AI to get the right answer, it just needs to, well, to think longer, to write out a more detailed internal monologue. And it's like, why? Why does just writing more make it smarter? For a long time, these things were just, they were mysteries.

0:36We knew reinforcement learning, RL, was making these models better at complex stuff, but the how was a complete black box. It really was. And what the sources were digging into today show is that the process was so opaque and, frankly, pretty inefficient because of a really basic assumption. Which was? We were treating every single word, every token the model generated as if it were the same. The optimization pressure was just, you know, spread evenly across the board, completely agnostic to what a word was actually doing. And that's our whole mission for this deep dive. We're going to peel back those layers because what's happening isn't just about the AI learning more facts.

1:12It's actually developing this hidden hierarchical way of reasoning that, well, it looks a lot like how you and I would tackle a hard problem. It's fascinating because that old inefficiency, it just came from that one oversight. I mean, think about it. If you're solving a puzzle, the part of your brain that says, OK, let's try a different strategy here is doing a totally different job from the part that does the simple addition. One is strategy. The other is just raw execution. Exactly. And once researchers could tell the difference between those two types of thought tokens, they could finally target the real bottleneck holding back AI reasoning.

1:50OK, so let's unpack that. Let's get into this internal structure. Before we talk about the new solution, this targeted RL approach, we have to understand these two different jobs that tokens are doing. You said it parallels human thought. It's a really clean parallel. You've got your high-level planning, the strategy, and then you have the low-level, almost mechanical execution. It's the architect versus the construction worker. I like that. So what do these look like inside an AI's chain of thought? Well, on one side, you have what they call low-level execution tokens. These are the building blocks, the pure mechanics of it.

2:23So like the numbers and the plus signs. Exactly. Arithmetic, putting a value into a variable, applying a formula you already know, even just formatting the answer correctly. Their only job is to be procedurally correct. They're the grunt work. Making sure 2 plus 2 actually equals 4. Precisely. And then on the other side, you have the really interesting stuff, the high-level planning tokens. The research has a great name for them. Strategic grams. Strategic grams. Okay. These are the tokens that are actually steering the whole process through the architect, drawing the blueprint. They handle the big logical moves.

2:57So not equals four, but something more like half. More like, given this fact, we can deduce. Or let's try a different approach. That's a branching move. Or even self-correction, which is a huge one. Oh, right, where it kind of catches its own mistake. Yeah, it will generate something like, wait, hold on. Earlier I assumed X, but the problem actually says Y. That's not a calculation. That's a high-level strategic decision. It's the model deciding what to do next, not just how to do the math. So how on earth do you get a computer to reliably identify something so abstract? How do you tell a strategic thought from just a random word?

3:35It's actually a really clever statistical trick. They define these strategic grams, or SGs, by their pattern of use. Think of them as reusable bits of scaffolding. They're short phrases, usually three to five words. They have a very specific signature. They pop up all the time across thousands of different problems, but within any one solution, you might only see it once or twice. I think I get it. So a phrase like, first, let's define our variables, would be a strategigram. You use it in tons of different math problems, but you only say it once at the beginning of each one. You've got it. That statistical pattern is what separates the structural guiding language from the repetitive low-level calculation language.

4:14And that new lens, just being able to see that difference, is what explains those aha moments we talked about at the start. The research suggests that the whole RL training process isn't this one continuous slog. It actually happens in two very distinct phases. Yes, and it shows that the learning bottleneck actually moves over time. Phase one is what they call forging reliable low-level skills. The construction worker phase. The construction worker phase, exactly. The model first has to get its tools in order. It has to become absolutely reliable at the basics. Because think about it, one tiny calculation mistake can make the most brilliant, complex strategy completely worthless.

4:54And you can actually see this happening in the metrics, right? You can watch the model getting better at this. You can. During this phase, you see a really sharp decrease in two key metrics for those execution tokens. Perplexity and entropy. Which basically means the model is getting more confident and more certain about how to do the simple stuff. Right. It stops being surprised by the need to write an equal sign. It just knows. It's drilling its scales, you know, like a pianist. The uncertainty about how to play a C major scale goes down to zero because it becomes muscle memory. So it's building its reliable toolbox.

5:28And there's this great little piece of evidence for this. For models that are already really capable, like some of the more advanced ones they tested, this first phase is either super short or sometimes it's not even there. Because they already have the toolbox. They've already done that work. which really supports this idea that these low-level skills are just the ticket to the game. They're not the game itself. Okay, so this is where it gets really good. Once the toolbox is built, once the model can do the math reliably, the whole learning problem shifts. And this is phase two, steering skills with strategic planning.

6:01This is where the magic happens. The model has mastered how to calculate. Now it has to learn what to calculate. It has to explore and master a whole playbook of high-level strategies. And the evidence for this is the opposite of phase one, instead of uncertainty going down. It actually goes up. But, and this is the critical part, it only goes up for the strategic tokens. You see a steady increase in something they call semantic entropy. And semantic entropy isn't just about word-level randomness. It's measuring the diversity of the actual plans the model is using. Correct. The model is actively trying out new things.

6:35It's expanding its strategic playbook. It's learning how to do deduction, and then learning how to do reasoning by contradiction, and then learning how to self-correct. This strategic diversification is what really drives the improvement in reasoning. But wait, that seems a little strange. If the goal is to get better and more correct, why would you want more entropy, more diversity? Wouldn't you want it to find the one best strategy and just get really certain about that? That is a fantastic question, and it gets to the core of it. For execution, yes, you want certainty. But for strategy, you need exploration.

7:10A composer doesn't want to know just one chord progression. They need a huge vocabulary of them to solve complex musical problems. So the model needs high strategic entropy. It has to try different logical paths. And yeah, a lot of them will fail at first. But that exploration is the only way it can discover that small handful of truly powerful general-purpose strategies that work everywhere. So the aha moment isn't the model suddenly learning a new math fact. It's the moment it discovers and masters a whole new type of strategic thinking. And that perfectly explains the length scaling thing, too.

7:45Yeah. A model with a bigger strategic playbook will naturally produce longer, more careful, more deliberate chains of thought. That leads to better answers. Okay, so if this high-level planning is the real bottleneck, then the way we used to do RL with methods like GRPO was just fundamentally inefficient. You're just wasting all this optimization pressure on the simple stuff. Totally inefficient. It's like giving your architect and your construction worker the exact same bonus, regardless of who came up with the brilliant design. So what's the solution? The solution they came up with is called Hierarchy Aware Credit Assignment, or for short, IICRA.

8:20And it's an RL method designed to do one thing, focus all that optimization pressure right on the strategic bottleneck. How does it do that? How does it separate the two? It uses what they call asymmetrical credit. So when the model gets a problem right, HIHE finds all the planning tokens, the strategic grams in that successful solution. The let's deduce or let's reconsider phrases. Exactly. And it gives them an extra boost of reward. It amplifies the credit for good strategic thinking. It's like giving the architect the bigger bonus. It is. But here's the really clever part. It's what it does when the model gets the answer wrong.

8:58OK. On a failed attempt, HITT goes back and finds the planning tokens again. But this time it dampens the penalty. It softens the punishment for trying a new strategy that didn't work out. So wait, if the model tries some wild new strategy and fails because it made a simple addition error, the penalty for the bad math is still high. Oh yeah, very high. But the penalty for the risky but creative strategic move is much, much lower. That is the whole philosophy. It encourages what they call anisotropic exploration. It's not just random exploration, it's pushing the model to explore specifically in the direction of new strategies.

9:34It makes strategic failure cheap and execution failure expensive. That's a huge change. You're basically teaching the model that it's okay to take risks with your planning because that's how you learn. And the results were, I mean, they were really clear. HICRA just consistently blew the old GRPO method out of the water on all sorts of complex reasoning tasks, from text-only problems to multimodal ones with images. And the error analysis just sealed the deal. Yeah, it was the final nail in the coffin for the old method. They looked at what kinds of mistakes RL was actually fixing. And the vast majority of the performance gain came from fixing flaws in high-level plans, not from fixing little calculation mistakes.

10:13The real leverage is in teaching the model how to think, not just how to add. So this changes everything about how we should even measure AI learning. The old way of tracking things just doesn't work anymore. It's totally misleading. If you just look at the average token entropy across the whole output, it's useless. In phase two, as the model gets super confident about all the low-level stuff, that average entropy goes down. Which makes it look like the model has stopped exploring, stopped learning. Right. It looks like it's plateauing, even while its actual problem-solving ability is still shooting up.

10:45It's because you're averaging the high entropy of the creative strategy with the near-zero entropy of the mastered execution. The pianist is still exploring new compositions, even though they're 100 % certain how to hit the keys. Perfect analogy. And that's why you have to switch to measuring semantic entropy. You have to look at the diversity of the strategic units themselves. That metric actually tracks with performance. It tells you when the model is really getting smarter. The sources also made a point about not just targeting any high entropy token, but specifically the planning tokens. What's the difference?

11:18It's about precision. Look, a lot of planning tokens are high entropy because they're decision points. Yeah. But the reverse isn't true. Most high entropy tokens are not planning tokens. They're just noise. They could be noise. They could just be slightly different ways of phrasing a simple calculation. You know, the result is versus which equals. Those don't change the high level plan at all. If you just reward all high entropy tokens, you're wasting a lot of optimization on that meaningless variation. The functional definition, the strategic grams, is way better because it targets the actual role of the token, not just its statistical weirdness.

11:54It targets the cognitive job it's doing. Exactly. This has been, well, a real deep dive. So to sum it all up for you listening, when you see these amazing leaps in AI reasoning, it's not just about the model getting faster at math. It's about the emergence of this hidden strategic planning layer. The AI is literally learning how to think. It's a total paradigm shift. The old idea of treating all tokens the same is just, it's over, it's too inefficient. And this opens up so many new directions for research, we can start thinking about valuing the process. You mean like rewarding a model for coming up with a really brilliant plan, even if it messes up the final calculation?

12:30Why not? That's how we teach people. We praise the correct line of reasoning. We need to rethink the whole action space for RL, not as individual words, but as these meaningful strategic ideas. This new hierarchical view of how these models think, it feels like we're just scratching the surface. We're just getting started. But it's the first real map we have of how intelligence actually emerges from these systems. And maybe it's something for you to think about, too, in whatever you're trying to learn. Recognizing what part of the skill is just execution that needs drilling, and what part is strategy that needs wide-open, messy, and creative exploration.

From the publisher

This paper discusses how a successful RL fine-tuning uncovers an emergent two-phase hierarchical reasoning dynamic in LLMs, mirroring human cognition by separating high-level strategic planning from low-level procedural execution. The authors argue that conventional RL methods, which apply optimization pressure agnostically to all tokens, are inefficient because they fail to concentrate learning efforts on the true bottleneck: mastering strategic planning tokens. The proposed method, HICRA, addresses this by selectively amplifying the learning signal for these high-impact planning tokens, with extensive experimental results demonstrating that this targeted approach significantly outperforms baselines like GRPO across various mathematical and multimodal benchmarks. The paper also introduces Strategic Grams and Semantic Entropy as diagnostic tools to accurately track this strategic exploration, revealing why common metrics like token-level entropy are often misleading.

More from Best AI papers explained

All 475 episodes
Emergent hierarchical reasoning in LLMs through reinforcement learningBest AI papers explained · 13 min
Listen in VO