Base models know how to reason, thinking models learn when

11 Oct 2025 · 12 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Why “thinking models” (e.g., DeepSeek R1, Claude 3.7 Sonnet) outperform base LLMs on hard math—arguing that base models already contain core reasoning skills, and post-training mainly teaches when/how to activate them.

Guest backgrounds

No guests are named in the transcript.

Key claims

Pre-training builds latent reasoning mechanisms; post-training teaches orchestration. Researchers map reasoning types using sparse autoencoders (bottlenecking internal activations). They then test causality with a “hybrid model” that injects steering vectors into the base model’s residual stream, timed by a classifier trained on successful expert (thinking-model) traces.

Notable examples

Math 500; hybrid steering recovered 91% of the performance gap while steering only ~12% of tokens (max ~21%). Random or mistimed steering reduced performance; different base models needed different steering targets (planning vs restatement).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Thinking Models

0:46 to 3:03

Discussion of how thinking models outperform base models in reasoning tasks.

“Which brings us to the core question, and this is really your mission, listening in.”

Cognitive Skills in Models

3:04 to 4:13

Exploration of whether new skills are learned or existing skills activated.

“What kinds of things showed up on that list?”

Using Sparse Autoencoders

4:14 to 5:52

Explanation of sparse autoencoders and their role in analyzing model outputs.

“Now, the really tricky part, how do you actually test the idea that the base models already have these skills, just hidden, without accidentally teaching them the skill while you're trying to measure it?”

Identifying Core Reasoning Skills

5:53 to 7:32

Discussion on identifying the core reasoning skills in models and their parallels to human thinking.

“No, it wasn't trained on the problems themselves.”

Testing Latent Skills in Models

7:33 to 9:49

Investigating how to test for skills in base models without altering them.

“Almost all of it recovered just by teaching the base model when to use the skills it apparently already possessed.”

Implications for Model Training

9:50 to 11:34

Examining the implications of timing and resource allocation in model training.

“The thinking model is just the master conductor knowing exactly when each section of the orchestra needs to play.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, let's unpack this. So today we're diving into something, well, pretty fascinating happening with large language models. You know how some of the newer models, the ones people are calling thinking models, like DeepSeek R1 or that new Claude 3.7 Sonnet? The ones that generate those really long step-by-step reasoning chains. Exactly. And suddenly they're just way better at really hard stuff, like, you know, competition-level math. They just blow the original base models out of the water. Yeah. That performance gap is, I mean, it's kind of what everyone's chasing in AI right now. When they use that extra thinking time, maybe driven by things like RLVR reinforcement learning with verifiable rewards, they just leap ahead on complex benchmarks.

0:41GS and 8K, MAT 500. The difference is dramatic. Right. Which brings us to the core question, and this is really your mission, listening in. Where does that incredible reasoning power actually come from? Did that extra training, that post-training phase, somehow teach the model completely new ways to think, new cognitive skills? Or, and this is the really mind-bidding idea, did it just teach the model when to use abilities it kind of already had? Abilities just sitting there, latent, in the base model. Well, the research we're looking at suggests a pretty nuanced answer. But it leans heavily towards that second idea.

1:18The basic thesis is this. Pre-training is where the model picks up the fundamental tools. Yeah. The ability to recall a fact or plan a step. That later stage, the post-training. That seems to be almost entirely about teaching the model how to deploy those tools efficiently. how to orchestrate them. So these thinking models aren't necessarily smarter in terms of raw knowledge. They're just better managers of what they already know. Okay. So if we want to prove that, that the skills are already in the base model, we first need to know what those skills even are. Right. I mean, looking at a long chain of thought output, just reading it and saying, okay, here it planned, here it calculated.

1:53That feels really subjective. How do you get around that? It totally is subjective. Yeah. That was the first big problem to solve. So to get an objective map, the researchers used a really clever bottom-up approach. They use these things called sparse autoencoders. Sparse autoencoders. Okay, sounds technical. For those of us not deep in the weeds of model architecture every day, what's an SAE doing here? Can you give us a picture? Sure. Think of the model's internal thought right before it spits out a sentence. It's like a really complex, high-dimensional signal. Very noisy. An SAE acts kind of like a filter or maybe like a a data compressor.

2:30Its job is to take all that complexity, all those activations, and force it through a really narrow bottleneck. It has to squash the information down. Ah, so it's forcing the model to reveal the most important parts, getting rid of the noise. Exactly. By making that bottleneck small, say, forcing the representation down to maybe just 20 dimensions, the SAE has to find the most fundamental patterns, the core reasoning steps. It filters out the linguistic fluff, the stylistic choices, and hopefully leaves you with just the, interpretable cognitive function happening underneath. Okay. So it gives us this clean, hopefully shorter list of what the model is actually doing.

3:06What kinds of things showed up on that list? What are these core functions? Well, they found a set of mechanisms that honestly look a lot like how humans solve problems, especially in math and logic. Things like recalling mathematical formulas, obviously, but also verifying intermediate steps, checking its work as it goes, and planning next steps. backtracking, like going back if something didn't work, and even conditional outcome projection, which is basically running a quick what-if scenario in its head. Wow. It's almost eerie how much that mirrors actual critical thinking skills. Plan, do, check, adjust.

3:41It really is. Yeah. And what's super interesting, maybe even a bit surprising, is how few of these core categories they found. Across different models, they tested the optimal number of these distinct reasoning types consistently landed somewhere between like 15 and 25. Only 15 to 25. That's it. Yeah. It suggests that complex reasoning isn't about having thousands of micro skills. It's more about having a fairly limited toolkit of core operations and just getting really good at sequencing them. Okay, so we have the map now, the 15 to 25-ish core skills. Now, the really tricky part, how do you actually test the idea that the base models already have these skills, just hidden, without accidentally teaching them the skill while you're trying to measure it?

4:25Right. That's the causality problem. You don't want to change the model while observing it. So to test this causally, they built something they call the hybrid model. And the whole thing relies on this concept of a steering vector. Steering vector. Let's break that down. What is that exactly? Okay. Imagine the model's internal processing state, its activations as moving along this path, the main residual stream. A steering vector is like a carefully calculated nudge, a direction in that high dimensional space. When you add this vector, this nudge, into the model's activation stream at the right moment, it reliably pushes the model's output towards a specific target behavior without changing the model's underlying weights.

5:07So let me see if I get this. If the base model is just about to say something generic, but you inject the planning next step steering vector, it'll instead output something like, okay, first I need to calculate X. That's exactly the idea. It's a targeted temporary intervention that activates a latent skill that's already there. So the higher model basically uses two parts working together. Okay, what are the two parts? First, you have the original base model. It doesn't get retrained. It holds all the potential skills that how to do things. The second part is this clever thing called the thinking model activation classifier.

5:39Think of it as the director or the orchestrator. Its job is to figure out the when. When is the right moment to deploy a specific reasoning step? Wait, how did this classifier know the right timing? Did they just look at the answer key for the math problems? That feels like cheating. Good question. No, it wasn't trained on the problems themselves. It learned by watching the expert, the fully trained thinking model. They basically analyzed the successful reasoning paths, the chains of thought from the superior thinking model. They looked at its internal activations when it was successfully solving problems and trained the classifier to recognize those patterns.

6:17So the classifier learned, okay, when the problem context looks like this, the successful thinking model usually activates its verify step mechanism right now. Ah, I see. So the classifier learns the timing from the expert model, and then when it sees a similar situation in the base model's processing, it says, okay, base model, now is the time, and injects the right steering vector. Precisely. It provides the orchestration cue. And the crucial thing, again, is that this is all happening in the activation space, that residual stream. They are not updating any of the base model's core parameters, its weights, which is strong evidence the capability was already there, just waiting for the right signal.

6:55Okay, the setup is ingenious, but the million-dollar question, did this actually work? Did this hybrid model, with its nudges, perform anywhere near the actual fully trained thinking model, or was it just a, you know, a cool theoretical idea? Oh, it worked. I mean, it really worked, especially when they tested it on the tough benchmarks. The results on Math 500, the competition level math stuff, were, well, frankly, they were stunning. Okay, lay it on us. What was the big number? How much performance did they get back? Get this. In the best case, they reported steering a Quinn 2.532B base model.

7:28Using the patterns learned from the QWQ32B thinking model, the hybrid approach recovered 91 % of the performance gap. 91 % just by adding these timed nudges. That's almost all of it. Almost all of it recovered just by teaching the base model when to use the skills it apparently already possessed. It's incredible leverage. That's amazing. But surely they must have been constantly steering it, right? Nudging it at every single step. That's the other kicker. No, you think so, right? But these huge gains came from steering only a tiny fraction of the tokens. Across all the different model pairs they tested, on average, they were only applying these steering vectors to about 12 % of the tokens in a given problem solution.

8:0812%. Yeah. And even the most intervention heavy case was only around 21%. So what? Intervening on maybe one token out of every eight, roughly? And that gets you 91 % of the way from the basic model to the expert model on super hard math. Wow. It really flips the script on needing massive retraining, doesn't it? It suggests reasoning improvement is much more about efficient timing and resource allocation. And the ablation studies back this up, too. They tried messing with it using random steering vectors or firing the right vector but at the wrong time. performance just tanked. So it proves both things matter.

8:43The specific skill you activate, the vector's direction, and the precise moment you activate it, the classifier's timing. Both are essential. And I thought it was interesting that the type of steering needed wasn't the same for all models. Like different base models needed different kinds of help. Absolutely. The orchestration wasn't a cookie cutter solution. For example, Lama models seem to benefit most from nudges towards planning next steps. It suggests maybe their bottleneck was organizing the sequence of operations. Okay. Whereas some of the Quinn models, they relied more heavily on getting skeered towards problem restatement at key moments.

9:18So the classifier learned to target the specific weak points of its base model. Okay. So let's pull back and synthesize this. If you put all these pieces together, what's the big takeaway for how we should think about these powerful LLMs? I mean, the biggest implication is that those fancy post-training methods, RLVR, distillation, whatever comes next, they're primarily teaching the model orchestration. They're building the conductor, not the instruments. The actual skills, the reasoning mechanisms, those seem to be largely baked in during that massive pre-training phase. The thinking model is just the master conductor knowing exactly when each section of the orchestra needs to play.

9:57That has some pretty profound implications for efficiency, right? If you don't need to retrain everything to get better reasoning. Huge implications. It suggests we might be able to get massive reasoning boosts with much more targeted, maybe even lightweight interventions. Instead of updating trillions of parameters, maybe we can focus on developing better ways to identify and activate these latent skills using things like steering vectors. Perhaps even dynamically at the moment you need them. Like installing a better control panel instead of rebuilding the whole engine. Exactly. We're not teaching a new skill.

10:28We're building a really precise, potentially cheap remote control for a skill that's already there. So the base model knew the steps. The thinking model just knew the choreography. Yeah, it's like the difference between knowing all the formulas for a physics test and knowing exactly which formula to use for which question when the clock's ticking. One is knowledge, the other is strategy. What a fascinating peek under the hood. So base models have the skills, thinking models master the timing, and crucially, that timing intervention is surprisingly sparse. Right. And, you know, it leaves you with a really provocative thought to chew on.

11:03If we now strongly suspect that pre-training is where these fundamental reasoning building blocks are acquired, well, how can we design that pre-training process differently? Can we structure it or monitor it to make sure those latent skills are laid down even more cleanly, more disentangled right from the start, so they're even easier to steer later? Instead of just teaching the timing after the fact, could we build models where the skills are perfectly modular and ready to be activated from day one? That feels like the next big challenge, doesn't it?

From the publisher

This paper argues that thinking language models (LLMs that reason step-by-step) do not acquire entirely new capabilities during post-training but rather learn when to deploy pre-existing reasoning mechanisms latent in their base counterparts. The authors use an unsupervised clustering methodology via Sparse Autoencoders (SAEs) to derive an interpretable taxonomy of distinct reasoning behaviors, such as numeric computation and planning next steps. They then implement a hybrid model that uses the base model for generation but is guided by the thinking model's activation patterns via steering vectors to activate specific reasoning behaviors. This hybrid approach successfully recovered up to 91% of the performance gap between base and thinking models on reasoning benchmarks like MATH500 while steering only a small fraction of tokens, supporting the idea that the primary benefit of complex training is teaching efficient mechanism deployment.

More from Best AI papers explained

All 475 episodes
Base models know how to reason, thinking models learn whenBest AI papers explained · 12 min
Listen in VO