The Path Not Taken: RLVR Provably Learns Off the Principals

23 Nov 2025 · 12 min · 4 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

White-box analysis of RLVR (reinforcement learning with verifiable rewards) showing a “sparsity paradox”: RLVR achieves big gains while leaving 36%–92% of parameters untouched.

Key claims

(1) sparsity is not random; RL updates show persistent, model-conditioned optimization bias (high update agreement; Jacquard overlap ~0.58; strike-like patterns in attention Q/K and output matrices). (2) The “sparsity” is amplified by BF16 precision limits (updates below ULP get rounded to zero). (3) Three gates explain the effect: KL anchor constrains step size, pre-trained geometry steers updates to low-curvature, spectrum-preserving subspaces, and BF16 filters tiny changes.

Notable examples

spectral singular-value profiles and principal subspace rotation ~0° for RLVR vs up to ~75° drift for SFT; scrambling transformer-layer geometry collapses the bias.

Guest backgrounds

No guests are named in the provided transcript.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding RLVR and the Sparsity Paradox

0:45 to 2:38

Exploration of the RLVR mechanism and its unexpected efficiency.

“they found RLVR gets these amazing results by changing shockingly little.”

The Mechanisms Behind Sparse Updates

2:38 to 6:12

In-depth analysis of the constraints that govern RLVR updates.

“Yes, because the sparsity we see is in many ways a superficial artifact.”

Empirical Evidence and Observations

6:12 to 8:02

Discussion of the empirical data supporting the theoretical claims.

“If the theory is correct that RL is preserving geometry, we should see it in the data.”

Implications for Fine-Tuning Strategies

8:02 to 10:55

Insights on how RLVR changes the approach to model fine-tuning.

“To prove the pre-trained geometry was the source of this bias, they went in and intervened.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive, the place where we try to cut through the noise and get straight to what's really going on inside this complex research. And today it is some really deep research. Oh, it is. We're doing a piece of, I guess you'd call it white box analysis. We're peering deep inside large reasoning models, LRMs. And specifically the ones that are tuned with reinforcement learning with verifiable rewards, RLVR. RLVR. And here's the paradox that we're really getting into today. We know RLVR is a huge deal. It's expensive. It's high compute. But it delivers these incredible gains in things like math, coding, complex reasoning.

0:38It's a step that takes a model from just being smart to being genuinely reliable. Exactly. But here's the bizarre part. When researchers actually looked at the model weights after all that training, they found RLVR gets these amazing results by changing shockingly little. Right. With something like supervised fine-tuning, SFT, you see dense updates. The whole network shifts. Maybe 80-99 % of the parameters move. But with RLVR, they found sparsity from, what, 36 % all the way up to an incredible 92 % of parameters, just untouched. It's the ultimate efficient upgrade. So our mission today is to figure out how you can change almost nothing and yet change everything about what the model can do.

1:19We need to get past just seeing the sparsity and find the where and the why. The underlying bias. The bias, yeah. What lets RL do so much with so little? Okay, so let's jump right into what the paper calls the sparsity paradox. The first big observation is that this isn't random. The RLVR updates, they're not just randomly sparse. Not at all. They show what's called a persistent model-conditioned optimization bias. Okay, so think of it less like a shotgun blast and more like a sniper hitting the same three targets every single time. That's a great way to put it. Researchers found that if they took the same base model and ran RL training on it multiple times independently...

1:56You'd expect some randomness. You would, but the weights that actually got updated showed this incredibly high agreement. They were landing in the same place over and over again. So even with the randomness of RL training, the update signature looked almost identical every time. Precisely. They measured it with things like Jacquard overlap, and it was around 0.58, which is way, way higher than random chance. And when you look at the visualizations, you see it. The updates form these clear, strike-like patterns. Not just noise. No, like row-wise updates in the query and key matrices and then column-wise in the output matrices.

2:29It's a very clear, predictable pattern. That concentration really suggests the model has these preferred pathways for learning. But that brings us to this really critical detail about what sparsity even is here. Yes, because the sparsity we see is in many ways a superficial artifact. It's a readout of the bias, but it's amplified by a hardware constraint. You're talking about B-Flow 16 precision. I am. So most large-scale training now uses BF16 or BF16 to be faster, but this format has a really limited number of bits for the fractional part of a number. Okay, so I know BF16 is for speed, but what's the implication for these really tiny updates the model might be making?

3:08It means there's a floor, the minimum size for a change to even be recorded. We call it the unit in the last place, or ULP. So if the optimization calculates this microscopic update, and that update is for a part of the network, the model doesn't want to change much anyway. The number is just too small to be stored. Exactly. It falls below the ULP threshold. The hardware can't represent it, so it gets rounded down to zero. Wait, so you're telling me this multi-million dollar RL process is effectively being filtered by a standard number format? It's a physical constraint acting like an amplifier.

3:41The optimization bias scares the update to be small in certain regions, and then the VF16 filter just says, nope, too small, and clips it. The weight looks sparse, but not because the gradient was zero, the update was real, just too tiny to register. That is a completely different way to think about it, which brings us to the core reason this happens, the implicit compass that's guiding RLVR in the first place. And this compass is really an interaction between three things. We can think of them as gates, constraint, geometry, and precision. Let's start with the first one, the constraint. This is the KL anchor.

4:15Right. So on policy RL is, by its nature, pretty conservative. Even if you don't have an explicit KL penalty, like in some algorithms, there is an implicit KL leash. Leash. It stops the new policy from drifting too far away from the old one in a single step. The math proves this directly limits the overall size of the weight update. It's a stability thing. If you take too big a step, your model can just destabilize. So the training itself forces the model to only take these tiny, careful steps. It's conservative by design. Exactly. And that small, bounded step then leads us to the second gate, which decides the direction of that step.

4:53That's the model geometry. The steering core. The steering core, yes. Because the step is small, the model's existing, pre-trained geometry acts like a really powerful compass. It points the update down the path of least resistance, the safest, most efficient path. So rather than trying to, I don't know, bulldoze a new path through the model structure. It finds a detour. A low resistance detour. Yeah, the mountain analogy here is perfect. SFT often has a very specific goal, and getting there might mean going right over the mountain, a high curvature area that takes a ton of energy and radically changes the weights.

5:26But RLVR. RLVR, constrained by that KL leash, gets steered towards these low-curvature, spectrum-preserving subspaces. It takes the easy way. It goes around the mountain. It's like a smart contractor who knows which walls are load-bearing and makes sure to leave them alone. That is the key implication. RL updates preserve the model's core spectral structure. They don't tear it down and rebuild it. And that finally brings us back to the third gate, precision. The filtering gate. All those small changes that were routed into low curvature areas are now subject to that BF16 precision limit. The tiniest ones get filtered out, and that's what makes the final result look so sparse.

6:04The paradox is basically solved by these three gates working together. Okay, for us to really buy into this, we need to see the data. Let's look at the empirical proof, starting with the spectral geometry comparison between RLVR and SFT. Right. If the theory is correct that RL is preserving geometry, we should see it in the data. And, well, it's pretty striking. Very easy. With RLVR, you see almost no spectral drift. The singular value profiles for the base model and the tuned model are nearly identical. And the principal subspace rotation, which is literally the angle of the model's main compass, it's almost zero degrees.

6:40The core structure is unchanged. Untouched. SFT, on the other hand, is a total overhaul. It causes these huge repatients, sometimes up to 75 degrees, and massive spectral drift. It's actively changing the model's fundamental directions. So SFT is turning the entire ship while RLVR is just tweaking the rudder. A very good way of putting it. Okay, that brings us to the next piece of evidence, which connects that low-curvature path directly to the model's geometry. Principal weights alignment. First, can you define principal weights for us? Sure. Principal weights are just the high magnitude weights.

7:14They're often in the most curved parts of the loss landscape. They're your load-bearing walls, the super influential pathways the model relies on. And they're what SFT usually targets because changing them has a big impact. Exactly. So if our theory is right and RL is avoiding these high curvature areas, it should also be actively avoiding these principal weights. And is it? Dramatically so. The research shows that RL updates have a sub-random overlap with the principal weight mask. It's not just avoiding them. It seems to be actively targeting the non-principal low-magnitude weights. It's deliberately working on the periphery.

7:51On the decor, not the foundation. Because that's the safest, most stable path for improvement when you're on that KL leash. This feels like the smoking gun. Yeah. And the researchers even provided some causal evidence to lock this down. They did. This is really cool. To prove the pre-trained geometry was the source of this bias, they went in and intervened. They deliberately scrambled the geometry of a few transformer layers using things like orthogonal rotations. So they just rearranged the internal compass for those specific layers. Exactly. And in the layers, they scrambled. The optimization bias, that predictable pattern of updates, completely collapsed.

8:25It went to random levels. But in the layers they left untouched, the bias was still there, strong as ever. It's definitive proof that the base model's innate structure is what creates the path of least resistance for RL. The geometry is the compass. And we should say this signature, this minimal rotation and off-principle updating it, holds up across all kinds of RL tasks. Math, coding, RLHF, even agentic tasks. This isn't a one-off fluke. No, it seems to be the fundamental way K-Lankard RL works. And that has some really profound practical implications. It means we have a white box understanding now.

9:03We know RLVR is playing a totally different game than SFT. Which means all those clever, parameter-efficient fine-tuning methods, PEFTs, that were designed for SFT, they might be completely wrong for RL. Let's talk about that misalignment. Research actually tested this, right? They tried to apply an SFT-style tuning strategy to RL. And the failure was immediate. It was catastrophic, really. They tried forcing the updates to only target the high-curvature principal weights the SFT way. The whole process just destabilized and accuracy tanked. You're forcing the agent to climb the mountain when its whole design is about going around it.

9:35Exactly. But then they tried the opposite. The RLVR native solution. Right. They created a safe mask where they only fine-tuned the non-principle low-magnitude weights the areas RL wants to go to anyway. And that sparse training almost perfectly tracked the performance of full, dense RLVR. Which is a huge insight for efficiency. Massive. It suggests you can get all the gains by updating maybe 70 % of the range. parameters as long as you pick the right 70%, the low energy pathways. And we see the same story with low rank adapters like LORO. We do. Standard LORO works pretty well with RL and it turns out that's because its structure naturally favors these off-principle updates.

10:15It accidentally aligns with RL's dynamics. But the moment you try to use a more advanced LORO variant designed for SFT, like PISA, which tries to force the updates to align with those high curvature principal directions, it breaks. It doesn't help at all. PISA gave zero gain over standard LoRa. And worse, if you pushed the learning rate higher to try and eke out more performance, PISA would often just cause the training to collapse entirely. Because you're fighting the model's fundamental nature. You're pushing against that KL stability boundary way too hard. You are. Forcing updates into those principal directions just fundamentally contradicts the geometry-preserving dynamics of RLVR.

10:53It's almost poetic, isn't it? We spend all this time trying to force models up the steepest hill, only to find out the best way to improve them is to respect their structure and take the easy path. The success comes from working with the neural geometry, not against it. And this really changes how we need to think about building these models going forward. We need our native PFTs, ones that are geometry aware. Okay, so let's wrap this up. The central takeaway from this deep dive is that RLVR isn't sparse by accident. it's a highly efficient process constrained by that KL anchor and guided by the model's own geometry to make these minimal low curvature changes.

11:30And that sparsity we see is just a side effect. It's the sign of a stable geometry preserving optimization strategy that's working as intended. We've moved from just guessing at how RL works inside these huge models to actually having a map of its internal compass. Which is vital for anyone trying to build the next generation of these things efficiently. And that leaves us with a final, I think provocative question for you to think about. If a pre-trained model's existing geometry dictates where it can safely and easily integrate new knowledge, if that geometry determines its RL friendliness, can we start designing pre-training to create a better internal compass from the start?

12:09Should we judge future-based models not just on their raw performance, but on how gracefully they can be fine-tuned? Something to think about.

From the publisher

This paper studies mechanistic explanation for the paradox that **Reinforcement Learning with Verifiable Rewards (RLVR)** reliably improves large language model reasoning while making only minimal, sparse changes to parameters. The authors introduce the **Three-Gate Theory**, arguing that sparse updates are a surface artifact of a **model-conditioned optimization bias**. **Gate I (KL Anchor)** constrains each update, while **Gate II (Model Geometry)** steers the updates off the principal, high-curvature directions favored by **Supervised Fine-Tuning (SFT)** and into low-curvature subspaces, thereby preserving the model's spectral structure. **Gate III (Precision)** amplifies the appearance of sparsity by masking small updates in non-preferred regions due to bfloat16 storage limits. Consequently, the work demonstrates that **RLVR learns in a distinct optimization regime from SFT**, which suggests that SFT-era parameter-efficient fine-tuning (PEFT) techniques are often ill-suited for RL applications.

More from Best AI papers explained

All 475 episodes
The Path Not Taken: RLVR Provably Learns Off the PrincipalsBest AI papers explained · 12 min
Listen in VO