In short
The episode explains a research paper (“The Invisible Leash: Why RLVR May Not Escape Its Origin”) arguing that reinforcement learning with reward models (RLVR) mainly sharpens a model’s existing capabilities rather than discovering fundamentally new reasoning.
Guests
None named as podcast guests; the episode discusses authors Fang Wu, Weihao Xuan, and colleagues.
Key claims
Support preservation—RLVR cannot assign non-zero probability to solutions the base model initially gave zero probability to; thus RLVR’s PAS@K can’t asymptotically beat the base model. RLVR acts like conservative reweighting (minimal distribution change) and reduces answer-level entropy (less diversity).
Notable examples
Experiments with ProRL starting from DeepSeek R1 Distill-Qwen 1.5B on Math 500/Minerva/Olympiad Bench/AM and QA benchmarks. Results show more empirical support shrinkage than expansion: e.g., Minerva+Olympiad Bench gained 3 new answers but lost 48; SimpleQA+SciBench gained 21 but lost 55. Some expansion occurs (e.g., new Olympiad Bench solutions; Reasoning Gym tasks like BoxNet/Arc1D), but it’s less common.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding RLVR
0:30 to 1:51
Exploring the concept and implications of RLVR in AI models.
“Basically, you take a model that's already been trained, a base model, and you fine tune it using reinforcement learning.”
Support Preservation Explained
1:52 to 4:28
Delving into support preservation and its constraints on RLVR.
“So to get at this core limitation, this leash idea, we need to start with a concept from the paper they call support preservation.”
Entropy and Its Trade-offs
4:29 to 6:31
Analyzing the entropy reward trade-off and its effects on model outputs.
“This is not about finding fundamentally new ways to solve a problem.”
The Experiment Findings
6:32 to 9:58
Reviewing the experiments conducted to test RLVR's effectiveness.
“And the paper even finds this effect, this empirical support shrinkage, happening even within math problems sometimes.”
The Implications of Findings
9:59 to 14:00
Discussing the broader implications of RLVR limitations in AI development.
“When they did a granular comparison, the shrinkage generally outweighed the expansion.”
Exploring the Limits of RLVR in AI
14:00 to 16:08
Learn how RLVR techniques are constrained by existing models and what it means for AI exploration.
“And this explains some practical things people are already doing.”
Transcript
Automatic transcript. May contain errors.0:28Welcome to the Deep Dive. with verifiable rewards. RLVR for short. Yeah, RLVR. Basically, you take a model that's already been trained, a base model, and you fine tune it using reinforcement learning. It gets rewards automatically calculated, telling it, yes, that's right, or no, try again. So it's like giving the AI specific feedback to get better. Sounds really powerful. It is powerful for optimization, definitely. But here's the big question, and it's causing quite a stir in the research world. Is RLVR actually expanding what these models can fundamentally reason about, pushing them into totally new cognitive areas?
1:05Or is it just making them sharper, you know, better at applying what they already kind of knew, just amplifying the high reward stuff for better precision? That is the absolute core question, and it's what we're diving into today. We're unpacking a really fascinating new study. It's called The Invisible Leash, Why RLVR May Not Escape Its Origin by Fang Wu, Weihao Xuan, and their colleagues. Okay, The Invisible Leash. So our mission for you today is to pull back the curtain on this RLVR stuff. We'll get into the theory, the limits the paper talks about, and some, well, pretty surprising results from their experiments.
1:40We want to show you why RLVR, even though it's successful, might be held back by this invisible leash and importantly, what that implies for where AI reasoning is headed. OK, let's dig in. Let's do it. So to get at this core limitation, this leash idea, we need to start with a concept from the paper they call support preservation. What exactly do they mean by support here? OK, so think of a model's support as like the entire universe of possible answers or solutions it could theoretically generate. anything it assigns even a tiny non-zero probability to. Okay, it's potential output space. Right.
2:16Now, the paper argues, and this is key, if a specific solution has zero probability according to the initial base model, then RLVR, because it learns by sampling things the base model generates and then getting rewards, well, it can never learn to assign that zero probability solution, a non-zero probability. Ah, I see. So it's like you said, if the base model thinks something is impossible from the start. RLVR can't make it possible. It's stuck within that initial realm of possibility defined by the base model. So the analogy they use, or that we're using, the dog on the leash, it really fits, doesn't it?
2:50It does. No matter how much you reward the dog for trying to reach something outside its leash radius, it just physically can't get there. Exactly. The base model's initial probability distribution sets the boundary. The paper actually formalizes this with theorem 2.2. It shows the support of the RLVR model is always a subset of the base model's support. Always a subset. Okay. Which means, yeah, RLVR can't discover solutions that the base model deemed impossible. No matter how good those solutions might actually be, how high the reward. And this has a direct consequence for how we measure performance, right?
3:22Like PASIC-K. Yes, exactly. Corollary 2.3 in the paper follows from this. It shows that even if you had infinite samples, an infinite amount of training time, the RLVR model can't asymptotically get better than the base model on PASIC-K. That's the metric for whether at least one correct answer is found within K-tries. If the base model couldn't find any correct answers in its support for a problem, RLVR can't magically create one. Wow. Okay. So that really reframes RLVR. It's not this explorer charting new territory. The paper calls it something like a conservative reweighting mechanism. That's the term they use, and it's a really critical insight.
4:01They look at it through a mathematical lens variational inference for the technically inclined. And what it shows is that RLVR actually tries to make the smallest possible changes to the base model's way of thinking, its distribution. Smallest possible changes. Yeah. The optimal policy it finds is basically the one that's closest, in a mathematical sense, to the original base model, while still managing to achieve whatever reward target you set. Even if you remove some of the mathematical constraints, the regularization, they show in Corollary 2.7 that the updates are still fundamentally anchored to that base model.
4:37This is not about finding fundamentally new ways to solve a problem. It's about taking the ways the base model already knows and making them much more prominent, more probable if they lead to rewards. Precisely. It's making the existing wheel turn really, really well, very efficiently. OK, makes sense. It ensures things are stable. The updates are efficient. But that efficiency comes at the cost of exploration. That's the trade off. And this conservative nature, this anchoring, it leads straight into another major finding, doesn't it? the entropy reward trade-off. What's entropy measuring here?
5:09Right. So in this context, entropy is basically a measure of diversity or randomness in the model's answers. Answer level entropy, specifically. The study finds quite clearly that RLVR systematically reduces this answer level entropy. Reduces diversity. So the model starts to converge on fewer distinct answers. Exactly. Theorem 2.8 states it formally. The entropy of the RLVR model's output distribution will be less than or equal to the base model's entropy. It generally goes down. Okay. Now, intuitively, less diversity sounds good if you want the one right answer, like in a math problem, right?
5:44Sure, sure. If there's a single best solution, like a mathematical proof or the optimal move in a game, then lower entropy is exactly what you want. It means the model is zeroing in on that correct answer, increasing the probability of getting it right on the first try. That's the pass at one metric we talked about. Better pass at one. More precision. Yes. But the flip side is, by collapsing onto fewer answers, it might ignore other perfectly valid, maybe slightly different, correct solutions, especially if you're sampling multiple times, hoping for variety. Hmm, I can see that. So great for math proofs, maybe, but what about tasks where there isn't just one right answer, like creative writing, or dialogue systems, or maybe even coding suggestions?
6:25That's where the concern lies. Reducing diversity in those more open-ended domains could be a drawback. You might suppress interesting, alternative, equally valid outputs. And the paper even finds this effect, this empirical support shrinkage, happening even within math problems sometimes. It can lose track of some correct answers it could previously find. Wow, okay. Shrinkage. We need to talk about the experiments then, because this isn't just theory. They tested this quite a bit. Oh yeah, extensive experiments. They used a pretty sophisticated RL-VR method, ProRL. starting with a decent base model, DeepSeq R1 Distilquan 1.5b.
7:00And they tested it across a whole range of tasks. Math reasoning benchmarks like Math 500, Minerva, Olympiad Bench, AM. The tough stuff. Right. And also non-math reasoning tasks, simple QA, live bench, side bench, reasoning gym, a real mix. And they introduced this idea of empirical support. How's that different from the theoretical support we discussed? It's a practical refinement. Theoretical support includes any answer with non-zero probability, even if it's astronomically small, like 10 on 100. Empirical support looks at the set of correct answers that the model assigns non-negligible probability to under its actual sampling process.
7:37The stuff you might actually see if you run the model. The stuff that's practically reachable. Makes sense. Yeah, because technically, with softmax outputs, everything has some tiny probability, but most are effectively invisible. Okay, so what did the experience show then? Did RLVR manage to break the leash and find lots of new stuff in practice? Well, the dominant finding was that RLVR mostly sharpens the probability distribution within the empirical support the base model already had. It makes the known good answers much more likely. For instance, they looked at the total set of correct answers found by either the base model or the RLVR model on Olympiad bench and side bench.
8:14The combined number was very similar, around 600 or so in both cases. So not much net gain in the total pool of reachable correct answers. Largely, no. It mainly reallocated probability towards the better answers within that existing pool. Figure 1 in the paper illustrates this quite well. But were there any cases where RLVR did seem to expand the horizon, find things the base model just couldn't? Yes, there were instances of what they term empirical support expansion. Moments where RLVR did assign noticeable probability to correct answers that were practically impossible for the base model to find.
8:50Uh-huh. Like what? They found, for example, three new solutions on Olympiad bench that the base model missed. 11 on simple QA, 10 on scibench. Not huge numbers, but evidence of some expansion. And some tasks in Reasoning Gym, like BoxNet and Arc1D, showed really striking expansion. RLVR got near-perfect scores where the base model was really struggling. Figure 3 shows this. Okay, so it can happen. The leash isn't perfectly rigid, maybe. It can stretch sometimes? It seems so, occasionally. But, and this is a big but. Uh-oh. The paper found that empirical support shrinkage was actually a more common and often larger effect.
9:23Shrinkage. So RLVR actually lost access to correct answers the base model could find. Yes. It failed to recover them. For example, on the AME math benchmarks, it missed three solutions the base model found in 2024, another three in 2025. And the losses were sharper on Minerva and Olympiad Bench, losing 22 and 26 correct completions, respectively, that the base model could generate. Wow. And on Simple QA and PsyBench, it forfeited 20 and 35 correct completions. Figure 2 in the paper gives examples of these lost solutions. So the tradeoff is real. Gain some precision, maybe occasion to find something new, but often lose some of the breadth you started with.
10:02That seems to be the pattern. When they did a granular comparison, the shrinkage generally outweighed the expansion. Like across Minerva and Olympiad Bench combined, RLVR found only three totally new answers but lost 48 that the base model knew. Three gained, 48 lost. That's significant. It is. And on SimpleQA and SciBench combined, it gained 21 new ones but forfeited 55. Okay, this is huge. What does this tell us then, fundamentally, about what RLVR is doing? It really reinforces the idea that RLVR is acting primarily as that sampling reweighting mechanism. It's reshuffling probabilities within the base model's known world.
10:40It delivers higher precision, which is valuable, but it doesn't seem to be a reliable engine for discovering fundamentally new reasoning paths. It's more about refinement than discovery. More precision enhancer than reasoning discoverer. Exactly. And this connects to other known issues, too, like temporal forgetting or catastrophic forgetting and continual learning. Right, where training on a new task makes a model forget how to do older tasks. Yeah. It seems this focusing effect of RLVR might be a related phenomenon concentrating on high reward paths can inadvertently prune away other useful, previously known paths.
11:14Okay, let's circle back to entropy for a minute because you mentioned the paper distinguishes between token-level and answer-level entropy. That sounds important. It is, and it's quite nuanced. Token-level entropy measures the uncertainty at each step of generating an answer. Like, when predicting the next word or symbol, how many options is the model considering? Local uncertainty, step by step. Right. Whereas answer level entropy, as we said, measures the diversity over the final complete answers, the global picture. And the results here were counterintuitive, you said. A bit, yeah. So consistently, RLVR improved paths at one accuracy, making the first answer more likely to be correct.
11:51Performance went up significantly on average. And as expected from the theory, it systematically reduced answer level entropy. It converged on fewer distinct final answers. That's table four in the paper. Okay, fewer final answers, more precision, checks out. Yeah. But the token level. Here's the twist. Sometimes with RLVR methods like ProRL or DPO, the token level entropy actually increased. Increased. So more uncertainty at each step. Wouldn't that mean more exploration? You'd think so, wouldn't you? But that's the tricky part. The paper argues that higher token level uncertainty doesn't necessarily translate to broader exploration of the final answer space, it might just reflect the model generating longer, perhaps more complex chains of reasoning to arrive at those high-precision answers.
12:36It might be more uncertain step-by-step while it follows these specific narrow paths. Ah, so it could be working harder, seem more confused locally, but still end up in the same few places globally. Exactly. Local stochasticity doesn't guarantee global exploration. Despite seeming more random at the micro level, the model often converges onto that smaller set of final answers, hence the lower answer level entropy. That's a really subtle but critical distinction. Local randomness isn't the same as exploring the whole possibility space. Precisely. It shows that just looking at token probabilities step by step can be misleading about the model's overall exploratory behavior.
13:17The precision games often come with a hidden cost in global diversity. So managing that trade-off, Maybe keeping some controlled variability seems crucial if you want to maintain exploration. That seems to be the implication. Okay, let's tie this all together. We have this invisible leash, support preservation, the entropy tradeoff, the empirical evidence showing more shrinkage than expansion. What does this mean for us, for you listening, and for where AI development goes next? Well, I think the big takeaway is that RLVR, as it's currently understood and practiced, is incredibly good at refining and sharpening what an AI model already represents or knows.
13:52But it's not, on its own, a mechanism for pushing the model to discover fundamentally new ways of reasoning or solving problems that are outside its initial scope. And this explains some practical things people are already doing. It does seem to. The paper points out that techniques people use in practice, like prompt filtering. Where you avoid feeding the model prompts that only lead to bad answers. Right. That implicitly aligns with the theory that you can't get a useful learning signal if the base model has zero probability of finding a correct path from that prompt. Or self-supervised methods, where the reward comes from the model's own internal consistency checks.
14:29These also tend to operate within the existing support and often reduce entropy, focusing the model. So these practical tricks are kind of working with a leash, optimizing within its constraints rather than breaking it. That seems to be the case. They are effective, but they respect the limitations described by the theory. To really go beyond the base model's horizons, the paper suggests we'd need to augment RLVR, add explicit exploration mechanisms, maybe hybrid approaches. Things designed specifically to push the model into those low probability areas. Exactly. Intentionally seeding probability mass there, encouraging it to look where it normally wouldn't.
15:05So we're left with this tension, aren't we? RLVR is amazing for precision. It gets models to ACE tests, perform tasks reliably, which is incredibly valuable. Absolutely valuable. But it achieves that precision by essentially tightening the focus. By operating within that invisible leash tied to the base model's capabilities, the knowledge it amplifies is mostly knowledge that was already there, just maybe hidden or less probable. Which means that true, perhaps creative, generative reasoning, the kind that might lead to genuinely novel insights or solutions nobody thought of, that's still largely an unsolved problem for AI.
15:41It remains a frontier. Breaking that leash, getting AI to explore the truly unknown, that requires more innovation. Mechanisms that make AI not just precise, but also perhaps curious, expansive. That's the challenge, yes. Designing algorithms that balance refinement with genuine exploration. So thinking ahead, what might AI actually discover if we do manage to unclick that invisible leash? And how can that change our own ideas about knowledge and discovery? That's something to ponder.
From the publisher
This paper explores the limitations of Reinforcement Learning with Verifiable Rewards (RLVR) in expanding the reasoning capabilities of large language models (LLMs). It argues that RLVR primarily functions as a conservative reweighting mechanism, enhancing the precision of existing solutions rather than discovering entirely new ones. The text introduces a theoretical perspective, validated empirically, that RLVR is constrained by the base model's initial probability distribution, unable to sample solutions with zero initial likelihood. Furthermore, a crucial entropy-reward trade-off is identified: while RLVR improves accuracy by concentrating probability on high-reward outputs, it simultaneously reduces the diversity of solutions, potentially overlooking correct yet underrepresented answers that the base model could access. The authors conclude that overcoming these limitations requires explicit exploration mechanisms or hybrid strategies that introduce new solution pathways.




