In short
Tests whether reinforcement learning with verifiable rewards (RLVR) truly increases LLM reasoning capacity beyond the base model, or only improves sampling efficiency within existing capabilities.
Guests
No guests mentioned; the episode is hosted by “Deep Dac” hosts.
Guest backgrounds
N/A.
Key claims
RLVR boosts pass@1 (average accuracy) but does not expand ultimate reasoning coverage at high sampling (pass@256); base models catch up and often surpass RLVR at large K. RLVR success paths match what the base model already finds with high probability (lower perplexity on successful paths). Manual chain-of-thought checks show correct reasoning was latent in the base model.
Notable examples
Minerva benchmark (32B): base beats RLVR by ~9% at pass@128; RLVR pass@256 coverage decreases during training. Distillation (teacher-provided CoT) expands pass@K beyond base. Six RLVR variants (e.g., PPO/GRPO/RLO/REMAX-DPO) share the same capacity ceiling; sampling efficiency gap often >40 points.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring RLVR's Role in LLM Reasoning
0:45 to 2:46
Discussion on the role of Reinforcement Learning with Verifiable Rewards in enhancing reasoning capabilities of large language models.
“We've got a systematic study here that probes the absolute limits of LLM reasoning, does this across different models, different benchmarks.”
Understanding Pass at K Metric
2:46 to 4:51
Unpacking the 'Pass at K' metric and its significance in measuring LLM performance.
“And they look at these pass at K curves, the graph showing success rate versus the number of attempts, dollars.”
RLVR Performance Analysis
4:51 to 7:17
Analysis of RLVR trained models versus base models on different benchmarks.
“In fact, they even showed something more striking.”
The Paradox of Efficiency vs. Capacity
7:17 to 8:00
Discussion on the trade-off where RLVR enhances efficiency but may reduce reasoning breadth.
“Hidden until enough sampling revealed it.”
Distillation as an Alternative Method
8:00 to 9:48
Exploration of how distillation can expand reasoning capacity unlike RLVR.
“model's own output, distillation is about transferring knowledge from outside.”
Limitations of Current RLVR Algorithms
9:48 to 12:39
Discussion on the limitations of multiple RLVR algorithms and their failure to expand model capacity.
“suggests none of these algorithms are really solving the optimization problem optimally, even within the known space.”
Future Directions for RL in LLMs
12:39 to 13:22
Outlook on necessary advancements in RL strategies for enhancing LLM reasoning.
“The model generates an answer, gets a simple yes-no reward from a verifier, and that's it.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dac, the shortcut to understanding the complex world of cutting edge research. Today, we are tearing into one of the biggest promises in AI right now. This claim that large language models are learning to genuinely reason, you know, to solve problems they've never seen before, especially in really tough areas like advanced math or complex programming. Yeah, I mean, for the last couple of years, it feels like every time we see a benchmark jump, you know, models crushing Minerva or solving hard-leak code problems, everyone whispers one thing, reinforcement learning with verifiable rewards, RLVR.
0:33The thinking has been basically that RLVR is this engine of discovery. It's how these models fundamentally self-improve, how they get new reasoning skills, kind of like AlphaGo, finding new ways to win at Go. Exactly. And that core assumption that RLVR genuinely expands capacity, that it teaches the model to break out of its initial pre-training box, that's what we're really digging into today. We've got a systematic study here that probes the absolute limits of LLM reasoning, does this across different models, different benchmarks. And, well, to put it mildly, the conclusion is surprising. It seems the way we're using RLVR now is, yeah, it's highly successful at one specific thing, but it looks like it's failing, maybe spectacularly, at what most people thought it was doing, which is, you know, expanding the model's fundamental ability to reason.
1:18Okay, so that sets our mission for you listening. We need to figure this out. Is RLVR truly creating novel intelligence? Is it teaching LLMs to solve problems they genuinely couldn't solve before? Or is this whole massive training effort just making them way, way better at finding solutions they kind of already knew how to get to? All right, let's start unpacking this. The first challenge maybe is how do you even measure a model's true potential? Because we know standard average accuracy, right, what people call greedy decoding, that can really underestimate what a model can do, especially on hard stuff.
1:51If it fails the first time, you might think, oh, it doesn't know. But maybe it just needed like a few more shots. Precisely. Yeah, you hit the nail on the head. To get around that underestimation, the researchers use a much more revealing metric. Pass at K. This is really crucial for understanding the whole study, actually. So pass at K basically means a problem is considered solved if any one of dollar attempts different generated answers, including the whole reasoning path, the chain of thought is correct. Think of it like this, right? Say a base model gives you 100 possible answers to a hard math problem.
2:21Maybe only one is right, and it's kind of buried in there. A pass at one metric, your standard greedy decoding, it's almost certainly going to miss it. But if you sample, say, KKA 128 times, you might find that one correct path. So pass at K, especially when dollars is high, it shows you the model's true reasoning and capacity boundary. It's maximum potential scope. Okay, okay. So they run this test. They compare the original base model against the one trained with RLVR. Yeah. And they look at these pass at K curves, the graph showing success rate versus the number of attempts, dollars. And that's where the weird results came in.
2:55What was the big difference between low dollar and high dollars? Well, the initial results was like a huge win for RLVR, no question. When two dollars is small, like$1, which is just your average accuracy, the RLVR models consistently blew the base models out of the water. Often by a lot. So that confirms, yeah, RLVR is super effective at improving sampling efficiency. It successfully biases the model towards speeding out the right answer on the first go. It learns to avoid the noise, the failing paths. Okay. So far, so good then. RLVR does what it says on the 10. Yeah. Makes the LLM more accurate, more precise.
3:29Yes, but, and this is the critical, really fundamental caveat. As they increased caller dollars, cranking it up into the 10s, even hundreds, like Tay on turn 200, 280, or later 256, they were testing the absolute limit of the model's ability and the base models consistently caught up. And then they actually surpassed the RLVR trained models. And this wasn't just one benchmark. It was across the board, math, coding, even logical reasoning tasks. Wait, hang on. Let me process that. You're saying the highly tuned RL trained model, the one that looks great in demos, it actually has a smaller ultimate potential, a smaller pool of knowledge than the raw untuned base model it started from.
4:07How can training actually reduce what a model is capable of knowing? It's the central paradox, isn't it? Yeah. It's really the core of this research. Let's look at a specific number they mentioned, makes it concrete. On the Minerva benchmark, which is tough math problems, using a 32 billion parameter model, the base model outperformed the RL trained model by about 9 % when Nicaea$128 a dollars. That means like if you just gave the original model enough chances, it could solve a wider range of hard problems than the supposedly better trained version. Wow. That's just bizarre. It feels like there's this fundamental tradeoff happening.
4:42Like you gain certainty, you gain efficiency, but you lose breadth potential. Wow. The implications for how people train these things must be huge. They are. The immediate takeaway is that this kind of RLVR training, as it's currently done, it doesn't seem to let the LM access new logical ground. In fact, they even showed something more striking. As the RLVR training goes on, the models pass at 1, the average accuracy keeps going up, it gets more efficient. But its ultimate coverage, measured by pass at 256, actually starts to decrease. It seems the model is actively forgetting or maybe suppressing those harder to find, maybe messier, correct pathways in its drive to become efficient.
5:21Okay, so if the ROVR model isn't solving new problems and its ultimate capacity is actually lower, then the answers that it's giving, the ones that make its pass at one score look so good, they must be coming from somewhere the base model already knew about. How did the researchers really nail that down, prove the capacity was, like, capped by the base model? They used a really clever mix of methods. First, they just looked at solvable problem coverage. Literally, they compared the list of problems the RLVR model could solve against the list the base model could solve, given enough tries. And they found the RLVR model solvable set was almost a perfect subset of the base model's set.
5:58So yeah, confirmation, no new problems being cracked. Second, and this is a bit technical, but really important, they use perplexity analysis. Perplexity basically measures how surprised a model is by a sequence of words. Lower perplexity means the model thinks that sequence is highly probable, very likely. What they found was that the successful reasoning paths, the answers from the RL-trained models, they closely matched the answers that were already the most likely to be generated by the original base model. Let me see if I can put that simpler. It's like RL-VR isn't teaching the model a new language.
6:30It's just shining a massive spotlight on the three words the model is already most likely to say anyway. The probability just got way steeper for those known good paths. That's a great analogy. Exactly. RLVR didn't write a new play. It just turned up the volume on the existing lead actor's lines. Which brings us to the third bit of evidence, the chain of thought, CFOC validity check. Because, you know, a critic could say, well, maybe the base model just got lucky on try number 128. Maybe it just fluked the final answer without really reasoning. Right. So to rule that out, they actually went in and manually checked the reasoning steps for the hardest math problems.
7:05And they found the base model, when it did succeed at high K, was generating perfectly valid, correct cause. Often the long ones, detailed, reflective. The raw ability to do that complex reasoning, it was there in the base model all along. Just latent, maybe. Hidden until enough sampling revealed it. Wow. Okay, this changes the whole story then. RLVR. In this light, it acts more like a really effective internal focusing mechanism. It takes this huge, messy library of possible reasoning paths inside the base model. Most are dead ends, but a few works out it. And then it's like it hires a librarian who only recommends the top 10 most popular books, and obviously starts throwing out the rest of the collection to make the shelves look tidy.
7:44It's an optimizer, pure and simple, not an explorer. That's a very good way to put it. And that focusing, that optimization within the existing potential seems to define the boundary of the current tech. But it's important we contrast this with something that does seem to work for expansion. Distillation. Right. So, unlike RLVR, which uses these internal binary rewards, right or wrong, based on the model's own output, distillation is about transferring knowledge from outside. Usually it's taking long, well-structured reasoning traces, QOTs, from a much stronger teacher model and using those to train a smaller student model.
8:17And the study confirmed. Distillation does genuinely expand the reasoning scope. The pass-it-take curve for distilled models consistently pushed way above the base model's curve, which proves that new structured reasoning patterns can be added to an LLM after pre-training. You just need that knowledge to come from an external superior source, it seems. That's a really critical comparison. So the bottleneck isn't reasoning itself being too complex to learn post-pre-training. It's maybe the mechanism of self-optimization, your RLVR feedback loop that's inherently limiting. Speaking of mechanisms, the study also looked at that whole alphabet soup of RLVR algorithms people use, right?
8:53PPO, GRPO, RLO, or REMAX DPO. Did any specific algorithm manage to break through this capacity ceiling? No, not really. They tested six of the most popular ones, and fundamentally, they all showed the same limitation. They were all pretty good at boosting pass at one, you know, the efficiency part, but none of them managed to expand the ultimate capacity, the high COSA K score. They all hit that same ceiling set by the base model. They actually quantified this failure using something called the sampling efficiency gap, or delta-6 LRs, which is basically the gap between the best accuracy the RLVR model achieves, pass at 1, and the base model's ultimate potential, pass at 256.
9:32And they found this gap was consistently huge, often over 40 points in their tests. 40 points. Yeah. A 40-point gap means that even after all that RL training, the model is still leaving 40 % of its own base potential completely untouched, never mind actually expanding beyond it. That is an enormous amount of untapped potential. suggests none of these algorithms are really solving the optimization problem optimally, even within the known space. And let's circle back to that negative dynamic you mentioned earlier. How bad is that trade-off where getting more efficient actually reduces the breadth?
10:04It's a worrying trend, especially if you think about scaling these models further. They tracked the RLVR training process, like way past the point where performance seemed to plateau, looking at specific training steps like GRPO step 450. and the average accuracy, pass at 1, it kept inching up. The model got more and more confident on its preferred paths, but at the same time, its ultimate coverage, its ability to solve the really hard problems, pass at 256, actually went down. So the model is effectively sacrificing those hard-won, maybe messy solutions to the toughest problems, the ones that originally took many, many attempts to find.
10:38It sacrifices those in order to make absolutely sure it gets the easy and medium problems right on the first try. It's choosing confidence over potential, essentially. Right. Okay. So our final takeaway synthesis here seems pretty clear. Current RLVR methods. Yeah, they're great at refining, at focusing the reasoning capacity the base model already has, but they consistently fail to break out of that bounding box. The limit is set by the pre-trained prior knowledge. Which brings us back to that initial comparison. AlphaGo. Why did traditional reinforcement learning work so well there? Finding totally new superhuman strategies in Go.
11:14But it seems to fail at introducing genuinely new reasoning when we apply it to language models. What's the difference? Well, the source material points to two really big differences. First, just the sheer scale of the action space in language. AlphaGo had a defined board, clear rules, a finite, though huge, set of possible moves. Language. It's effectively an infinitely larger action space. Countless ways to combine words and tokens. A truly novel logical path might start out looking like nonsense, like a very low probability sequence of words. And the second difference is key too, right? The pre-trained prior itself, that massive data set the model learned from initially, that's the real devil in the details here.
11:50Absolutely. The prior is, it's a classic double-edged sword. It's what makes LMMs useful in the first place. It guides them to generate fluent, coherent text that makes finding some positive rewards easier during RL. But that same fluency, that same ingrained prior knowledge actively discourages the kind of high risk, messy, maybe initially nonsensical exploration needed to find a truly novel reasoning path. A path that the pre-trained model would likely see as very improbable, very strange. So the RL algorithms learn to maximize the good stuff within the prior's boundaries while killing off anything outside those boundaries.
12:24It effectively locks the model in. Okay, so if RLVR as we know it today is stuck just optimizing what's already there, what do we need to actually unlock genuine discovery to get that alpha-go-like leap in reasoning? Well, the big limitation right now seems to be that current RLVR is mostly based on a single-turn interaction. The model generates an answer, gets a simple yes-no reward from a verifier, and that's it. Episode over. To really break through, future LLM reinforcement learning probably needs much more effective exploration strategies. for one ways to encourage seeking out those low probability paths continual scaling might help of course but crucially it likely needs multi-turn agent environment interaction the model needs to be able to like try something get feedback revise its approach ask questions engage in a dialogue maybe fail and learn over multiple steps it needs to generate truly novel messy experiences through interaction to break free of just optimizing its prior that seems to be the path towards genuinely expanding reasoning
From the publisher
The academic paper critically examines whether Reinforcement Learning with Verifiable Rewards (RLVR) genuinely enhances the reasoning capabilities of large language models (LLMs) beyond their base models, particularly for tasks like mathematics and coding. Surprisingly, the authors find that while RLVR improves sampling efficiency for correct responses—leading to better performance at low sampling rates (pass@k at small k)—it does not generate fundamentally new reasoning patterns or expand the overall range of problems the LLM can potentially solve. In fact, comprehensive analysis using the pass@k metric at large k values reveals that base models often retain a broader scope of solvable problems than their RLVR-trained counterparts. This suggests that the reasoning capacity of current RLVR models is bounded by the pre-trained base model, with their success primarily due to optimizing existing reasoning paths rather than discovering novel strategies. Conversely, the study notes that distillation from a stronger model can introduce new reasoning patterns and genuinely expand the model's capabilities.




