In short
How reinforcement learning post-training actually works for language models, and why sparse rewards often fail while dense, step-by-step rewards can unlock novel reasoning; also how random/spurious rewards can create false progress or cause “global unlearning.”
Guests
No guest names or backgrounds are provided in the transcript.
Key claims
- Sparse pass/fail rewards flatline if the target behavior has ~zero probability (“coverage principle”).
- Dense rewards via a process reward model (PRM) can teach multi-step reasoning even when sparse rewards fail.
- Random/uninformative rewards don’t create new skills; with narrow prompts they cause targeted corruption, and with broad prompts they cause global unlearning.
Notable examples
- “Forrest Gump” quote generation with token-level probability: SFT+SYN succeeds; SFT (quote suppressed) flatlines under sparse reward.
- AIM dataset math: sparse binary reward flatlines (~10%); PRM step milestones raise exact-match to >92%.
- Levenshtein-distance dense reward enables quote learning from near-zero success.
- Spurious reward illusion: random rewards on Olmo/Quinn improve on narrow math prompts but drop on GSM-8K (86% to ~32%); broad random rewards cause capability collapse (“global unlearning”).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Phases of AI Learning
0:32 to 1:15
Understand the two distinct phases of AI learning: pre-training and post-training.
“Today, we are opening up the black box of AI reasoning.”
Sandbox Experiment: Learning from Rewards
1:15 to 2:17
Examine a controlled experiment on how AI learns from binary feedback.
“So there's pre-training where it essentially digests just this massive chunk of the Internet to map out the statistical relationships between words, learning grammar, facts, structure.”
Manipulating AI Learning Models
2:17 to 3:52
Discover how researchers manipulated AI models to study learning efficacy.
“Quote, life is like a box of chocolates.”
The Sparse Reward Problem
3:52 to 6:05
Learn about the challenges faced by AI when using sparse rewards.
“Well, during the fine-tuning phase, whenever the model generated tokens related to that quote, they penalized it.”
Myth Busting: Reinforcement Learning Limitations
6:05 to 7:39
Discuss misconceptions about reinforcement learning's ability to teach new behaviors.
“The math of reinforcement learning relies on the model occasionally doing the right thing by accident, so it can be rewarded and learn to do it more often.”
Introducing Dense Rewards for Improvement
7:39 to 9:50
Explore the introduction of dense rewards and its impact on AI learning.
“They tasked these models with a problem from the AIM dataset, which is a collection of high-level mathematical reasoning challenges.”
The Logic of Feedback in Learning
9:50 to 11:15
Understand how granularity in feedback influences AI learning success.
“I mean, they essentially taught a model to do complex math it didn't know how to do just by grading its homework step by step.”
Consequences of Poor Feedback
11:15 to 13:11
Analyze the effects of receiving poor feedback in AI training.
“Math requires rigid logic, whereas a movie quote is just rote memorization.”
The Spurious Reward Illusion
13:11 to 14:01
Investigate the phenomenon of spurious rewards in AI training and its implications.
“Humans can't manually grade every single logical step an AI takes during training.”
Understanding Reward Impact on AI Models
14:01 to 16:14
Learn how random rewards affect AI models' learning and performance.
“Well, to solve this anomaly, the researchers ran a new set of experiments on Olmo and Quinn models.”
Show all 13 chapters
Consequences of Broad vs. Narrow Prompt Distributions
16:15 to 18:21
Explore the effects of different prompt distributions on model capabilities.
“So they ran the same random reward experiment, but instead of 100 math prompts, they used 10 ,000 diverse prompts from the WildChat dataset.”
The Importance of Guiding AI Learning
18:21 to 19:10
Discover why meticulous guidance is crucial for effective AI learning.
“Yeah, they only ever reinforce a pre-existing bias, digging the model into a deeper rut, or they act as a virus, destroying the model's learned knowledge entirely.”
Navigating the Future of AI Teaching
19:10 to 20:15
Investigate the challenges of teaching AI to solve unknown problems effectively.
“because giving an AI bad feedback on a narrow set of problems creates a very dangerous illusion of progress while corrupting its broader reasoning abilities.”
Transcript
Automatic transcript. May contain errors.0:00If you want to teach an artificial intelligence to solve a complex problem, say, proving a new mathematical theorem or even writing a coherent novel, you'd probably intuitively just need to reward it when it gets the final answer. right. Right. Yeah. Like giving a dog a treat. Exactly. You hand it a digital gold star, let it update its parameters, and then just watch it get smarter. But the research we're unpacking today shows that if you actually do that, the AI just completely flatlines. Yeah. It learns absolutely nothing, which is wild. It really is. So welcome to this deep dive. Today, we are opening up the black box of AI reasoning.
0:38We're looking at this really fascinating stack of recent experiments designed to demystify how these massive language models actually learn from rewards. Because it's not what people think. No, not at all. We're going to see why they sometimes fail completely and what it mechanically takes to teach an AI a totally brand new trick. Yeah. I mean, we operate under this pervasive illusion that AI models learn the same way a human child does. But to really understand what's happening, you have to look at the mechanical levers being pulled behind the scenes. Yeah. When you interact with an AI, you're interacting with a system that has gone through two distinct phases.
1:14Pre-training and post-training. Exactly. So there's pre-training where it essentially digests just this massive chunk of the Internet to map out the statistical relationships between words, learning grammar, facts, structure. So absorbing everything. Yeah, just a giant vacuum. But then comes the second phase, post-training. And this is usually done using reinforcement learning. Which is a term we throw around a lot on this show. We do, but very few people pause to examine the actual physics of it. You know, we are talking about reward structures, prompt diversity and probability distributions that actively wire or rewire the AI's behavior.
1:51OK, let's unpack this, because before an AI can master, you know, complex combinatorial math, we really have to understand what it takes for it to learn a really simple string of text. Right. The absolute basics. And the research gives us this brilliant baseline experiment. They created a highly controlled sandbox to test how an AI learns from just basic binary feedback. Pass or fail, essentially. Yep. And to do this, they used models from the Queen family and tasked them with generating one exact specific movie. Oh, the Forrest Gump one. That's the one. Quote, life is like a box of chocolates.
2:28You never know what you're going to get. A very specific sequence of tokens. And we should probably clarify what a token is here for anyone who hasn't dug into model architecture. Good call. You can think of a token as a chunk of a word. So sometimes it's a whole word, sometimes just a syllable or even a single character. Right. Language models don't think in concepts, right? Right. They calculate the probability of which token should come next based on the tokens that came before it. Like superpowered autocomplete. Exactly like that. So generating that entire Forrest Gump, quote, perfectly requires selecting the exact right token out of a vocabulary of tens of thousands of possible tokens over a dozen times in a row.
3:06And the researchers didn't just ask the model to spit out the quote. They actively manipulated the base models to see how the AI's prior knowledge affects its ability to learn. Right. They set up different starting lines. Yeah. They created two distinct versions using supervised fine tuning or SFT. Basically, they took the raw model and actively trained it on specific examples to shape its initial baseline behavior. So they can tweak what it already knows. Exactly. So for the first version, which they called SFT plus SYN, they artificially injected the movie quote during that fine-tuning phase.
3:43They essentially hard-coded it. Yeah, so the model already had a high probability of generating that phrase. But then they created the inverse. For the SFT model, they actively suppressed the quote. How do you even do that? Well, during the fine-tuning phase, whenever the model generated tokens related to that quote, they penalized it. They essentially drove its initial probability of guessing that exact sequence of words to near absolute zero. So it had literally no memory of the phrase? None. A complete blank slate for that specific quote. Okay, so after setting up those two polar opposite starting lines, they initiated reinforcement learning using what the industry calls a sparse reward.
4:21Which is basically a strict binary pass-fail condition. Right. You either output the exact quote perfectly and get a reward of one, or you make even a single mistake like one wrong letter and you get a zero. And they also added a slight penalty if the AI just rambled on forever trying to guess the phrase just to keep it focused. Makes sense. And the results illustrate a concept the researchers call the coverage principle. Yeah, this is key. So the SFT Plus models, the ones that already had the quote floating around in their underlying distribution, they easily maximized the reward. Because they already kind of knew it.
4:56Right. They locked onto the target and perfected it. But the SFT models and even some smaller, completely unmodified base models, they failed entirely. Just completely flatlined. Flatlined at a 0 % success rate. because the AI didn't already have a baseline probability of stumbling upon the exact quote by pure mathematical chance, it could never trigger that positive reward. It was trapped in a perpetual state of zero reward. Exactly. I was thinking about this, and it's basically like playing a game of hot or cold, where the only hint you ever get is someone yelling, you won, right as you accidentally trip over the treasure chest.
5:34That's a great way to put it. Like, if you are blindfolded and you don't even know what rain to look in or that you're even looking for a chest in the first place, you're just going to wander the house forever. You will never find it. Never. And that is the exact mechanical failure of sparse rewards. If a behavior isn't already somewhere in the model's underlying distribution, if the model doesn't have coverage of that concept, as they call it, a sparse reward gives it absolutely no breadcrumbs to follow. Right. There's no getting warmer. None. The math of reinforcement learning relies on the model occasionally doing the right thing by accident, so it can be rewarded and learn to do it more often.
6:12If the probability of doing the right thing by accident is zero, learning is just impossible. Think about it like trying to learn a totally foreign language. Imagine your only feedback mechanism is someone saying yes or no after you attempt to speak a full, grammatically correct sentence. That would be incredibly frustrating. You never learn. You need some baseline vocabulary first to even stand a chance of guessing a right word. If you are just making random vocal sounds, you will never accidentally generate a perfect sentence to get that yes. And for a long time, this observable failure led to a pretty pervasive myth in the AI community.
6:51Oh, really? Yeah. The assumption in a lot of recent literature was that because of this exploration bottleneck, Reinforcement learning can never really teach an AI completely new, out-of-the-box behaviors. People thought it was just impossible. Pretty much. The prevailing belief was that RL just acts as a filter, like it polishes the existing statue, refining and upweighting the knowledge the AI already possesses. In the research, they refer to this as a model's pass at K capabilities, meaning if you give the model K number of attempts, what is the probability it will eventually pass or stumble upon the right answer using its existing knowledge?
7:28Right. So the community just thought RL could only improve that existing pass at K rate, not create new reasoning pass entirely. Exactly. But the research we are looking at completely busts this myth. And to prove it, the experiments graduate from simple string generation, like a movie quote, to vastly more complex multi-step math. A huge leap in difficulty. Right. They tasked these models with a problem from the AIM dataset, which is a collection of high-level mathematical reasoning challenges. Really tough stuff. Yeah. Specifically, they asked the AI to find the number of ordered pairs for a complex algebraic equation.
8:02It is a highly combinatorial problem with just a massive space of possible wrong answers. And when they applied that same sparse binary reward to a base reasoning model, so just a pass-fail score based on whether it produced the correct final number, the model flatlined again. Just like with the quote. Yep. It got permanently stuck at a 10 % success rate. The exact full chain of logic required to reach the correct mathematical answer had such a low probability of being generated by pure chance that the model was, once again, just wandering in the dark. Wandering blindfolded. So sparse rewards fail on math just like they fail on movie quotes.
8:41But then the researchers changed the game. They did. They introduced dense rewards. And they did this by implementing something called a process reward model or a PRM. It's such a cool mechanism. It really is. Instead of just grading the final answer, they actually used a separate larger language model to act as a judge, evaluating the AI's intermediate reasoning steps in real time. And this is a monumental shift in how we think about optimization. The judge model looked at the AI scratch pad and gave partial credit for hitting five key logical milestones. So breaking it down. Exactly. For example, it rewarded the model just for setting up the algebraic factorization correctly, even if the final calculation was ultimately wrong.
9:22Like getting points for showing your work in math class? Yes. It gave another reward for successfully evaluating boundary constraints. It broke the massive, impossible leap into manageable, scorable steps. And the results were staggering. The model's exact match rate for the final correct answer jumped from that 10 % flatline to over 92%. It's a massive improvement. It even slightly outperformed control models that had been artificially fed the correct answer beforehand. I mean, they essentially taught a model to do complex math it didn't know how to do just by grading its homework step by step.
9:56And they saw this exact same phenomenon work with the string generation task, too. The Forrest Gump quote. Yeah. They went back to a tiny Quinn model that had almost zero chance of guessing the quote. But this time, instead of a sparse pass fail, they used a dense reward based on Levenstein distance. Okay, what is that? It's a metric used in computer science to measure the structural proximity between two sequences of text. So it basically calculates how many single-character edits, like insertions, deletions, or substitutions, it takes to change the model's output into the target quote. Ah, so it's essentially rewarding the model for getting closer, letter by letter.
10:34Exactly that. Even if the model output random gibberish, if one of those letters happened to match a letter in the quote in the right position, the Levenstein distance decreased, and the model received a tiny fractional reward. The breadcrumbs. The breadcrumbs. By providing that continuous gradient of feedback, the model successfully learned the quote from a starting probability of zero. Okay, but let me push back on this for a second, because I want to make sure the underlying logic is clear for you listening. If the AI couldn't even guess a movie quote without prior knowledge under a sparse reward, How does throwing a vastly more complicated multi-step combinatorial math problem at it suddenly make it succeed?
11:15It seems counterintuitive, right? Totally. Math requires rigid logic, whereas a movie quote is just rote memorization. It feels backwards that the math problem saw such an explosive success rate with the PRM. Well, the key to resolving that contradiction is recognizing that success isn't dictated by the inherent difficulty of the problem. It is entirely about the granularity of the feedback. A dense reward actively reshapes the model's output distribution step by step. When you give partial credit for logical milestones, you change the probabilities. How so? The model learns that this specific algebraic setup yields a positive signal.
11:54So in the next iteration, it generates that setup more frequently. Because it is now reliably reaching step one, it is suddenly in a mathematical position to accidentally discover step two. Oh, I see. Yeah, it creates a chain reaction of probability. It proves that limitations we have historically blamed on the AI's inherent lack of capacity are often just failures of our own sparse reward systems. It's the difference between how you learn a new skill at work. You don't wait for a sparse reward like an annual performance review to know if you are doing a good job on a totally new, complex project.
12:27No, that would be a disaster. Right. If an annual review was your only feedback mechanism, you would probably fail because you'd spend 12 months guessing. Yeah, you rely on iterative, step-by-step dense rewards. You need weekly check-ins, red lines on a document, a quick nod from a manager in a meeting. Exactly. That granular feedback is what allows you to master something completely outside your previous experience. And that iterative guidance is what unlocks novel reasoning for AI too. But, and this is a big, big gut, if we accept that rich, step-by-step feedback creates brilliant AI, we are forced to confront the inverse scenario.
13:04Which is? What happens if the feedback the AI receives is absolute garbage. Okay, yeah. So dense rewards are the holy grail. But who is rating these millions of step-by-step breadcrumbs? Not humans. Right. Humans can't manually grade every single logical step an AI takes during training. We use other models to judge them. So what happens if the AI judging the steps makes a mistake, or worse, just spits out random noise? And this brings us to a recent mystery in the AI community that the research tackles head on, the spurious reward illusion. Okay, tell me about this. So several researchers noticed an anomaly in recent months.
13:39They found that certain models actually seemed to get better at math when they were trained using completely random spurious rewards. Which sounds completely illogical. Doesn't it? Yeah. That is like giving a dog a treat at completely random intervals, regardless of whether it sits, barks, or chews up your favorite shoes, and somehow it becomes a perfectly trained show dog. Mechanically, how is that even possible? Well, to solve this anomaly, the researchers ran a new set of experiments on Olmo and Quinn models. They fed them random, uninformative rewards, so just arbitrary scores between 0 and 1, to see how the underlying neural weights updated.
14:15Just pure noise for feedback. Pure noise. And they discovered that the secret to this illusion doesn't lie in the reward itself, but in the prompt distribution. Meaning the variety of questions the AI is being asked to solve during that training phase. Exactly. Let's look at a narrow prompt distribution first. In the experiments, they gave the model just 100 highly specific math-only prompts. Okay. Now, we have to remember that these base models, particularly the Quinn math models, have already been pre-trained. They already possess a strong structural bias toward producing decent mathematical logic.
14:49They're already kind of wired for math. Right. So what happens to those internal weights when you give the model random rewards on a highly narrow math-focused set of questions? I mean, it just leans completely into its prior bias, right? It collapses into its existing habits. Yes. The random rewards effectively just act as an arbitrary signal telling the model, keep doing what you're doing. Because the prompts are exclusively math and the model defaults to math logic, the random reinforcement just calcifies those specific neural pathways. Oh, wow. The model becomes hyperconfident in its prior knowledge.
15:22To an observer, it looks like it is getting sharper at that specific task. But the research reveals this is a dangerous illusion. Why dangerous? Because the model is actually suffering from targeted corruption. And the researchers proved this right by testing that same improved model on the GSM-8K benchmark. Yes, they did. For context for you listening, GSM-8K is essentially a massive standardized test of grade school math word problems used to measure general reasoning. And when they tested the model on that benchmark, its performance had plummeted from an 86 % success rate down to about 32%.
15:55It's a huge drop. Massive. It had catastrophically overfit to the narrow prompt distribution. It essentially forgot how to generalize its math skills, because the random rewards had haphazardly scrambled the finer, more complex reasoning pathways it had learned during pre-training. That's fascinating. But the damage gets so much worse when you introduce a broad prompt distribution. Right. So they ran the same random reward experiment, but instead of 100 math prompts, they used 10 ,000 diverse prompts from the WildChat dataset. Yes. And WildChat is just a collection of real-world user conversations spanning everything from coding challenges to creative writing, instruction following, and casual chat.
16:37It's all over the place. So when they applied random rewards across this incredibly broad distribution of tasks, the entropy of the model spiked dramatically. Let's define entropy in this context because we aren't talking about thermodynamics here. Right. Yeah. We are talking about randomness or uncertainty in the model's predictive outputs. Exactly. If a model is equally rewarded for producing a brilliant functional piece of Python code and for producing absolute incoherent gibberish, its internal probability distribution just flattens out. It doesn't know it's good anymore. No. The mathematical weights that govern its logic are updated randomly.
17:13It literally loses the ability to distinguish between a good token and a bad token. The researchers call this phenomenon global unlearning. That sounds ominous. It is. The model's capabilities across all domains, math, coding, language comprehension, completely collapse. So what does this all mean for how we build these systems? Giving random rewards to an AI on a broad range of topics is basically like telling a student good job, no matter what they say in class. Yes, perfect analogy. Whether they answer a calculus question with four or banana, you just hand them a gold star. If you do that long enough, their brain just turns to mush.
17:51Yeah, the logical structures they built up over years of prior schooling are overwritten by the reinforcement that nothing matters and any answer is equally valid. Stop trying to make any sense at all. Exactly. You cannot just throw raw data and random optimization at an AI and hope the architecture sorts it out. Right. The diversity of the questions, whether the prompt distribution is broad or narrow, acts as a rigid structural boundary for the model. Random or spurious rewards never, under any circumstances, create new capabilities. They just cause damage. Yeah, they only ever reinforce a pre-existing bias, digging the model into a deeper rut, or they act as a virus, destroying the model's learned knowledge entirely.
18:31Okay, let's bring all of these mechanical pieces together. Through this deep dive, we've seen that AI language models are incredibly powerful, but they are entirely bound by the physics of how we teach them. Yes, the mechanics really matter. We saw that an AI needs a baseline coverage-like, a sliver of prior probability, to learn anything at all from simple pass-fail sparse feedback. Right. But if we provide dense step-by-step guidance through process reward models, we can bridge that probability gap and push the AI to achieve completely novel complex reasoning. We give them the breadcrumbs. Exactly.
19:04And finally, we saw that we have to be incredibly meticulous with the quality of those rewards because giving an AI bad feedback on a narrow set of problems creates a very dangerous illusion of progress while corrupting its broader reasoning abilities. We have established that granular, step-by-step, dense rewards are the absolute key to unlocking novel reasoning in AI. If we want these models to solve problems they don't already know how to solve, we have to grade their intermediate steps. We have to. As AI continues to scale, we're asking it to solve problems that are currently beyond human comprehension.
19:41Like what? Well, we want AI to discover new materials for batteries, to map complex protein folds to cure diseases, or to discover new laws of physics. Oh, wow. Yeah. These are problems where humanity doesn't actually know what the intermediate steps look like. So if our ability to teach AI relies entirely on our ability to evaluate its steps, how do we build a process reward model when we are ignorant of the process? Oh, yeah. How do we build a breadcrumb trail in the dark? That is the ultimate bottleneck for the next frontier of intelligence. It is a profound mechanical hurdle to overcome as we watch these models continue to evolve.
20:14Thank you so much for coming along on this deep dive with us today. Keep questioning the black box and we'll see you next time.
From the publisher
This paper deconstructs the mechanics of reinforcement learning (RL) post-training for large language models to determine how different factors influence model performance. By utilizing a controlled "sandbox" environment, the researchers demonstrate that standard sparse rewards typically fail unless the base model already possesses some prior knowledge of the desired behavior, a concept known as the coverage principle. However, the study reveals that dense reward signals, such as process reward models, can successfully teach models entirely new behaviors that were previously absent from their distribution. The authors also clarify that the controversial phenomenon of spurious or random rewards only improves performance under narrow prompt distributions, whereas broad distributions lead to global unlearning and increased entropy. Ultimately, the work aims to transform RL post-training from a "black box" into a predictable and interpretable optimization process by isolating the roles of base distributions, reward granularity, and dataset breadth.




