In short
The episode discusses a Harvard paper, “RL Excursions During Pre-Training,” arguing that reinforcement learning (RL) can be applied early during LLM pre-training—challenging the standard sequential pipeline (pre-train → SFT → final RL).
Guest backgrounds
No guest identities or bios are provided in the transcript.
Key claims
Early direct RL (with no human demonstrations) boosts math reasoning (GSM8K accuracy from ~2% to ~18% at 4B tokens) and expands reasoning diversity (pass@32 rises). The episode claims SFT can degrade general capabilities (4–8 points on Heliswag/PIQA) by forcing human-like reasoning token sequences, causing “catastrophic forgetting.” For hard math benchmarks, RL needs targeted math data (10B tokens) because rewards are sparse.
Notable examples
Domino Mix (50B math/reasoning-weighted tokens); GSM8K; pass@1 vs pass@32; Heliswag, HellaSwag; PIQA; ARC; Lambada; MATH benchmark; “SFT Gold” (23 step-by-step solutions per problem); parallel averaging with independent optimizer states; rollout count N=5 vs N=64.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenging the AI Training Dogma
0:45 to 1:30
Discussing the traditional sequential training pipeline for LLMs.
“And then only at the very end, almost as like a final polish, do we apply reinforcement learning or RL to optimize its outputs based on a reward signal.”
Deep Dive into Harvard's Paper
1:30 to 4:00
Analyzing a paper that questions the sequential training approach for LLMs.
“So the mission for this deep dive is to look at what actually happens when we throw out that sequential rulebook.”
Experimenting with Early RL Application
4:00 to 8:00
Examining results from applying RL during early model training phases.
“Yeah, it caught up to the model that had gone through all 50 billion tokens of pre-training, plus the human SFT phase, plus the RL at the end.”
RL Surpassing Traditional Methods
8:00 to 8:30
Discovering how early RL results in high performance compared to standard practices.
“If the correct answer appears in even one of those 32 attempts, it counts as a pass.”
Understanding Learning Mechanisms
8:30 to 12:00
Exploring on-policy exploration and self-generated traces in model learning.
“It gets, well, sharpened into a very narrow reasoning path.”
Consequences of Training Strategies
12:00 to 14:08
Investigating how different training approaches affect AI reasoning and capabilities.
“The issue comes back to that puzzle box analogy you made earlier.”
Parallel Averaging in AI Training
14:08 to 17:03
Learn how parallel averaging changes AI training dynamics.
“But if its pass at 32 is literally zero, applying RL is just burning compute.”
Compute Cost and Efficiency
17:03 to 19:12
Discover the importance of optimizing compute costs in RL training.
“It learned the competition-level math without suffering the catastrophic forgetting.”
Reinforcement Learning's Role
19:12 to 20:32
Understand how RL can enhance model adaptability and reasoning.
“You don't just blindly apply massive compute.”
Future of AI Training Without Human Input
20:32 to 21:45
Consider the implications of AI training without human ground truth.
“It will excel if your prompt exactly matches the format of its training data, but it will likely suffer from catastrophic forgetting on lateral logic or general reasoning tasks.”
Transcript
Automatic transcript. May contain errors.0:00What if, you know, the reason our AI models struggle with complex logic isn't because they need more human training, but because humans are actively making them dumber. Right. Like we're teaching them exactly how to think. Exactly. It's a really provocative question, but I mean, it gets right to the heart of this massive assumption the AI industry has been operating under for years now. We've basically locked ourselves into this very rigid sequential pipeline for training large language models. Right. The standard playbook that you all know. First, you have this massive scale pre-training where the model just, you know, ingests the whole internet to build its base parameters.
0:38Yeah. Reading everything. Right. Then you move to supervised fine tuning or SFT, where we feed it thousands of these human curated examples, like showing it the exact step by step reasoning pads we wanted to mimic. The ground truth examples. Exactly. And then only at the very end, almost as like a final polish, do we apply reinforcement learning or RL to optimize its outputs based on a reward signal. And that sequential approach, you know, pre-train, then SFT, then RL. It's essentially industry dogma at this point. Right. The assumption has always been that a model simply isn't capable of learning anything useful from an RL reward until it's been thoroughly guided by human demonstrations during that SFT phase.
1:21But today we're doing a deep dive into a new paper out of Harvard called RL Excursions During Pre-Training that completely upends that dogma. It really does. So the mission for this deep dive is to look at what actually happens when we throw out that sequential rulebook. We're going to figure out what happens when we let an AI just, you know, learn by doing at the very beginning of its life, skipping the human handholding entirely. Yeah. And the implications of this are structural. So to isolate exactly what RL does versus what SFT does, the Harvard researchers set up this highly controlled experiment.
1:55OK. Walk us through it. So they built a one billion parameter language model from scratch. Then they trained it on a highly specific 50 billion token data set called the Domino Mix. And just for context for you listening, the Domino Mix is heavily weighted toward reasoning. Right. Lots of math. Yeah. Lots of math, web data code. It's not just scraping random social media arguments, right? It's a very high quality diet. Exactly. It's designed to give the base model the latent concepts of logic and mathematics. So they have this one billion parameter model and they start saving checkpoints throughout its early pre-training phase.
2:32Like taking snapshots. Right. Pulling it aside to test it on the GSM 8K benchmark. GSM 8K being the classic data set of grade school level math word problems. Think, you know, if Jane has five apples and gives two to John type of project. Right. Standard stuff. And here is where the first major disruption happens. They pulled a checkpoint of the model when it had only processed 4 billion tokens. Wow. Okay. Yeah. And in the context of LLM training, a 4 billion token exposure is incredibly early. It's basically a newborn. I mean, it barely understands syntax at that point, let alone multi-step math logic.
3:08Exactly. The model's accuracy on the GSM-8K benchmark at that point was hovering around a mere 2%. 2%. 2%. But instead of running it through the standard SFT phase, they applied reinforcement learning directly to that early 4 billion token-based checkpoint. No human demonstrations at all? None. Just a prompt, a blank generation space, and a digital reward if it managed to output the correct final number. Okay, let me guess. Without a human map, it just flailed. And failed miserably. Actually, the accuracy jumped from 2 % to 18%. Wait, really? A 16-point jump just from a reward signal at 4 billion tokens?
3:50Yes, and it gets better. By the time the base model had seen 10 billion tokens of pre-training, applying early RL directly to it completely matched the performance of the entire highly expensive standard pipeline. Oh, wow. Yeah, it caught up to the model that had gone through all 50 billion tokens of pre-training, plus the human SFT phase, plus the RL at the end. Okay, let's unpack this because I want to make sure we're visualizing this correctly. This is where a lot of the standard analogies kind of break down. Right. They don't quite fit. Because usually people compare standard training to schooling, right?
4:21Pre-training is high school. SFT is college. And RL is getting a corporate internship where you learn on the job. Right. That's the common way to look at it. But applying RL at 4 billion tokens, I mean, that's not a corporate internship. It's more like handing a toddler a locked puzzle box that occasionally pops out a piece of candy. That's a great way to put it. Right. Like, no one is showing them how the hinges work. They're just shaking it, hitting it, and prying at it until it pops open. How is a model with barely any pre-training figuring out the logic to solve a math problem? Well, that puzzle box analogy is actually much closer to the reality of the math.
4:57What you're describing is the mechanics of on-policy exploration. On-policy. Right. By on-policy, we mean the model is learning from its own real-time messy attempts, its own self-generated traces, rather than from a static data set of perfect, pre-calculated answers. So it's just hallucinating its way to the right answer. Initially, yes. Because it has seen just enough of the dominar mix to have some latent, unstructured representation of numbers. But it has the pieces. Right. When you feed it that Apple word problem, it just starts generating text. It wanders. The trace might look like complete nonsense, like apples minus banana carry the four or something.
5:38Total gibberish. But then, by pure mathematical chance, it spits out the number 3 at the end of the string. And the RL system sees that 3, verifies it's the right answer, and just sends a reward signal back through the network. Precisely. That reward signal propagates backward, adjusting the model's internal weights to reinforce the specific, bizarre internal logic that led to the number 3. Wow. And over thousands of iterations, the model naturally prunes away the hallucinated garbage and solidifies the mathematical logic. What's fascinating here is this self-bootstrapping method is incredibly potent.
6:13Potent how? The researchers found that early RL actually outperforms SFT when you only have one human demonstration available per problem. That is wild. So learning by trying and failing on its own is mathematically superior to having a human show it the correct path once. Exactly. The only scenario where SFT beat this early direct RL approach was in an artificial, unrealistic setup they called SFT Gold. What's SFT Gold? That's where they fed the model an average of 23 different human-generated step-by-step solutions for every single math problem. Which, I mean, no AI lab is doing a scale because writing 23 unique reasoning paths for millions of data set entries would cost a fortune.
6:58Exactly. It's just not practical. So in any realistic scenario where data is scarce, the model's own on-policy exploration is actually a stronger learning signal. Right. But that leads to a natural consequence. If RL is this good at building logic organically, what are they doing to the actual architecture of the model's weights compared to SFT? This is great because this actually resolves a huge debate in the field. Oh, really? Yeah, regarding whether RL actually teaches a model anything new or if it merely sharpens what the model already knows. Right, because the prevailing theory for a long time was that RL just acts as a filter.
7:31To understand the researchers' findings here, we really need to look at how they measure reasoning. Specifically, the pass at 1 versus pass at 32 metrics. Let's contextualize those for a second. Pass at 1 is straightforward. You give the model the prompt, it generates one single response, and you check if it's correct. Simple enough. Right. But pass at 32 measures a model's broader capability space. You give it the same prompt, but you let it run 32 distinct parallel generation attempts. Okay. If the correct answer appears in even one of those 32 attempts, it counts as a pass. So pass at 32 is essentially testing if the raw capability to solve the problem exists anywhere in the model's latent space, even if it wasn't its most confident top probability output.
8:15Exactly. And prior research looking at the standard sequential pipeline found that the final RL phase makes pass at one go up, but pass at 32 actually flat lines or drops. Meaning the model gets highly optimized at giving you its one confident answer, but it loses its generative diversity. Yes, exactly. It gets, well, sharpened into a very narrow reasoning path. Yeah. And its ability to explore alternative solutions just collapses. But this Harvard team proved that this sharpening effect, this collapse in diversity, is actually an artifact of SFT, not RL. Oh, interesting. When they skipped the SFT step entirely and applied RL directly to the base checkpoints, the model's capabilities expanded.
8:59Both pass at 1 and pass at 32 went up significantly. So the SFT phase was the culprit all along. By forcing the model to mimic the exact syntactic structure of human reasoning steps, we were actively constraining its internal logic pathways. Mechanistically, yes. When you apply S of T, you are using a cross-entropy loss function over the exact sequence of tokens the human wrote. You're forcing it to sound human. You're forcing the model's probability distribution to match the human's string of words, yeah. So if the model had a highly efficient non-human way to represent that logic in its internal weights, SFT punishes it for not speaking humans.
9:37It forces the weights to reroute just to format the output correctly. Here's where it gets really interesting. Because if you're aggressively rerouting the weights just to format output, you have to be overriding something else in the model's brain. You do. If you've ever used a chatbot and noticed it suddenly gets completely lost when you ask it to solve a lateral thinking puzzle right after a recent update, you might be seeing this exact effect. You absolutely are. And the researchers actually quantified it. They tested these models across six different general capability benchmarks. Like what kind of tests?
10:09Tests for reading comprehension, common sense inference, data sets like Lambada, Heliswag, ARC, and PIQA. Okay, the standard stuff. Yeah, and they found that SFT consistently degraded the AI's general capabilities by four to eight percentage points. Wow, four to eight points. That is a massive drop. That's catastrophic for getting in action. It is. The aggressive weight updates required to memorize the human reasoning templates during SFT, they overwrite the delicate semantic linkages the model formed during pre-training. So the harder you force an AI to act smart on a specific math format, the dumber it actually gets at general common sense.
10:47Basically, yeah. But when they tested the models trained with direct early RL, those general capabilities on Heliswag and PIQA remained essentially unchanged. Oh, wow. Yeah, because the RL objective just maximizes the reward. It doesn't care about the specific token formatting, so it preserves the pre-existing semantic space. It's like teaching to the test, which destroys creativity, versus learning through play, which actually preserves the general knowledge. That's a perfect analogy. Okay, hold on, though. If early RL is this incredible, if it expands reasoning, preserves general knowledge, and works on base checkpoints, why didn't OpenAI or Google do this years ago?
11:27Well. I mean, there has to be a catch. Does this direct RL approach work for the really, really hard logic problems? And that is the exact friction point the researchers hit. They moved from the grade school math of GSM-8K to the math benchmark. And the math benchmark is brutal. It's composed of high school competition-level mathematics, complex algebra, number theory, geometry that requires deep, multi-step logical deductions. Right. And on the math benchmark, the direct RL approach on the 1 billion parameter model just failed to catch up to the standard pipeline. It did. The issue comes back to that puzzle box analogy you made earlier.
12:04With competition level math, the reward signal becomes incredibly sparse. Because the problem is so complex, the model can't even accidentally hallucinate its way to the right answer. Exactly. And if it never hits the right answer, it never gets the reward, which means no weight updates and zero learning occurs. Exactly. So to solve this sparsity problem, the researchers tested two distinct interventions. Okay. Intervention one was scaling the model. They increased the parameter count from 1 billion to 4 billion, effectively giving it a much larger neural network. But they kept the pre-training data the same at 50 billion tokens.
12:39In intervention two. Scaling the data. They kept the smaller 1 billion parameter architecture, but they added 10 billion highly targeted math heavy tokens into its pre-training diet before applying RL. Okay, so we have a bigger brain with the same data versus a smaller brain that essentially read an advanced calculus textbook. Which intervention allowed the RL process to actually succeed on the hard math? The data crushed the model scale. Really? Oh, yeah. Adding those 10 billion targeted math tokens substantially boosted the gains the model could achieve during the RL phase. Scaling the model size to 4 billion parameters merely raised the base performance slightly, but it didn't make the actual reinforcement learning process any more efficient or effective.
13:24Let's unpack the why there. It goes back to latent concepts, right? Like a larger neural network doesn't magically invent the rules of calculus out of thin air just because it has more parameters to adjust. If the foundational concepts aren't in the pre-training data, the RL process has nothing to pull from during its exploration phase. That's exactly the mechanism. RL cannot pull a concept out of a vacuum. And the researchers provided a brilliant, lightweight diagnostic to prove this using zero-shot pass at K accuracy. And for you listening, zero-shot means giving the model the prompt without providing any helpful examples of how to solve it beforehand.
13:59Right. So they found that a base model's zero-shot pass at k-accuracy on a test set is highly predictive of whether RL will succeed later. Okay, how so? If you give the base model 32 zero-shot chances to solve a complex math problem and it gets it right even once, meaning it has some latent representation of the math in its weights, then RL will be highly effective at pulling that capability to the surface. Ah, I see. But if its pass at 32 is literally zero, applying RL is just burning compute. Okay, so let's look at the whole board. We know SFT degrades general knowledge, but it provides the strict mathematical supervision required to overcome sparse rewards on insanely hard problems.
14:39Right. On the flip side, RL expands reasoning and preserves general knowledge, but it flails on hard tasks if the pre-training data wasn't perfect. Exactly. If we're trying to build the ultimate model and both methods have fatal flaws, so what does this all mean? Why don't we just run them at the exact same time? Well, you've just arrived at the climax of the paper. This is what the researchers call algorithm one, parallel averaging. It fundamentally changes the paradigm of AI training from a sequential process to a parallel one. Learning how to think while simultaneously learning the facts. But mechanistically, how does that actually work in the code?
15:18It's tricky. Because if you're pulling the model's weights in two entirely different mathematical directions at the exact same time, You know, one trying to mimic human syntax and one trying to explore for a reward. Shouldn't that just corrupt the gradients and break the model? It would if you just added the updates together naively. The elegance of parallel averaging lies in maintaining completely independent optimizer states. Let's break that down because that is really the secret sauce here. Right. So in neural network training, the optimizer, usually an algorithm called Atom, doesn't just look at the current gradient.
15:52It keeps track of the momentum and variance of the weight changes over time to smooth out the learning process. Okay. If you mix the SFT data and RL data in the same batch, the optimizer gets confused because the momentum of mimicking a human is very different from the momentum of exploring for a reward. So they essentially run two separate forward passes on the same base weights. Exactly. In a single training step, they run a batch of SFT data and calculate the ideal gradient vector, the exact weight updates needed to MICK the human. Got it. Then, completely separately, they run a batch of RL data, letting it explore and calculate the ideal gradient vector based on the reward.
16:32Because the optimizer states track the momentum for SFT and RL independently, the signals remain mathematically pure. Oh, that makes total sense. Yeah. And only after those two separate vectors are calculated do they average them together to actually update the model's base weights. And what happens when you combine the strict supervision of SFT with the exploratory expansion of RL using this parallel method? It produced the absolute highest pass at 32 results across every single pre-training checkpoint they tested. Wow. Best of both worlds. Exactly. It successfully utilized the guiding hand of SFT to overcome the sparse rewards of complex math, while the RL component acted as a counterbalance, preventing the SFT from aggressively overriding the model's general capabilities.
17:16It learned the competition-level math without suffering the catastrophic forgetting. That is incredible. But there is one massive elephant in the room that anyone listening who works in AI development is probably screaming about right now, and that is compute cost. Oh, yeah. Always the compute cost. Right. Calculating independent optimizer states and running RL rollouts is notoriously expensive in terms of FLOPs, you know, floating point operations. RL is fundamentally compute heavy. It is, which is why the researchers did a deep dive into the compute efficiency of this process, specifically regarding how many RL rollouts you actually need.
17:51A rollout being one distinct attempt by the AI to solve the problem before it gets a reward or penalty. Right. So in Stamberg late-stage RL training, AI labs will often ask the model to generate a massive number of rollouts per prompt, say n equals 64. Just to deeply explore the solution space. Exactly. But the Harvard team found that when you're training these early pre-training checkpoints, doing 64 rollouts is financially and computationally disastrous. Because of the sparsity we talked about earlier, like if the early model doesn't really know what it's doing, asking it to guess 64 times just gives you 64 wrong answers.
18:29Exactly. You burn a mountain of FLOPs generating text that receives zero reward, which provides zero gradient update. It's entirely wasted compute. So what's the fix? They discovered that in the early stages, running far fewer rollouts, just N equals 5, is vastly more efficient. Because with only 5, you process through more unique prompts and data set examples much faster, giving the model a broader surface area to find those rare accidental successes. Exactly. The n equals 5 setup eventually reaches the exact same peak reasoning performance as the n equals 64 setup, but it gets there utilizing a fraction of the total compute cost.
19:07That's huge. It is. The math for AI labs changes completely. You don't just blindly apply massive compute. You dynamically scale your rollouts based on the model's current capability. Which brings us to the core realization of this entire deep dive. We're constantly fed this narrative that the only path forward in AI is building infinitely larger, trillion-parameter models that require the energy output of a small nation to train. Right. The scale-is-all-you-need mindset. that. Yeah. But this paper proves that we don't necessarily need brute force scale. We just need to use our mathematical tools smarter.
19:42We can apply reinforcement learning much earlier in the pipeline. We can run it in parallel with SFT to prevent the model from forgetting its general knowledge. And we can optimize our FLOPs by tuning the rollout counts to the model's maturity. Exactly. It's a structural shift. Reinforcement learning is no longer just the final coat of paint you put on a finished product. It's a fundamental capability expanding mechanism that should be integrated from the very beginning. So as you integrate these AI tools into your daily workflow, keeping this training architecture in mind is crucial. It tells you exactly where a model's blind spots are.
20:19That's the practical takeaway. If we connect this to the bigger picture, if you know you are using an open source model or a commercial tool that relies heavily on standard sequential SFT, you should expect it to be a rigid test taker. Right. It will excel if your prompt exactly matches the format of its training data, but it will likely suffer from catastrophic forgetting on lateral logic or general reasoning tasks. Yeah, it'll just stumble. Conversely, models trained using early concurrent RL are fundamentally different under the hood. They will be much more adaptable and capable of exploring novel solutions when you throw them a curveball.
20:53Which is exactly what we need as we push these systems into more complex unmapped territories, like coding and scientific research. But looking at the sheer power of early RL leaves me with a final thought, something this paper kind of hints at, but leaves open for the future. Oh, what's that? Well, if reinforcement learning allows an AI to successfully teach itself complex logic through messy, self-generated traces at the very beginning of its life, do we even need human ground truth demonstrations for the next generation of models? That is the million-dollar question. Right. If forcing an AI to mimic human syntax actively constrains its internal reasoning and causes catastrophic forgetting, are we approaching the point where our rigid human way of thinking is actually the bottleneck holding true artificial intelligence back?
21:41Something to mull over as you explore the latent space of your own day. Thanks for joining us on this deep dive.
From the publisher
This research investigates the effectiveness of integrating reinforcement learning (RL) earlier in the large language model training pipeline rather than treating it solely as a final post-training step. The authors demonstrate that RL is effective remarkably early, often matching the performance of standard sequential pipelines after only a small fraction of pre-training is complete. Unlike supervised fine-tuning (SFT), which tends to degrade a model's general capabilities and narrow its output, direct RL preserves general skills and expands the diversity of reasoning paths. The study also identifies that targeted data composition is more critical for RL success than simply increasing model size. Finally, the researchers propose a parallel averaging method that combines RL and SFT updates to achieve superior results across all training stages. Together, these findings suggest that the current standard of isolating RL to the end of training is an unnecessary design choice that limits model potential.




