In short
How pre-training, mid-training (continued pre-training), and RL post-training interact to improve reasoning in language models, and when RL truly adds new reasoning vs merely polishing existing skills.
Guests
None mentioned in the transcript.
Guest/author backgrounds
None provided.
Key claims
RL improves reasoning only when tasks target the model’s “edge of competence” (not too easy, not too hard). Cross-context generalization requires a small pre-training “seed” of new-domain primitives (~1% exposure). Adding mid-training plus RL beats RL alone by 10%+ under fixed compute. Use process-aware rewards to prevent reward hacking.
Notable examples
Synthetic DAG-based arithmetic reasoning tasks with OPG complexity; “OEDGE” (11–14) yields up to 42% pass@128 gains; “OPT-20” cross-context transfer fails with 0–0.1% seed but succeeds with ~1% seed; process-aware rewards improve pass@1 by 4–5% on 15–20 tasks.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Reinforcement Learning in Language Models
0:45 to 3:42
Exploration of how reinforcement learning impacts language model reasoning.
“invent that solution or did it just find a similar pattern in some obscure corner of its training set?”
Success Metrics in AI Training
3:42 to 7:40
The critical axes for measuring model intelligence and performance.
“People have sort of assumed it does for years, but proving it was.”
The Role of Mid-Training in AI
7:40 to 10:42
Importance of mid-training as a bridge between pre-training and RL.
“They found that introducing even a really sparse exposure, we're talking 1 % or more of the new context's most basic operations.”
Preventing Reward Hacking in AI Models
10:42 to 12:30
Strategies using process-aware rewards to enhance model reliability.
“They probably have a really well-designed mid-training stage.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive, where we cut through the noise and deliver the essential insights from complex research. straight to you. Today, we're diving deep into the engine room of AI. We're wrestling with a question that honestly has been bugging researchers for a while now. When we use reinforcement learning RL to fine tune a language model, are we really teaching it new ways to reason? Or are we just, you know, making it a bit better at using skills it already learned during pre-training? It's this huge conceptual conflict, isn't it? Is RL a creator? Or is it just a really magnificent tuner.
0:33And for the longest time, we just couldn't know. Modern LMs are, for all intents and purposes, black boxes. Totally opaque. Exactly. They're trained on these vast, messy piles of internet data. So when a model solves a new complex problem, you have to ask, did it invent that solution or did it just find a similar pattern in some obscure corner of its training set? And that's why this new research from Carnegie Mellon University's Language Technologies Institute is such a big deal. They've essentially built a fully controlled lab for training these models. Every token, every reasoning step, it's all tracked.
1:08Which means they can finally isolate and measure the causal contribution of every single stage of the training process. So our mission for this deep dive is to lay out the strategic map that comes from this study. We're going to untangle the roles of the three main phases. Pre-training, this middle phase called mid-training, and of course the RL-based post-training. We'll see how they have to be sequenced, how they have to be balanced, if you want a model that can truly generalize its reasoning abilities. Okay, let's unpack this. The researchers measured success along two really critical axes.
1:39Think of them as the depth and the breadth of the model's intelligence. The first one is extrapolative generalization. That's the depth. It's asking, can the model solve problems that are structurally more complex than anything it saw during its training? And they measure that complexity with a really precise metric called OPG. It's basically the number of arithmetic operations in the problem's dependency graph. So if you train it on 10-step problems, can it handle a 15-step or even a 20-step one? That's the real test of depth. Right. And then the second axis is contextual generalization. This is breadth.
2:15The transferability. Exactly. Can the model learn a logical structure from, say, problems about animals in a zoo and then apply that exact same logic to a totally different narrative, like one about teachers in a school. The context is completely different, but the math is identical. Can the skill make that leap? And the whole framework is built on these synthetic reasoning tasks that use dependency graphs or DAGs. The beauty of this is that they have precise granular control over every variable. They can dial up the complexity LPG to any level they want. And the big benefit for us, you know, the people who actually want to use or build these things, is that this gives them what they call observable, parsable reasoning processes.
2:53Which means they aren't just looking at the final answer. They can check every single intermediate step the model takes. Which is so critical because it basically eliminates the risk of reward hacking. Right, where the model stumbles into the right answer for all the wrong reasons. Here, they force it to show its work. So just for context, what kind of model are we talking about here? It's a smaller decoder-only model. Think something in the style of a Quen 2.5, about 100 million parameters. And it was pre-trained on 10 billion tokens of what they call in-distribution problems. So basic reasoning tasks from OP equals 2 up to 10.
3:32So that's the foundation. That's the basic toolkit the model starts with. And RL then has to either polish it or build on top of it. Precisely. OK, so this brings us to the core finding. Does RL actually extend a model's reasoning ability? People have sort of assumed it does for years, but proving it was. Tricky. Impossible, really. But this study gives a definitive yes. but, and this is the key part, this is where the strategy comes in, it only works under very, very specific conditions. The number one condition is task calibration. RL only generates real capability gains when the task difficulty is aimed squarely at what the paper calls the model's edge of competence.
4:11So it has to be hard enough to force new learning, but not so hard that the model is just, you know, guessing. Exactly. And to understand those gains, we have to be clear on the metrics they used. There are two. Pass at one. Right. That's the first one. Right. Pass at one is about reliability. Did the model get it right on the very first attempt? That's what you care about in a real-world application. But then there's the other one, pass at 128. And that's a measure of potential. They let the model try the problem up to 128 times to see if the ability to solve it even exists inside the model's parameters.
4:43If your pass at 128 score goes up, it means you've learned a genuinely new capability. You haven't just gotten better at expressing an old one. Got it. Okay, so with those metrics, let's look at the different difficulty zones they tested. First up, the easy stuff, the in-distribution tasks, off 2 to 10. The stuff the model already saw in pre-training. And if you use RL here, all it does is sharpen existing capabilities. Your pass at 1 might go up a little. So it gets a bit more reliable. A little. But crucially, there's zero improvement in pass at 128. The skill was already maxed out. It's like asking a pro chef to practice making scrambled eggs.
5:21They might get a tiny bit faster, but they're not becoming a better chef. Then you jump to the other extreme, the ODARD tasks. And 15 to 20. Right. And if the task is too far beyond what the model knows, it just doesn't have the basic building blocks to even start. RL is completely ineffective there. Like asking that chef to build a skyscraper. Perfect analogy. They just don't have the foundational skills. But then there's the breakthrough, the middle ground, the Oded Edge tasks, up 11 to 14. This is the sweet spot. When they applaud RL right here, just beyond the model's pre-training comfort zone, they saw genuine capability gains.
5:56How much are we talking? Up to a 42 % increase in pass at 128. It's a huge jump. And it proved the model wasn't just replicating something it had seen. It was forced to take its existing opt-end skills and compose them into deeper, structurally novel solutions. So to stick with the analogy, the model knows how to make a five-layer cake. At the OEDGE, you're asking it to make a seven-layer cake. It has all the primitive skills, baking, frosting, but it has to invent the structural engineering to stack those extra layers without it all falling over. And that new structural composition is the new reasoning skill that RL taught it.
6:32So this gives us our first big piece of practical guidance. To get the most out of RL, you have to design your data to hit that sweet spot. Exactly. Target problems where the model fails on the first try, so pass at 1 is low, but succeeds after a few tries, so pass at K is high. That's the very definition of the edge of competence. Okay, so we've established RL can make a model reason deeper. Let's shift to breadth. Contextual generalization. Can it take that math skill it learned from the zoo animal problems and apply it to the school teacher problems? And this is where the importance of that initial pre-training phase really comes roaring back.
7:08RL, it turns out, is a brilliant composer. But it cannot create something from nothing. It needs a seed. It needs a minimal seed of primitives from that new domain to work with. The numbers they found here were honestly just stunning to me. I'm very credible. They tested what happens if pre-training provides next to no exposure, like 0%, or even just 0.1 % to the new context. And RL completely failed. The model could not make the jump no matter how much fine-tuning they threw at it. But then they plant this tiny, tiny seed. And this is the amazing part. They found that introducing even a really sparse exposure, we're talking 1 % or more of the new context's most basic operations.
7:48Just the simplest op equals two examples. That's all it took. That was enough. Once that minimal seed was planted, RL could robustly amplify it and achieve strong cross-context generalization on even the hardest OPT-20 tasks. Just 1%. That feels so counterintuitive. Aren't you worried that by trying to sprinkle in 1 % of every possible domain, you'd just dilute the whole pre-training process? It's a fair question, but the data here suggests the tradeoff is absolutely worth it. The goal isn't to teach complex things in every domain. It's just to provide that basic coverage. You plant the seeds of these basic atomic building blocks.
8:26And then RL, which is incredibly efficient at this, comes in and does the heavy lifting of composing them into complex structures later. So for the simple stuff, the model is mostly just replicating patterns. But as things get more complex, it starts generating truly novel structures, but only if that little contextual seed was planted right at the very beginning. And that leads us to our second piece of practical guidance. When you're designing a pre-training curriculum, prioritize broad coverage of basic primitives over depth. Plant as many seeds as you can, even at a low density, like 1%. Right, because RL is the efficient compositor that we'll build on them later.
9:02Okay, we've covered the start of the process, pre-training, and the end, RL. Let's talk about that underexplored middle step, mid-training, sometimes called continued pre-training. Right, CPT. It's supposed to be this bridge between the super broad knowledge of initial pre-training and the very sharp specialization of RL. And the core finding here, especially when you're working with a fixed compute budget, which let's be honest, everyone is. Everyone is. Is that adding a dedicated mid-training phase makes a huge difference. It just substantially strengthens generalization. A pipeline with mid-training plus RL consistently beat RL alone.
9:40We're talking an average improvement of over 10 % on those really hard OD tasks. So you can't just skip it. But if your compute is limited, how do you decide how to split it? And that brings us to guidance number three. The allocation has to be task-aware. Precisely. If your main goal is reliability on those OD-age tasks, you're optimizing for pass at one, then you should put most of your budget into mid-training. Then just use a bit of light RL at the end. You're building a stable, reliable model. But if your goal is pushing the absolute limits of generalization, maximizing that pass at 128 score on Odie Hard tasks, you do the opposite.
10:19You'd use a modest budget for mid-training, just enough to get the model ready, and then you pour the rest of your compute into heavy RL exploration. Yeah, you can think of mid-training as putting the training wheels on the bike. It stabilizes the model, it conditions its priors, and it makes it ready for RL so the learning is efficient and it actually sticks. It kind of explains why some of the big proprietary models seem to respond so much better to RL fine-tuning than others. They probably have a really well-designed mid-training stage. Almost certainly. Okay, last major point. Let's loop all the way back to that fear we talked about at the beginning.
10:55Reward hacking. The model getting the right answer by cheating. How did they stop that? They solved it with something they call process-aware rewards. They basically took their very strict evaluation rule. You need correct intermediate steps and the correct final answer. and they baked it directly into the reward function itself. Okay, so walk me through that. Imagine you're teaching a kid math. Standard RL is like only giving them a gold star if the final answer is right. Right. That's a sparse reward, what they call R out. The kid could have just guessed or scribbled nonsense and gotten lucky.
11:25But process-aware rewards are like looking over their shoulder and checking their work on the paper. Exactly. That's the dense reward, RRPV. You're giving feedback on every single step, even if the final answer is right. If they skipped a step or made a logical error along the way, you penalized them because the process was wrong. So the final reward is a blend of both the outcome and the process. And did it work? It did. Integrating that dense process feedback consistently led to measurable gains. We're talking a 4 % to 5 % improvement on pass at 1 for those really tough op 15 to 20 tasks. And I imagine it just makes the model more trustworthy.
12:02It fundamentally shifts the model away from finding shortcuts and towards faithful, verifiable reasoning. It teaches it how to think, not just what to spit out. Which gives us our fourth and final piece of guidance. If you have a way to verify the intermediate steps, if that process supervision is high quality, you should absolutely be blending that sparse final answer reward with dense process-level feedback. It's your best defense against reward hacking, and it builds a more consistent, reliable model. So if we pull this all together, we've got a clear strategic recipe. We do. First, RL needs to target the edge of competence.
12:37Not too easy, not too hard. Second, pre-training has to provide that minimal 1 % seed of basic knowledge for any new context you want the model to generalize to. And third, that mid-training stage is an essential bridge. It provides the scaffolding that makes RL stable and effective, especially when your compute is limited. What this study really does is it takes the kind of vague art of LM fine-tuning and turns it into a clear, measurable engineering discipline. It really does. It shows that true innovation in AI reasoning isn't about brute force. It's not just bigger models and more data. It's about strategically sequencing and balancing complexity across the entire training pipeline.
13:17It makes everything fit together perfectly. Exactly. It really makes you think, though, if just the tiniest exposure, that 1 % of basic operations during pre-training, can unlock this massive deep generalization later on through RL, what latent skills are just sitting there, dormant, inside the huge general purpose models we use every single day, just waiting for the right perfectly calibrated RL data set to come along and wake them up? That is definitely something to think about. Thank you for joining The Deep Dive.
From the publisher
This paper details a controlled experimental framework used to examine the interaction between pre-training, mid-training, and reinforcement learning (RL) on the reasoning abilities of language models (LMs). Researchers from Carnegie Mellon University and the Language Technologies Institute utilized a synthetic dataset with explicitly defined reasoning complexity and contextual templates to isolate the causal effect of each training stage. Key findings indicate that RL yields true capability gains only when targeting the model's "edge of competence," where tasks are difficult but still within reach of generalization. Furthermore, minimal pre-training exposure to long-tail contexts is critical for RL to induce robust contextual generalization, and incorporating a mid-training phase substantially improves performance under a fixed computational budget. Finally, the study confirms that process-aware rewards effectively mitigate reward hacking and enhance reasoning fidelity.




