In short
When to teach LLM reasoning—whether to “front-load” reasoning data during pretraining or bolt it on later via supervised fine-tuning (SFT).
Key claims
Reasoning must be learned early; SFT can’t fully catch up if pretraining lacked reasoning. Early gains compound through later stages (including reinforcement learning). Data allocation should be asymmetric: pretraining needs diverse reasoning at scale, while SFT needs high-quality, deep chain-of-thought (CoT) examples.
Notable examples
A control model with zero reasoning in pretraining gained only after doubling SFT data, but still lagged models that had reasoning early; reasoning-in-pretraining improved ~19% on hard reasoning tests, ~9.3% after SFT, ~18.57% final accuracy gap after reinforcement learning, and ~39.32% on tough math competition tasks.
Guests
None named in the transcript.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Importance of Timing in AI Training
0:45 to 2:28
Discussing the timing of incorporating reasoning in AI model training.
“You start general, train on tons of text to get language down.”
Benefits of Front-Loading Reasoning Data
2:28 to 4:19
Exploring the advantages of integrating reasoning data early during training.
“A huge head start before you even start fine-tuning.”
Strategic Data Allocation in Training Phases
4:19 to 5:48
Analyzing the difference between data strategies in pre-training and SFT.
“That initial choice, what goes into pre-training, it basically sets the performance ceiling for the final model.”
Quality vs. Quantity in SFT
5:48 to 8:08
Examining the significance of high-quality data in the SFT phase.
“The foundation needs exposure to many different kinds of reasoning.”
The Synergy of Pre-Training and SFT
8:08 to 10:17
Understanding how high-quality data in both phases reinforces model skills.
“So maybe not worth it then for pre-training.”
Implications for Future AI Development
10:17 to 12:28
Discussing the impact of early reasoning data on broader AI capabilities.
“Using the high-quality data in both phases reinforces the skills.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we're really cracking open the playbook for building these incredibly powerful AI models. That's right. Our mission today is to figure out maybe the most crucial decision in training LLMs, timing. When does reasoning get baked in? Yeah, it's the question right now, isn't it? If you're building these systems, you have to know, is reasoning, you know, a specialized skill you teach later? Like an add-on. Exactly. Or is it something fundamental, something that needs to be there right from the start during that massive pre-training phase? And the stakes are, well, huge.
0:35Get the timing wrong. And you could spend millions training a model that just hits a wall. A performance ceiling you can't break. An unfixable ceiling, yeah. So historically, there's been a kind of standard approach, right? You start general, train on tons of text to get language down. Trillions of tokens, yeah. Basic knowledge, grammar, syntax. And then reasoning was sort of bolted on later, often using what's called supervised fine-tuning, SFT. Right. SFT uses that really high quality, often human checked data, instruction following stuff. Expensive stuff. Incredibly expensive and time consuming to create, especially those complex step by step solutions, the chain of thought examples.
1:15So, yeah, the thinking was save that precious data for the end. Refine the model. Treat reasoning like a finishing touch. Seemed logical. Economical anyway. It did. It was the convention. But the research we're digging into today really challenges that. It asks, okay, what if we put that sophisticated reasoning data in much earlier during pre-training? Does it mess things up? Make the model too narrow too early? Or is it the only way to build a solid foundation for real intelligence? And the answer seems pretty clear now based on this work. It's a direct challenge to how things have been done.
1:51Which is? Front loading that reasoning data, it seems absolutely critical. Pre-training with it builds these foundational logic skills that you just can't seem to add effectively later with SFT. So you're saying SFT can't make up for lost time. That's what it looks like. You're essentially building the engine for reasoning early on, and it has to be installed from the start. Okay, let's unpack that. This idea of a durable advantage. Because the common wisdom might have been, oh, the model can just catch up later. Right, catch up hypothesis. And this research says, nope, false. Early investment pays off, and it keeps paying off.
2:25It compounds. That's the core finding. They saw that putting diverse reasoning data into that initial pre-training led to this immediate, really significant jump, plus 19 % on average on tough reasoning tests. 19%. That's huge. It's massive. A huge head start before you even start fine-tuning. And they really tested that catch-up idea directly, didn't they? They took a model with a weak start. Yep. The baseline model, let's call it the control, zero reasoning data in its pre-training. Just general text. And they tried to force it to catch up how? They threw the kitchen sink at it in the next phase.
2:59They gave it double the amount of that high-quality SFT data compared to the models that did get reasoning early on. Trying to compensate. Exactly. And the results, well, they were pretty stark. Really undeniable. It didn't work. It didn't even come close. Doubling the SFT data for that control model still wasn't enough to match the performance of even the worst performing model that had some reasoning data during pre-training. Wow. So SFT is limited by what happened before. Fundamentally constrained. You can't fix a weak core reasoning engine later. It doesn't matter how much fancy targeted instruction you give it.
3:35I like your analogy. It's like building a race car. Yeah. If you get the chassis and engine block wrong at the start, putting on expensive racing tires later isn't going to suddenly give it horsepower it never had. That's a perfect analogy. And you see that compounding effect you mentioned. The advantage doesn't just appear and vanish. It gets amplified. How so? Okay. So take the group of models that did get reasoning in pre-training. After they also went through SFT, they still outperformed that control group, the ones that started weak by 9.3 % on average. So SFT helps everyone, but it helps the ones with a better start more.
4:09Exactly. SFT acts like a turbocharger, boosting performance, but it only works effectively if the engine, that core logical processing, was built correctly during pre-training. And the final proof seems to come after the last stage, reinforcement learning. Right. The ultimate alignment phase. That initial choice, what goes into pre-training, it basically sets the performance ceiling for the final model. And the gap was still there. Oh, yeah. A math and 18.57 % final accuracy gap between the model fully pre-trained with reasoning and that weaker control model. Almost 20%. And it gets even wider on really specialized stuff.
4:47They tested on these super hard, AIM-y competition math problems. Okay, those are tough even for humans. Extremely tough. And the models with the reasoning front-loaded, they showed a, get this, 39.32 % improvement over the baseline. Wow, nearly 40%. That's not just better, that's a different class of performance. It's the difference between actually solving the problem and just failing. It really confirms it. Those earliest steps, what you feed the model when building its foundation, that literally dictates the upper limit of what it can ever achieve. Okay, so principle one established. Front load the reasoning data.
5:18It's crucial. But what data? Just dump it all in. Ah, good question. Because you can't just throw everything in and hope for the best. That leads to the second key idea, this asymmetric principle of data allocation. Asymmetric, meaning it's not the same strategy for every stage. Exactly. It's not simply more data is always better. It's about being strategic. What works best depends entirely on whether you're in pre-training or SFT. The strategy needs to be different, asymmetric. Okay, let's break that down. Phase one, pre-training. What's the priority there? Diversity and scale. Think broad. The foundation needs exposure to many different kinds of reasoning.
5:57Like building up a wide base of understanding. Precisely. Like building a reservoir, you need a wide catchment area. The research showed the best results came from a large, diverse reasoning data set. It had about 56 % of the number of tasks. A real mix. A real mix. And that diverse set gave the biggest initial boost, that plus 11 % advantage we mentioned earlier. Pre-training needs breadth. It needs to seal lots of problem types to learn the underlying abstract structures of logic. Okay, so pre-training diversity, what about the next phase, SFT, the refinement phase? The priority completely flips.
6:35Once that broad foundation is there, SFT is all about quality. Specifically, quality defined by depth and complexity. Quality over quantity now. Definitely. If pre-training is the reservoir, SFT is the high-precision filter and pump system. They saw a huge plus 15 % gain just from using specialized high-quality SFT data. And quality here means what exactly? Long answers. It largely means long, detailed chain of thought or CO-T examples. Right, CO-T, where the model doesn't just give the answer. It shows its work step by step. Exactly. And that depth seems critical for SFT, the high-quality data they used.
7:10The average answer length was over 10 ,000 tokens. 10 ,000. Compared to what for the mixed stuff? Compared to only about 550 tokens for the average mixed quality example. So it's the difference between a quick answer and like a detailed textbook explanation walking you through everything. That's a huge difference in complexity. Massive. And they even proved this point another way. They took a big noisy data set and just filtered it. They kept only the examples where the answer length was over, say, 4096 tokens. Just filtering by length to get deeper examples. Right. Creating a quality filtered set.
7:44And that set, even though it was much smaller, gave huge improvements in SFT. It really hammers home. Reasoning depth is the key quality marker for that targeted SFT phase. Okay, but there's a nuance here, right, about synergy. You mentioned high-quality data. It wasn't immediately amazing in pre-training. Yeah, this was really interesting. Adding that super high-quality, long-coatee data early during pre-training, it gave only a very small initial benefit, barely noticeable sometimes. Huh. So maybe not worth it then for pre-training. We'll wait for it. Here's the twist. After that same model went through SFT, that early high quality pre-training suddenly unlocked an extra performance boost.
8:24An additional plus 4 % gain compared to models that only got the diverse mixed quality stuff early on. So it was like latent potential, lying dormant. It really seems like it. It suggests pre-training with high quality stuff builds this kind of hidden capacity. The model learns the patterns of deep reasoning, but it needs the focused SFT phase to really activate that potential and bring it out. Fascinating. So the foundation matters, but SFT activates it. Which brings us to a potential pitfall. If quality is so crucial for SFT, what happens if you just try to scale up using, well, junk data or mixed quality data?
9:03Yeah, the research is crystal clear here. Just blindly scaling up your SFT using mixed quality data. It's not just ineffective. It can actually hurt performance. Really? Actively harmful? Yes. This is the big more is not better warning sign. They tried it. They doubled the amount of mixed quality SFT data. The resulting model showed basically zero average improvement overall. And worse, it actually damaged its mathematical reasoning skills performance dropped by 5 % on average in math tasks. So you're diluting the signal. Exactly. That shallow, noisy reasoning, when you scale it up, it just washes out the precise, useful signal the model needs for refinement.
9:39SFT is about precision, following complex instructions. Feed it diluted instructions, you get worse results. Okay, that makes sense. But it raises another question. This high-quality data, the really good COT stuff, it's expensive and hard to get. Very. Should you use it twice? Once in pre-training to build that latent potential, and then again in SFT to activate it? Or does the model just forget the first time it saw it? Catastrophic forgetting, isn't that a concern? It's often a valid concern, yes. But here in this context, the research suggests that kind of strategic reuse is actually good.
10:12It acts like reinforcement, consolidating the skill, not making it forget. So seeing it again helps. It seems so. Using the high-quality data in both phases reinforces the skills. Think of it like this. Pre-training lets the model internalize the abstract patterns of complex reasoning. Then SFT comes along and provides really strong, focused practice on that prepared foundation. Like learning a concept in class and then doing intense homework problems on it. Exactly that. It locks it in. Okay. So bringing it all together, if you're out there actually building these models, there seems to be a pretty clear blueprint emerging.
10:47Absolutely. The core takeaway, the actionable plan, is pretty straightforward now. For that foundational pre-training phase, go broad. Prioritize scale and diversity of reasoning types. You're building the general logic engine. Get the foundations wide. Right. But save your most precious resources, the specialized, super long, high quality chain of thought examples for the targeted refinement phase of SFT. Don't waste them early. Use them for focus later. Precisely. That asymmetric strategy, hitting diversity early and quality later, seems to be the key to maximizing your results, getting the most bang for your buck computationally and data wise.
11:25And as we wrap up this dive, there's one last implication here that's really interesting. You mentioned the models got better not just at math and code. Yeah, that was intriguing. Injecting reasoning data early didn't just boost the scores in the areas you'd expect. It seemed to help the model develop better internal ways of representing abstract logical structures generally. Leading to? Unexpected gains. Even in science domains that weren't the main focus of the reasoning data, it suggests the model learned something more fundamental about how to reason, which transferred across domains. Okay, so if that early diversity is the key to building these more general, domain-agnostic reasoning abilities, that makes you wonder, doesn't it?
12:04It really does. What's the optimal breadth then? To build a truly robust generalist logic engine, do we need reasoning data covering, I don't know, law, philosophy, history, art, criticism? Or is focusing on the big ones, math, code, science enough to cover the essential structural patterns of logic? How broad do you need to go to build that foundation for a true generalist AI? That's the next frontier, isn't it? Something for you, the listener, to think about. We know timing is critical and we know the type of data matters for each phase. But just how wide that initial exposure needs to be, that might define the next big leap in AI capability.
12:43a fascinating question to end on we'll be back soon for another deep dive looking forward to it
From the publisher
This research paper, by authors affiliated with NVIDIA, Carnegie Mellon University, Boston University, and Stanford University, focuses on the optimal strategy for incorporating reasoning data into Large Language Model (LLM) training. The central finding challenges the conventional approach of relying solely on post-training, demonstrating that "front-loading" reasoning data during the pretraining phase is critical, yielding a durable 19% average performance gain on expert-level tasks. The research establishes an asymmetric principle for data allocation: pretraining benefits most from broad diversity and scale in reasoning patterns, while supervised fine-tuning (SFT) is most sensitive to high data quality. The study concludes that early investment in reasoning creates a foundational capacity that cannot be fully replicated by later-stage fine-tuning, advising against naively scaling mixed-quality SFT data.




