In short
TailSFT (Filtered Fine-Tuning) fixes the “SFT trap” where standard supervised fine-tuning collapses probability mass onto easy majority answers, hurting later reinforcement learning (RL) that needs diverse candidate reasoning paths.
Guest backgrounds
No guest identities are provided; the episode is a two-host “Deep Dive” discussion with no named external guests.
Key claims
Minimizing cross-entropy during SFT boosts pass-at-1 but destroys coverage (pass-at-K). TailSFT runs a baseline “dry run,” then masks easy sequences by zeroing their loss once improved, forcing learning on the hard tail. This improves RL with GRPO because GRPO requires diverse sampled responses.
Notable examples
Graph-navigation experiment (pass-at-1 drops, pass-at-8 stays strong). Benchmarks using OMO3-7B: Cruxivalo coverage +16.8% (pass-at-16), AM +3.1%, RL gains up to +3.9% pass-at-1; RL reward rises 2.5x faster. Coverage ratio diagnostic: ROE16 > 1 predicted benefit in 10/11 settings.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the TAILSFT Paradox
0:55 to 2:16
Exploring the flaws in standard AI training and how it limits complex reasoning.
“Looking at the cutting edge of AI development today, we are unpacking a fascinating paradox that looks a lot like that basketball player.”
Phases of AI Model Training
2:16 to 3:18
Breaking down the three phases of training and focusing on supervised fine-tuning.
“I mean, to force the AI to get smarter, developers are having to artificially stop it from studying the easy stuff.”
The Impact of Cross-Entropy Loss
3:18 to 5:32
Discussing how the minimization of cross-entropy loss affects AI performance.
“What is actually happening mathematically during standard supervised fine-tuning that causes this massive problem?”
The SFT Trap Explained
5:32 to 7:45
Examining how the SFT phase compromises the model's capability for complex tasks.
“Which brings us to a really crucial distinction in how we evaluate these models.”
Introduction to TAILSFT
7:45 to 8:30
Introducing TAILSFT as a solution to the SFT trap without major overhauls.
“Okay, so if standard SFT is basically poisoning the well for the reinforcement learning phase, developers must be scrambling for a way to intervene.”
Mechanics of TAILSFT
8:30 to 11:39
Explaining how TAILSFT operates to improve model training efficiency.
“Like, how does it decide what to filter?”
Real-World Applications of TAILSFT
11:39 to 14:01
Discussing the performance of TAILSFT in various benchmarks and its implications.
“And the resulting metrics are what developers really need to wrap their heads around.”
Understanding TAIL-SFT's Impact on RL
14:01 to 16:43
Learn how TAIL-SFT improves model performance in reinforcement learning.
“And again, just to reiterate the paradox here, TAIL-SFT achieved this massive coverage expansion while sometimes slightly lowering the intermediate pass-at-1 accuracy, right?”
The Coverage Ratio Diagnostic Tool
16:44 to 19:17
Discover how the coverage ratio helps determine TAIL-SFT's effectiveness.
“So the model that started the RL phase looking slightly worse on greedy accuracy ended the RL phase significantly smarter.”
The Philosophical Implications of TAIL-SFT
19:18 to 22:09
Explore the deeper questions surrounding memory and learning in AI.
“And if we connect this to the bigger picture, this diagnostic entirely validates a stage-aware approach to AI model development.”
Transcript
Automatic transcript. May contain errors.0:00Imagine a star athlete, let's say a basketball player. They're in the gym practicing free throws and they shoot a hundred times and they hit 99 of them. Pretty good stats. Right. So the coach looks at those stats and says, incredible. Your accuracy is perfect. Just keep practicing free throws. Oh, I see where this is going. Yeah. So the athlete spends the next month doing absolutely nothing but shooting free throws. Yeah. I mean, they never run complex plays. They never practice defense. Right. Just standing at the line. Exactly. Then game day arrives. The opponent double teams them, forces them to the perimeter, and suddenly our star athlete is completely paralyzed.
0:41Because they only know one highly predictable scenario. Exactly. They've optimized so heavily for that one specific thing that they crippled their ultimate potential to actually play the chaotic, you know, dynamic game of basketball. It's a great analogy. Welcome to the Deep Dive. Looking at the cutting edge of AI development today, we are unpacking a fascinating paradox that looks a lot like that basketball player. Yeah, it really does. We're looking at this phenomenon where trying to make an AI model, quote unquote, better in the middle of its training actually destroys its ability to do complex reasoning later on.
1:15It is a phenomenal paradox, and it challenges a lot of our basic assumptions about machine learning. Right, because we usually think of improvement as a straight line. Exactly. We are so used to thinking that if an AI model gets better scores on an intermediate test today, well, it's just going to be a smarter, more capable model tomorrow. But that's not the case here. Not at all. Today is really about stepping back and seeing the forest for the trees. To understand what's happening at the bleeding edge for you, the listener, we have to look at the full training trajectory of a model. Instead of just obsessing over those isolated short-term metrics.
1:51Exactly. Okay, let's unpack this. We are doing a deep dive into a breakthrough training method called TAILSFT. Yes, TAILSFT. We're going to explore exactly why the current middle step of standard AI training is fundamentally flawed, and how simply forcing an AI to stop practicing what it already knows unlocks these massive gains in complex math and coding. Which is such a wild concept. It really is. I mean, to force the AI to get smarter, developers are having to artificially stop it from studying the easy stuff. Right. And to appreciate why TAIL-SFT is such a necessary intervention, we need to zoom in on the standard assembly line of how a modern language model is built.
2:32Okay, let's lay that foundation. So most people following AI know the basic three phases. First, you have pre-training on massive data sets. Right. That's just reading the whole internet, basically. Pretty much. Then phase two is supervised fine-tuning, or SFT, which teaches instruction following and formatting. Teaching it to actually act like a chatbot. Exactly. And finally, you have phase three, reinforcement learning, or RL. And that's what drives a really complex reasoning. Okay. Pre-training, SFT, then RL. Right. But the failure point we are examining today sits squarely in the transition between stage two and stage three.
3:07The handoff from SFT to reinforcement learning. Yes. So assuming we all understand that pre-training builds the raw knowledge base, let's focus on that middle phase, the SFT phase. Okay. What is actually happening mathematically during standard supervised fine-tuning that causes this massive problem? Well, during standard SFT, the algorithm's primary goal is to minimize what we call cross-entropy loss. Cross-entropy loss. Right. You are feeding the model an instruction and a target response. And the model is constantly being rewarded for getting more and more confident at producing that exact sequence of tokens.
3:45Just nailing that one specific answer. Exactly. The mathematical objective is to lower the cross-entropy loss, which directly translates to pushing the model's confidence in that specific answer as close to 100 % as possible. Wait, I have to push back here for a second. Go ahead. I mean, if I am a developer spending millions of dollars on compute, isn't getting a lower error rate and higher confidence exactly what I want? It seems totally intuitive. Right. Like, why is a low cross entropy loss a bad thing? I want my model to be highly confident in the correct answer. I get that. But you have to look at the mechanics of how a neural network actually outputs language.
4:23Okay. When an AI predicts the next token, it doesn't just pick one word and move on. It has options. Right. It generates a probability distribution across its entire vocabulary. It assigns a probability score to every single possible next word. And that's what you call the probability mass. Exactly. The probability mass. Now, when SFT minimizes cross entropy loss, it is forcefully squeezing almost all of that probability mass into one single expected token. Oh, I see. And at the same time, it's pushing the probabilities of all the alternative tokens down to near zero. Wow. So it's not just learning the right answer.
4:59it is actively unlearning or at least suppressing any alternative ways to answer. That is the mathematical reality of it, yeah. SFT shifts the probability mass away from weird, complex, or unusual reasoning paths and just dumps all of it into the most obvious standard majority path. It's the basketball player shooting free throws again. Exactly. It gets hyper-focused on producing responses it can already produce easily. Okay, so the model becomes this highly confident but very one-dimensional thinker. Which brings us to a really crucial distinction in how we evaluate these models. The difference between PAS at 1 and PAS at K.
5:38Ah, okay. The researchers refer to PAS at K as coverage, right? Right. Let's clearly define those two metrics for the listener because they are the hinge for this whole tail SFT breakthrough. Sure, so PAS at 1 is what we call greedy accuracy. Greedy accuracy. Yeah, if you ask the model a question and it generates one single response, does it get the answer right on that first try? That is your pass at one. And because SFT pushes all the probability mass into the most obvious answer, it's probably really good at that. Exceptionally good. It drives up pass at one for standard tasks very quickly.
6:11But pass at K is a totally different way of measuring success. Right. PASECK asks a broader question. It says, if you let the model generate K number of different responses, say 16 different attempts at the same prompt. Okay, 16 different tries. Yeah. Is at least one of those 16 attempts the correct, brilliant answer? And this represents the model's coverage. Exactly. It measures the breadth of the menu of options the model retains in its repertoire. are. And if we look ahead to the final stage of training reinforcement learning, that's where the model is supposed to learn advanced reasoning, right?
6:46Complex coding, difficult math. Right. RL is the final boss. An RL literally requires a broad menu of options to sample from because in RL, you don't give the model the answer. No, you just give it a prompt and let it generate attempts. And then you reward the successful one. Exactly. But if standard SFT has completely destroyed the model's coverage by pushing all the probability mass onto a single basic response. The menu is empty. Right, the RL phase is mathematically starved. If the RL algorithm can't randomly sample a complex nuanced answer from the model's repertoire, it has absolutely nothing to reward.
7:23So the training just stalls out. Essentially, yes. The model won't even attempt the harder reasoning paths because the SFT phase beat that out of it. Man, so optimizing for that short-term metric low cross entropy and high pass at one literally sabotages the model's potential for the final most crucial stage of learning. And that right there is what the research calls the SFT trap. The SFT trap. Okay, so if standard SFT is basically poisoning the well for the reinforcement learning phase, developers must be scrambling for a way to intervene. They are. But they probably don't want to throw out the whole training architecture, right?
7:57No, completely rebuilding the pipeline is too expensive. Which brings us to TAILSFT. Right, the solution. TAILSFT is presented as a sequence-level filtering algorithm. And the elegance of it is in its architectural simplicity. It functions as a direct drop-in replacement for standard SFT. So you don't need a massive overhaul. Exactly. You don't need to completely rebuild your data pipeline or design a new neural network. So let's break down the mechanics for a second. How does TAILSFT actually force the model to stop practicing the data it already knows? Like, how does it decide what to filter?
8:32It relies on dynamic comparison. So before the SFT training phase even begins, TAIL-SFT takes the raw base model and runs it over the training data. Just a dry run. Yeah, to record the initial loss for each sequence. It establishes a baseline of how difficult a specific piece of data is for the model right out of the gate. Okay, so we have a starting benchmark for every single sequence. Right. Then, during the actual training process, TAIL-SFT constantly monitors the current loss of a sequence and compares it back to that initial baseline. I see. If the model has improved sufficiently on that specific sequence, meaning the loss has dropped past a certain threshold, TAIL-SFT dynamically masks it out of the training batch.
9:14Wait, when you say masks it out computationally, what happens? Are we literally deleting the data from the dataset? No, we aren't deleting it. During the backward pass of the neural network... Which is the step where the model actually updates its internal weights based on its errors, right? Exactly. During that step, tail SFT just zeros out the loss for that specific sequence. The model still sees the data, but it is physically prevented from updating its weights based on that easy example. So it basically mutes the update. Yes, it mutes the update. And by muting the update on the easy sequences, the algorithm is forced to redirect all of its computational effort toward the undermodeled tail of the data distribution.
9:53The difficult edge cases. The complex reasoning paths. That is brilliant. What's really fascinating here is how clearly this dynamic plays out when you strip away the complexity of language entirely. Oh, right. They did a purely mathematical test. Yeah, the research highlights a controlled experiment with graph navigation that perfectly illustrates this retention of probability mass. The directed graph experiment. Let's walk through that for a second. You basically have an AI agent that has to navigate from point A to point B across a series of connected nodes. Right. And in this controlled setup, they designed the graph.
10:26So there is an easy, highly visible majority path that works most of the time. The free throw. Exactly. And then there's a more obscure, complex minority path. So when they ran this under standard SFT, I'm guessing the model's behavior was exactly what we just talked about. Completely predictable. It found the easy majority path and aggressively anchored on it. All the probability mass shifted to that one single sequence. So its pass at one was incredibly high. Oh, yeah. It could solve the graph on the first try, using the basic method almost every time. But what about pass at eight? Like, if it got eight attempts?
11:00When researchers looked at its ability to find alternative, complex paths if given eight attempts, the metric just flatlined. The model had fundamentally unlearned the minority path. Because the SFT trap squeezed it out. Exactly. But, and here's the kicker, when they applied tail SFT to the exact same graph. The algorithm dynamically zeroed out the loss on the majority path once the model had learned it. Right. And because of that, the model was forced to retain probability mass across multiple different edges of the graph. It had to keep exploring. Yeah, it kept those alternative complex solutions alive within its neural network weights.
11:39And the resulting metrics are what developers really need to wrap their heads around. Because with TaylorSFT, the path at one actually dropped, right? It did. The greedy, single-sample accuracy went down a bit because the model was no longer hyper-confident in just the easiest path. But the pass at 8, the coverage remained robust. It did. It traded a fraction of short-term accuracy for a massive expansion of its creative problem-solving menu. And that is the exact state you want to model in before you hand it over to reinforcement learning. Exactly. Okay, so a controlled graph experiment is a great way to isolate the math.
12:14But, you know, if you're listening to this and you build or deploy models, you want to know if this actually works in the real world. Of course. Does this scale to massive language models with billions of parameters? The short answer is yes. It scales remarkably well. Let's talk about the proof. So to prove the real world viability, the researchers applied TAIL-SFT to the OMO3-7B base model. Which is a state-of-the-art 7 billion parameter open weights architecture. That's a serious model. Very serious. And they tested it in some of the most rigorous environments possible. Domain-specific math and coding tasks.
12:48Right. Let's dig into those benchmarks, starting with coding. They used the Cruxivalo benchmark. Yeah. And for context, Cruxivalo doesn't just ask a model to write a generic script. It gives the model a Python function and a specific input and asks it to predict what the exact output of that code will be. It requires, like, execution in the head. Right. It is a phenomenal test of actual reasoning. So what happened when they swapped out standard SFT for tail SFT on this benchmark? The coverage gains, specifically measuring pass at 16, were massive. On Cruxivallo, tail SFT drove up to a 16.8 % absolute gain in coverage over the standard model.
13:25Whoa, almost a 17 % jump. Yeah, it's huge. I guess because coding is so full of edge cases and alternative logical routes, retaining that probability mass just gives the model so many more chances to arrive at the correct execution state. Exactly. And the games were mirrored across other complex domains, too. On the AM benchmark, which consists of highly difficult competition-level mathematics. Super hard math. Yeah, TLSFT saw a 3.1 % absolute gain in coverage there. And on MBPP PlusWild, which tests synthesizing Python programs from descriptions, they saw consistent expansions of the model's repertoire.
14:04And again, just to reiterate the paradox here, TAIL-SFT achieved this massive coverage expansion while sometimes slightly lowering the intermediate pass-at-1 accuracy, right? Right. So if a developer is blindly looking at a standard leaderboard, the TAIL-SFT checkpoint might actually look worse in the short term. It might, which really requires a fundamental shift in how we evaluate intermediate training checkpoints. Okay, here's where it gets really interesting for me. The TAIL-SFT model has a broader menu of options. It has retained its probability mass. But does that actually translate to the final fully trained model being smarter?
14:37That is the million dollar question. Because reinforcement learning is the final boss. The whole point of TAIL-SFT is just to prep the model for RL. Right. If the final RL model isn't demonstrably superior, then these intermediate coverage metrics are meaningless. Exactly. But the data proves the ultimate payoff. They put both the standard SFT models and the TAIL-SFT models through a rigorous reinforcement learning protocol called GRPO. GRPO, which stands for Group Relative Policy Optimization. Let's spend just a second on the mechanics of GRPO because I feel like it explains why TAILS-FT is so incredibly effective here.
15:15Definitely. GRPO is a highly efficient RL method. Instead of relying on a massive, memory-heavy, separate value model to score every single response... Which is what older methods like PPO do, right? Exactly. Instead of doing that, GRPO samples a group of different responses to the same prompt directly from the model. Just pulls a batch of attempts. Yeah. It scores all the responses in that group, calculates the average score, and then updates the model's policy based on how each response performed relative to that group average. Ah, okay. So if a response in the group scores higher than the group average, its probability is increased.
15:49Oh, yeah. And if it scores lower, its probability is decreased. Exactly. But here is the catch. For GRPO to work effectively, the initial group of sampled responses needs to be diverse. It needs variety. Yes. It needs to contain a variety of reasoning paths so the algorithm can actually find a relatively better answer to reward in the first place. Which perfectly aligns with TAIL-SFT. Exactly. Because standard SFT collapsed all the probability mass, a GRPO sample group from a standard model is going to be incredibly homogenous. It's looking at 16 variations of the exact same basic, probably wrong answer.
16:24But a sample group from a TAIL-SFT model is going to contain a wide variety of distinct reasoning paths. And because of that diversity, initializing the GRPO phase with the TAIL-SFT checkpoints consistently improved the final pass-at-one performance by up to 3.9 % across the math and coding benchmarks. That is amazing. So the model that started the RL phase looking slightly worse on greedy accuracy ended the RL phase significantly smarter. It finally had the raw materials necessary to actually learn. And this brings us to a massive practical benefit for developers too, right? Computational efficiency.
17:02Huge benefit. Because TAIL-SFT preserves so much coverage, the GRPO algorithm had a rich set of rewarding responses to pull from immediately. It didn't have to wander around in the dark. Exactly. It didn't have to spend thousands of compute hours randomly mutating a homogenous output, hoping to accidentally stumble upon a better reasoning path. The complex answers were already simmering, right? below the surface as a result the early training reward during the RL phase rose up to 2.5 times faster for the tail SFT models compared to the standard models 2.5 times faster I mean in an industry where compute time costs tens of thousands of dollars an hour that is a massive operational advantage it really is but you know with results this profound I imagine the temptation is to treat tail SFT as a magic bullet like let's just apply it to every model every data set, every single training run?
17:53Well, the research is very careful to point out that this isn't a blind fix. It's a targeted intervention. Exactly. And thankfully, they developed a very lightweight diagnostic tool to determine mathematically exactly when TAIL-SFT will actually help. Oh, that's super useful. Yeah, it relies on a metric called the coverage ratio, denoted as ROE 16. ROE 16. Let's break down the math of this coverage ratio for the listener so they can apply it. It's essentially a balance sheet of what happens during a standard training run, right? Basically, yeah. You take your base model and you measure its initial coverage.
18:25Then you run a standard SFT training pass on it. You then look at what standard SFT destroyed. You measure the base reachable coverage that was lost and you divide it by the new coverage that was gained. Coverage lost divided by coverage gained. Right. And if that ratio is greater than one, meaning the standard SFT process is actively destroying more coverage than it is creating, tail SFT is highly likely to fix the problem. It is a remarkably accurate predictor, actually. Across the testing environments, out of 11 specific settings where that ratio is greater than one, 10 saw positive coverage gains by switching to tail SFT.
19:01Wow, 10 out of 11. And the 11th was statistically unchanged. None of them degraded. That's incredible. It's a clean, empirical diagnostic. Just run a standard pass, check your balance sheet, and if the ratio is above one, swap in tail SFT to unblock your reinforcement learning phase. It's that straightforward. And if we connect this to the bigger picture, this diagnostic entirely validates a stage-aware approach to AI model development. Stage-aware. Yeah, we can no longer judge intermediate checkpoints by how well they perform on standalone local metrics like cross-entropy loss. We have to judge these intermediate steps strictly on how effectively they support the subsequent training.
19:41Because the objective that makes a model look strong at stage 2 might be the exact mathematical objective that poisons it for stage 3. Perfectly said. Let's take a breath and recap the journey we've been on today. We unmasked a hidden flaw in the transition between supervised fine-tuning and reinforcement learning. The SFT trap. Right, where optimizing for short-term confidence collapses a model's probability mass and destroys the broad menu of options it needs to truly reason. And we explored how tail SFT's dynamic loss masking acts as a highly effective countermeasure against that trap. By muting the weight updates on sequences the model has already mastered, it literally forces the AI to explore the under-modeled tail of the data.
20:22Retaining his probability mass across multiple reasoning paths. And we saw the real-world proof. By preserving that coverage, that pass at 16 variety tail SFT provides a vastly superior foundation for the GRPO phase. Unlocking massive reasoning games in complex math and code and learning up to 2.5 times faster. All monitored by a clean mathematical diagnostic in the coverage ratio. Exactly. So what does this all mean? For you, the listener, understanding TAIL-SFT is a profound reminder that optimizing for a short-term metric can fundamentally sabotage long-term potential. It really is. And it raises an important question, actually, one that pushes into the philosophy of artificial intelligence itself.
21:06Oh, let's hear it. Well, we usually think of forgetting in machine learning as a bug, right? We call it catastrophic forgetting when a model loses access to old data. Right. The goal is always maximum retention. Don't forget anything. But TAIL-SFT proves that actively preventing a model from anchoring on its most deeply ingrained knowledge is actually a prerequisite for advanced reasoning. Wow. So if we extrapolate this out toward the pursuit of artificial general intelligence, is active, controlled unlearning going to be a requirement? That is a wild thought. I mean, will a future superintelligence system need to possess a mechanism to deliberately drop its mastery of the mundane?
21:44It might have to. Like if a neural network is going to synthesize entirely new physics or solve unsolvable equations, perhaps it mathematically cannot do so unless it is forced to forget the majority paths that we human beings have already mapped out. It's entirely possible. Maybe true intelligence isn't just about accumulating knowledge. It's about having the architectural discipline to stop practicing the free throws. Well said. Thanks for joining us on this deep dive. We'll see you next time.
From the publisher
Researchers introduce TailSFT, a modified supervised fine-tuning algorithm designed to better prepare language models for subsequent reinforcement learning. Unlike standard fine-tuning that minimizes overall cross-entropy, TailSFT filters out sequences that the model has already mastered to focus training on the under-modeled "tail" of the data distribution. This approach prioritizes coverage, ensuring the model retains a diverse range of correct responses that reinforcement learning can later identify and amplify. Theoretical analysis and experiments on the OLMo-3 7B model demonstrate that TailSFT significantly boosts performance in math and coding tasks, particularly by improving pass@K metrics. Ultimately, the authors show that a higher-coverage initialization leads to faster learning and superior final accuracy after reinforcement learning. This work advocates for a stage-aware approach to AI development, where intermediate training phases are optimized specifically to benefit the next stage of the pipeline.




