The Coverage Principle: How Pre-Training Enables Post-Training

24 Oct 2025 · 16 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains why next-token pre-training that minimizes cross-entropy (or sequence-level KL) can still fail after fine-tuning for coding/reasoning, and introduces the “coverage principle.” It argues cross-entropy is anti-correlated with downstream performance because it averages over common tokens, while success depends on whether the model assigns enough probability mass to rare “golden ticket” answers.

Guest backgrounds

No guests are mentioned; it’s a solo “Deep Dive” discussion.

Key claims

Coverage profile (tail/CDF of log density ratio) predicts post-training potential; MLE implicitly optimizes coverage; cross-entropy worsens with sequence length H but coverage stays flat after convergence; coverage relates to “inherent variance” (few critical decision points).

Notable examples

Best-of-N sampling and RL fine-tuning depend on coverage; if golden answers have near-zero probability, BON fails even with many samples. Experiments on complex graph reasoning show KL grows with length while coverage does not. Training fixes: normalized SGD, truncated SGD for distillation, and coverage-aware checkpoint selection via a “simple tournament” minimax procedure.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Optimization Paradox

0:45 to 1:30

Discussion about the paradox of LLMs performing poorly in fine-tuning despite low loss during pre-training.

“But the data, well, the data often says otherwise.”

Cross-Entropy and Its Flaws

1:30 to 2:30

Exploration of cross-entropy loss and its misleading correlation with model performance.

“This, we argue, is the real predictor of post-training potential.”

Introducing Coverage Principle

2:30 to 4:50

Introduction of the coverage principle as a missing link between pre-training and post-training.

“But the tasks we actually care about, solving a hard math problem, writing genuinely new code, generating creative text, those depend on getting maybe just a few crucial tokens, right?”

Understanding Coverage Profile

4:50 to 6:20

Explanation of the coverage profile and its importance for model performance.

“The CDF, however, is much more sensitive to what's happening out in the tails.”

The Relationship Between Coverage and KL Divergence

6:20 to 7:20

Discussion about the statistical distinctions between KL divergence and coverage profile.

“Always a tricky factor when models have to generate longer outputs.”

Implications of Coverage on Model Training

7:20 to 8:30

How the coverage principle affects the training of large language models based on empirical evidence.

“KL divergence, it gets worse as the sequence length H increases.”

Practical Solutions for Improving Coverage

8:30 to 10:15

Discussion of practical methods to improve model training for better coverage.

“Okay, that sounds less arbitrary than just counting tokens.”

Enhancing Distillation Techniques

10:15 to 11:50

Improvement proposals for distillation methods to enhance coverage.

“The proposal is quite elegant, actually.”

Choosing Checkpoints Based on Coverage

11:50 to 13:25

A method for selecting model checkpoints that focuses on coverage rather than loss.

“And the key result here is that this method is shown to actually match the best possible theoretical rate for achieving good coverage via MLE.”

The Coverage Principle and Practical Steps

14:00 to 14:44

Learn how the coverage principle informs practical techniques for model training.

“That's the real link between pre-training and whether things like best event or RL fine tuning will actually work well downstream.”
Show all 12 chapters

Understanding Semantic Coverage

14:44 to 15:22

Explore the concept of semantic coverage and its implications for pre-training success.

“These are tangible techniques practitioners can start experimenting with.”

Measuring Pre-Training Success

15:22 to 15:54

Discover how to measure the success of pre-training beyond just probabilities.

“Ah, so not just did it assign non-zero probability to this exact string of good tokens, but maybe did it capture the underlying idea even if phrased differently?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're really cracking open something that trips up a lot of people working with large language models. It's this, well, optimization paradox. You know, why is it that an LLM that gets better at predicting the next word, lower loss, technically superior, often just, well, fails spectacularly when you try to fine tune it for specific things like coding or maybe complex reasoning? Yeah, it's exactly where the theory seems to just hit a wall with practice, doesn't it? You spend all this compute, all this time on the big pre-training phase. Huge amount, yeah. Where it's all about minimizing that cross-entropy loss.

0:38Then you go to post-training, maybe using RL, maybe best event scaling. And you'd think starting with a model that's aced the pre-training test would set you up for success. It feels logical. Right. But the data, well, the data often says otherwise. That metric we all look at, cross-entropy loss, it can sometimes be, believe it or not, anti-correlated with how well the model actually performs downstream. Anti-correlated. That's just wild. It's like training a marathon runner by measuring how fast they can, I don't know, peel potatoes and then being surprised they don't win the race. We're fundamentally using the wrong yardstick.

1:09That's a great analogy, actually. And that's precisely what this theoretical investigation we're looking at today digs into. It aims to really characterize that relationship properly. OK. And the key idea, the missing link, if you will, is something called the coverage principle. And alongside that, we're introducing a new way to measure things, the coverage profile. This, we argue, is the real predictor of post-training potential. Okay, coverage. So this is the secret ingredient, the thing that connects that massive pre-training effort to actually getting good results later. We need to unpack this.

1:44What is coverage exactly? And, you know, how can people working on these models actually optimize for it? Move beyond just chasing that cross-entropy score. All right. So let's start with the villain of the piece, maybe cross entropy or, you know, technically sequence level KL divergence. It's all about minimizing the average prediction error. Right. So that's the goal people aim for in pre-training. That's the standard approach. Yeah. And well, therein lies the problem. When you minimize the average, you're mostly rewarding the model for getting the easy common stuff right. Imagine like generating code.

2:17Maybe 90 percent of the tokens are boilerplate semicolons, brackets, common keywords. Super predictable. Cross-entropy loves when the model nails those, gives it a great score. So it gets really good at the boring bits, which inflates its score. But the tasks we actually care about, solving a hard math problem, writing genuinely new code, generating creative text, those depend on getting maybe just a few crucial tokens, right? The rare insightful ones, the golden ticket responses, let's call them. Exactly. That golden ticket response might have a vanishingly small probability according to the model.

2:54Maybe so low it practically never can be sampled. But the model can still have a fantastic overall cross-entropy score because it aced all the easy parts. Right. And that's where the coverage profile comes in. It's a refinement, you see. It specifically looks at whether the model assigns enough probability mass, even if it's small, to those rare but really high-quality possible responses. Ed asks, are the good answers, even the long shots, still somewhere in the realm of possibility for this model? Okay, and this isn't just theoretical navel gazing. You're saying there's a direct practical benefit for listeners here.

3:25Absolutely. The research shows that this coverage profile, let's call it cove-in, is actually necessary and sufficient for certain post-training methods to work well, specifically things like best of n or BON. Ah, BON, where you generate, say, 100 different outputs and pick the best one. Precisely. Bond's success relies entirely on the model, sometimes generating that high-quality answer within those 100 attempts. It's often measured by pass. At N, did you get at least one good one in N tries? And how well Bond works is a really strong indicator for how well reinforcement learning fine-tuning will go later.

4:02Because if the model has terrible coverage. If those golden tickets have effectively zero probability. Then Bond is useless. You can generate a million samples and never see a good one. And RL won't have anything good to reinforce. Exactly. You need that spark, that possibility for these methods to latch on to. So, okay, conceptually, what's the difference? They're both based on probabilities, right? Cross entropy and coverage. What's the statistical distinction? It's about how you look at the distribution of probabilities, or more technically, the log density ratio. KL divergence, or cross entropy, corresponds to the mean of that ratio.

4:36It averages everything out. Okay, the average. The coverage profile, on the other hand, corresponds to the cumulative distribution function, the CDF, of that same ratio. Now, forget the jargon for a second. Think of it like this. The mean KL divergence is easily dominated by the big, bulky middle of the distribution, all those common easy tokens. Right. The CDF, however, is much more sensitive to what's happening out in the tails. It specifically tells you about the probability of the worst-case errors, or in our case, the probability assigned to those rare high-value outputs. It focuses attention on whether the model keeps those crucial, difficult answers in play, even if they have low probability overall.

5:17All right, so we know what coverage is now and why it seems to matter more than cross-entropy, but why does pre-training, which explicitly optimizes cross-entropy, sometimes lead to good coverage anyway? That feels like the next piece of the puzzle. Right, and that brings us to the coverage principle. This is a really core finding in the research we're discussing. It basically says that the standard pre-training process, maximum likelihood estimation, or MLE, just predicting the next token, actually implicitly pushes the model towards having good coverage, even though the explicit loss function, cross-entropy, isn't a great measure of it.

5:51So it's like a happy accident or hidden benefit of MLE. You could say that. It's a fundamental property of the process. It helps explain why massive pre-training does often work, despite the flaws in how we typically measure its progress with cross-entropy. It's optimizing for coverage under the hood, so to speak. That's actually quite comforting, knowing that the default method has this beneficial side effect. It is, but there's always a but, isn't there? When people tried to formalize this link using the traditional scaling laws we use for LLMs, they hit a snag, a big one. Related to sequence length.

6:23Ah, sequence length. Let's call it H. Always a tricky factor when models have to generate longer outputs. things tend to break down. They really do. And the standard scaling laws that try to connect cable divergence to coverage made a pretty scary prediction. They suggested that the amount of compute you'd need at test time, like the N in best of N, would need to grow exponentially with the sequence length H. Exponentially. That sounds bad. Like really impractical. Exactly. It basically implies that building models that can handle long, complex tasks effectively is almost impossible if you're relying on those scaling laws, or at least incredibly inefficient.

6:59It paints a very bleak picture. And does the evidence back that up or does reality behave differently? Well, the research provides some clear empirical evidence showing this prediction is, let's just say, overly pessimistic. They looked at specific tasks like complex graph reasoning problems, which require understanding long sequences of steps. Okay. And sure enough, if you just track the sequence level KL divergence, it gets worse as the sequence length H increases. It seems to scale linearly with H. The model gets penalized just for generating longer outputs, essentially. Which seems wrong if the output needs to be long to be correct.

7:34Precisely. But here's the kicker. When they measured the coverage profile for those same models on those same tasks, they found it showed basically no dependence on the sequence length h once the model had converged. Wow. So KL divergence gets worse with length, but coverage stays flat. That's what they observed. The model had learned to cover the important decision points within the sequence, regardless of the total length. This suggests coverage generalizes much more effectively, faster and smarter, you might say, than cross-entropy. So coverage isn't fooled by the length. It cuts through the noise.

8:09It avoids that, what did the paper call it, spurious dependence on things like age? Exactly. It's not getting penalized for what they term missing mass, basically, all the predictable low-impact tokens in a long sequence. It focuses the generalization effort where it counts. And just to add a bit more technical detail, this leads to a more refined way of thinking about complexity. Instead of just raw sequence length H, the analysis suggests coverage bounds depend more on something called the inherent variance. Inherent variance. Okay, that sounds less arbitrary than just counting tokens. How should we think about that?

8:42What is it capturing? Think of it as maybe the effective complexity or the number of truly hard decisions in the sequence. Instead of just counting every single token H, Inherent Variance tries to quantify how many points in this sequence are actually pivotal, where there's high uncertainty or variation in what the correct next token should be, given the context so far. Ah, I see. So a thousand token sequence might only have, say, five really critical choice points. Exactly. And the Inherent Variance would reflect those five points, not the full thousand. This focus on the high variance decision critical tokens is why the coverage profile, which is sensitive to them, turns out to be a much better predictor of downstream success than metrics swamped by the sheer length H.

9:24This is fascinating. So the theory lines up. MLE implicitly optimizes for coverage. Coverage is what actually predicts downstream success, and it avoids the pitfalls of sequence length that plague cross entropy. But what does an engineer do with this information? If standard training methods like SGD still get tripped up by sequence length H in practice, how do we actually, you know, force the optimization to prioritize coverage? That's the crucial next step, moving from understanding to action. We need practical ways to tweak the optimization process itself, to really lean into this coverage principle and make sure our training directly encourages good coverage.

10:03Because you're right, standard stochastic gradient descent, while its guarantees regarding coverage, still unfortunately inherit that suboptimal dependence on age. So vanilla SGD isn't quite enough on its own. Even if MLE has the right underlying tendency, what's the first practical fix suggested? Okay, solution one, normalized SGD. The proposal is quite elegant, actually. It involves a modification to the standard SGD update rule where you normalize the gradient. Normalize it. How does that help? By normalizing, you essentially remove the influence of the magnitude of the gradient, which can get skewed by long sequences of easy predictions.

10:38This technique is shown, provably, to improve the resulting model's coverage by eliminating that problematic dependence on sequence length age. It helps the model achieve a learning rate in terms of coverage that's independent of how long the sequences are much closer to the ideal theoretical MLE guarantee. So you're basically telling the optimizer, look, don't get distracted by how long the sequence is. Just focus your updates on improving the quality where it matters, especially after those tricky high variance decision points. That's a great way to put it. You're preventing the signal from the few important tokens from being drowned out by the noise from the many easy ones.

11:15OK, that makes sense for general training. What about specific scenarios like distillation? Yeah. Where you have a powerful teacher model providing guidance. Ah, yes, the expert distillation setting. That offers even more opportunities. This leads to solution two, an improved gradient normalization specifically for distillation. When you have access to those token-level probabilities from a strong Peacher model, you can use a more sophisticated update rule. The paper proposes a specific type of truncated SGD update. Truncated. Yes. It essentially focuses the update even more sharply on the most informative parts of the sequence, guided by the teacher's probabilities.

11:55And the key result here is that this method is shown to actually match the best possible theoretical rate for achieving good coverage via MLE. It fully leverages the teacher's knowledge to make the student model's learning incredibly efficient and coverage focused. Wow. So potentially much faster convergence towards good coverage in that setting. Potentially, yes. All right. Two interventions on the training process itself. But there's also the immediate problem. You've finished training. You have multiple checkpoints. Which one do you actually use for fine-tuning? We know picking based on the lowest cross-entropy score might lead you astray.

12:28Exactly. That's a huge practical issue. This brings us to solution three, coverage-aware checkpoint selection. If minimizing KL divergence can point you to a bad checkpoint, you need a different way to choose. The paper proposes a neat method called the simple tournament procedure. A tournament. Okay, sounds more interesting than just looking at a loss curve. How does it work? It's actually pretty intuitive. Instead of comparing each model checkpoint against some absolute loss value, you compare the models against each other based on coverage. Ah, head-to-head. Precisely. The procedure works by finding the model checkpoint that performs best in a sort of minimax sense.

13:06You select the model that minimizes the maximum empirical coverage loss against any other candidate model in your set of checkpoints. Okay, so you're looking for the model that's least likely to be significantly outcovered by any of its rivals, the most robust one. That's the idea. You want the model that minimizes the risk that another candidate found a high-quality answer that it completely missed. This approach is guaranteed to find models with good, robust coverage properties, even if your validation data set isn't a perfect mirror of the real-world data distribution you'll face later. It's a much safer way to select your foundation model for post-training.

13:42Hashtag tag tag outro. Okay, let's try and bring this all together then. It feels like the really big takeaway here is shifting our focus away from just raw next token prediction accuracy measured by cross entropy and towards this idea of the coverage profile. That metric, the one that specifically asks, does the model keep the rare good answers in play? That's the real link between pre-training and whether things like best event or RL fine tuning will actually work well downstream. That's the core message. And the coverage principle gives us confidence that pre-training is doing something useful in this regard, even if implicitly.

14:18It's not just about predicting the average token well. It has this hidden, crucial bias towards covering the important tail of the distribution. Right. And knowing this isn't just academic, you've outlined practical steps. Using normalized SGD during training to fight that sequence length dependency. Using that advanced truncation method if you're doing distillation. And crucially, using that tournament method to pick the right checkpoint based on coverage, not just loss. Exactly. These are tangible techniques practitioners can start experimenting with. But, you know, while this research offers a really clean, mathematically grounded solution by focusing on the probabilities assigned by the model.

14:55The field is, as always, looking ahead. Oh, what's bubbling up next? Well, the source material itself gives a little hint. It briefly mentions the need to think about semantic coverage. Semantic coverage. Okay, that sounds deep. What does that imply? It implies moving beyond just the raw probability scores for specific token sequences. It raises a really interesting and challenging question for all of us, and maybe for you listening. How should we measure pre-training success? if we need to account for the actual meaning, the concepts, the representations the model has learned. Ah, so not just did it assign non-zero probability to this exact string of good tokens, but maybe did it capture the underlying idea even if phrased differently?

15:39Precisely. If the model understands the concept well enough to express it in multiple ways, how do we measure that kind of coverage? It's a much harder problem, probably involving looking at the model's internal activations, its embeddings. But cracking that could lead to a far deeper understanding of what these models actually learn and ultimately LLM intelligence itself. Definitely something to think about. Moving beyond surface probabilities to the meaning underneath. Well, that's a fascinating place to leave it. Food for thought as you evaluate or build your next language model. Thanks for joining us for The Deep Dive.

From the publisher

This paper provides a theoretical analysis of next-token prediction in language models, introducing the concept of the coverage profile ($\text{Cov}_N$) as a superior metric to cross-entropy for predicting downstream performance with Best-of-N (BoN) sampling. The authors establish a "coverage principle," demonstrating that maximum likelihood, or next-token prediction, implicitly optimizes the coverage profile, leading to faster generalization that avoids the spurious dependence on sequence length seen in cross-entropy/KL divergence. The research shows that achieving a good coverage profile is necessary and sufficient for BoN success and derives scaling laws relating cross-entropy to coverage, while also exploring various optimization methods like stochastic gradient descent (SGD) and gradient normalization to provably improve coverage bounds. Finally, the text proposes tournament-style estimators for selecting models with optimal coverage, particularly in scenarios where the true data distribution is unknown.

More from Best AI papers explained

All 475 episodes
The Coverage Principle: How Pre-Training Enables Post-TrainingBest AI papers explained · 16 min
Listen in VO