Test-Time Scaling Makes Overtraining Compute-Optimal

7 Apr 2026 · 21 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains a new paper, “Test-Time Scaling Makes Overtraining Compute Optimal,” arguing that when models are allowed to use extra test-time compute to sample many candidate answers (high K), the optimal strategy is to pre-train smaller models far beyond Chinchilla’s 20:1 data-to-parameter rule.

Guest backgrounds

No guests are named; the episode is presented as a host-led discussion.

Key claims

Chinchilla laws optimized only training loss under a compute budget; they ignored inference cost and multi-sample search. Overtraining increases output-distribution entropy/coverage so repeated sampling finds correct multi-step solutions. The paper proposes “T-squared” scaling laws balancing N (model size), D (data), and K (samples), predicting loss and pass@k.

Notable examples

Meta Llama 2.7B trained on ~2T tokens (~290:1) and Google Gemma 7B on ~6T tokens (~857:1). Experiments: 21 new models (5M–901M params), benchmarks including LAMBADA, RCEZ, PsyQ, OpenBookQA, plus synthetic multi-step datasets from Claude/GPT. Results: ~2.8% prediction error; under fixed inference compute, overtrained small T-squared models beat Chinchilla-optimal models across tasks, and advantages persist after supervised fine-tuning on ARC-Easy, PsyQ, and OpenBookQA (though reduced).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of Test-Time Scaling

1:40 to 3:40

Explore the concept of test-time scaling and its implications for AI architecture.

“We have a fascinating new paper from researchers at UW-Madison and Stanford.”

Chinchilla Scaling Laws Explained

3:40 to 6:20

Understand the Chinchilla scaling laws and their significance in model training.

“Over the last year or so, we've watched these major tech labs just openly defy that math.”

Overtraining and Economic Logic

6:20 to 9:30

Discuss why tech companies are overtraining models beyond Chinchilla's recommendations.

“The T-squared scaling laws introduce this joint optimization framework.”

Introduction to T-squared Scaling Laws

9:30 to 12:30

Learn about T-squared scaling laws and their implications for model design.

“Right, but modeling this across a vast, varied dataset is incredibly tricky because of the natural variance in problem difficulty.”

Empirical Validation of T-squared Laws

12:30 to 14:00

Discover how researchers validated the T-squared laws through experiments.

“Pre-training 21 distinct architectures from scratch requires serious compute allocation.”

Understanding Multi-Step Search in AI

14:00 to 15:10

Explore how multi-step reasoning improves model performance.

“But they also needed to test tasks that strictly mandate multi-step search.”

Performance Comparison: T-Squared vs Chinchilla

15:10 to 17:31

Learn about the head-to-head comparison of overtrained models against optimal models.

“and they compared it head-to-head against a smaller, heavily overtrained T-squared model.”

Fine-Tuning and Its Challenges

17:31 to 19:12

Understand the impact of fine-tuning on overtrained models and their performance.

“There is an architectural friction that happens during fine-tuning with these specific models.”

The Shift in AI Model Architecture

19:12 to 20:29

Discover the implications of overtraining on the future architecture of AI models.

“We've traced the mathematics from NLL power laws to beta distributions, and we've mapped the physics of overtrained checkpoint surviving alignment.”

The Future of AI and Local Computing

20:29 to 21:28

Contemplate the potential of advanced AI running directly on local devices.

“The hardware limitations of the phone sitting in your pocket right now might suddenly be irrelevant.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine for a second that you're hiring an assistant and you've got two candidates sitting right in front of you OK. So the first candidate is just incredibly fast, like lightning fast. You ask them a really complex multi-step question and they instantly, I mean, the millisecond you finish talking, blurt out the very first answer that pops into their head. Right. No hesitation at all. Exactly. And, you know, sometimes their logic is actually brilliant. But other times they completely hallucinate a random detail that just derails the entire solution. But hey, they are undeniably quick. Fast but risky.

0:35Yeah, exactly. Now, the second candidate operates totally differently. You ask them the same exact question and they pause. They pull out a notebook. They draft, I don't know, a hundred different structural approaches to the answer. They cross out the logical dead ends. They combine the best paths. And then, sure, after a little bit of time, they hand you the absolute best, most rigorously vetted response possible. I mean, if your work relies on any kind of mission-critical accuracy like complex reasoning or coding, you're definitely hiring that second candidate. 100%. You want the one that's configured to utilize that extra time, that test time compute, to really sample and verify multiple paths.

1:15Right. But the crazy thing is, for a really long time, the artificial intelligence industry has basically been obsessed with scaling up that first type of candidate. Building those massive monolithic models. Yeah, these giant models designed to just execute a single zero shot forward past the exact millisecond you hit enter. But that entire paradigm is kind of fracturing right now. It really is, yeah. Which brings us to today's topic. We have a fascinating new paper from researchers at UW-Madison and Stanford. It's titled Test Time Scaling Makes Overtraining Compute Optimal. And in this deep dive into the research, we are going to explore the mechanics behind what is essentially a massive revolution in AI architecture.

1:57Moving away from those giant single-pass models. Exactly. Moving toward smaller, hyper-trained models that essentially run these deep, iterative search algorithms before they even respond to you. Okay, let's untack this. Because before we can get to the new framework, we really have to look at the baseline constraints that the industry has been rigidly adhering to for a while now. Right. To really understand this shift, you have to look at the chinchilla scaling laws from back in 2022. The famous chinchilla rules. Yeah, because for years that paper provided the, well, the ironclad mathematical consensus for optimal model pre-training.

2:33It established this very precise equilibrium. The rule was basically that for every single parameter in your model's architecture, you should allocate approximately 20 tokens of training data. Right. The 20 to 1 ratio. Exactly. So it's kind of like a baking recipe, right? Like the idea was that scaling the parameters, the size of the brain and scaling the data you feed it had to move in lockstep. Yeah, like flour and sugar. Right. If you add way too much flour, meaning the data without enough sugar, the parameters, the cake is supposed to be completely ruined. The math suggested that if you just kept pumping trillions of tokens into a fixed-size architecture, the model's loss curve would just asymptote.

3:13It would just hit a wall. Yeah, a wall where the parameter count simply couldn't compress any more generalized representations from all that extra data, meaning you were literally just burning expensive GPU clusters for zero gain. That was the accepted truth. I mean, the Chinchilla framework was entirely predicated on minimizing the pre-training loss under a strict compute budget. You only have so much money to spend on training, so you follow the 20 to 1 rule to get the best model possible for that budget. But then reality hit, right. Over the last year or so, we've watched these major tech labs just openly defy that math.

3:46Oh, completely. Look at Meta. They dropped Llama 2.7b. So they kept that 7 billion parameter architecture, but they pushed the training data to 2 trillion tokens. Which is, wait, let me do the math. That's a roughly 290 to 1 ratio. Right. And then Google released Gemma 7B and they pushed the envelope even further. They trained that on a staggering six trillion tokens. Six trillion. So that's an 857 to one ratio. Exactly. They completely blew past the chinchilla recipe. OK, I'm trying to map out the economic logic here because at face value, Google and Meta are intentionally burning millions of dollars on compute that the chinchilla rule literally dictates is suboptimal.

4:30They are. But they aren't doing it out of ignorance, obviously. They have the smartest engineers in the world. So why intentionally overtrain these smaller models? What's fascinating here is that the original scaling laws optimized purely for the training phase. Oh, interesting. Yeah. They treated the model's creation as the finish line. They completely ignored the lifecycle cost of actually using it, the inference cost. Oh, okay. The deployment reality. Exactly. When you deploy a model to millions of users running a 100 billion parameter model for every single prop that comes in, That is a massive ongoing operational expense.

5:01Right, because inference costs scales linearly with the parameter count. If you have a smaller parameter count, your memory bandwidth constraints and your compute overhead during deployment drop drastically. Precisely. So by aggressively overtraining a small architecture, sure, you spend heavily up front. You move way past that chinchilla optimal point. But you yield this dense, highly capable model that costs a fraction of a cent to run at inference for the end user. Okay, that makes perfect economic sense. But the research we're looking at today takes that economic reality and fuses it with a completely new functional paradigm, right?

5:39It does. it's not just about making that single pass zero shot inference cheaper anymore. Right. It's about architecting models specifically designed to generate multiple samples to solve hard problems. Which brings us to the core contribution of this UW-Madison and Stanford paper, the T-squared or train to test scaling laws. Yes, T-squared. And the premise here completely flips the script. It says, hey, if we know in advance that we aren't going to use the model for a single zero shot answer, If we know we're going to allocate test time compute to let it think, to let it sample k times to find the best path, then the underlying math of how we build the base model has to fundamentally change.

6:18Because the actual mechanics of multi-sample generation require a totally different type of latent space. Right. The T-squared scaling laws introduce this joint optimization framework. It balances three variables under one unified compute budget. You've got N for model size, D for training data, and crucially, K for the number of inference samples. Hold on. I want to pause on the mechanics of K for a second because this is crucial for the listener. Why does overtraining specifically benefit a model when it is allowed to sample multiple times? Right. That's the big question. Like, I get why a small model is cheaper to run a thousand times.

6:55but mathematically speaking, why does a hyper-trained small model yield a better variety of accurate samples than a larger model that was strained to the chinchilla standard? It really comes down to the entropy of the output distribution. The entropy. Yeah. When you deeply overtrain a smaller model, when you expose it to vast amounts of diverse data that go far beyond its theoretical capacity, you force its internal representations to become incredibly dense and smooth. Okay, so it's packing a ton of context into a tiny space. Exactly. It internalizes so many weird edge cases that its probability distributions across the logits become very, very well calibrated.

7:36Right. So when you crank up the temperature and you sample from that model 1 ,000 times, you aren't just getting random noise. And you aren't just getting the exact same answer slightly rephrased over and over. Oh, wow. So you get highly distinct, logically valid trajectories through the problem space. Exactly. Ah, so you're actually broadening the search tree. A larger chinchilla optimal model might just ruthlessly converge on one single path, but the dense overtrain model has this richer manifold of potential solutions to pull from. That's a great way to visualize it. It has more paths to explore.

8:10That makes so much sense. So how did the researchers actually mathematically map this out? Because, you know, optimizing N, D, and K simultaneously, that sounds like a non-convex nightmare. Oh, it is a nightmare, which is why they had to develop two distinct mathematical approaches to solve it. So approach one is based on modifying the continuous loss function, specifically the negative log likelihood or NLL. OK, NLL. That's the traditional metric for basically how surprised a model is by the next token. Right, exactly. They took the standard Chinchilla power law formula, which predicts NLL based on the model size and the data.

8:47And they introduced a new subtractive power law term that is driven by K. Driven by the number of samples. Yeah. This mathematically formalizes the intuition that as your sample budget K increases, your expected loss actually decreases. Okay, I follow that. But calculating NLL doesn't always translate perfectly to actual downstream accuracy. Like being slightly less surprised by a token doesn't guarantee the AI can actually solve a multi-step logic puzzle. That is the exact problem with NLL, yes. And that's where approach two comes in. Approach 2 tries to model the actual task accuracy directly.

9:21They specifically looked at a metric called pass at k. Pass at k, meaning the probability that at least one of your k-generated samples actually contains the correct answer. Right, but modeling this across a vast, varied dataset is incredibly tricky because of the natural variance in problem difficulty. Oh, sure. Because if a question in the data set is purely deterministic and easily retrievable-like, what is the capital of France? Sampling it 100 times does nothing. Right. The model either knows Paris on the first try or it doesn't. Exactly. The probability distribution is basically a flat line or a really sharp spike.

9:58But if the task requires multi-hop reasoning like a tough math problem, the probability curve changes entirely. Repeated attempts actually help there. Exactly. And that is why approach two utilizes beta distribution. Beta distribution. Yeah, the researchers use beta distributions because they are these remarkably flexible probability distributions defined by two shape parameters, alpha and beta. By fitting these parameters to the empirical data, the researchers could mathematically model the underlying variance in the quote unquote true pass probability across all the different prompts in a data set.

10:30Oh, I see. So instead of treating a data set like one giant monolithic difficulty level, the beta distribution essentially maps the terrain. Right. It maps the terrain perfectly. It accounts for the fact that some questions have a steep probability gradient where repeated sampling is highly effective while others are just flat. So by integrating that distribution, Approach 2 predicts the true expected pass at k for an architecture before you even allocate the compute to train it. And the implications of that specific formula are just staggering. How so? The mathematics show definitively that as your inference budget K grows, the optimal allocation of your compute shifts violently away from model size and directly into training data.

11:13Wait, really? Yes. If you know you are going to sample heavily at test time, the math proves you must design a much smaller architecture and just saturate it with data from day one. Wait, so if we know in advance that the AI is going to try to answer a question 1 ,000 times before giving you the final answer, this math tells us we should literally design its brain differently from the very beginning. Exactly. You intentionally build a smaller brain and give it way more books to read. That is a profound shift. We're talking about restructuring the fundamental anatomy of the base model based on an anticipated test time search algorithm.

11:49But, you know, mathematical proofs are just that. They're just proofs on paper. You have to validate them against the physics of actual neural networks. You do. And the empirical validation in this research is where the scale of their work really shines. It's incredible. Yeah, the empirical test bed they built is absolutely massive. They evaluated 106 different model checkpoints, ranging from 5 million all the way to 901 million parameters. But the critical thing here is that they couldn't just scrape existing checkpoints from Hugging Face to prove this math. Right. The open source ecosystem simply didn't have models that were systematically overtrained to the extreme degrees required to map this new T-squared curve.

12:29So to actually test their scaling laws, the researchers had to literally pre-train 21 brand new models from scratch and push them deep, deep into the overtraining regime. Which is wild. Pre-training 21 distinct architectures from scratch requires serious compute allocation. It's a huge investment. Yeah. They used their t-squared math to chart out exactly what performance these theoretical, heavily over-indexed architectures should achieve, and then they ran the clusters to see if the real-world loss curves actually aligned with their equations. You know, here's where it gets really interesting for me.

13:04It reminds me of early explorers or even like mapping aerodynamic drag in structural engineering. How so? Well, Chinchilla mapped the fluid dynamics for single pass efficiency, right? And everyone just assumed that if you pushed an architecture beyond that 20 to 1 ratio, the drag coefficient would stall the model's learning completely. You'd fall off the edge of the map. Right. You hit the wall. But these researchers essentially sailed off the edge of the Chinchilla map into these uncharted waters of overtraining. Their new T-squared compass predicted a completely new aerodynamic envelope that only opens up when you introduce high-frequency sampling.

13:40So they built the 21 models, put them in the wind tunnel, and tested them to see what they'd find. And they put them through a real gauntlet of benchmarks. They evaluated them on eight distinct tasks to ensure the findings weren't brittle. Right. They used standard real-world tests like Lambea for language modeling and RCEZ, PsyQ, and OpenBookQA for factual reasoning. But they also needed to test tasks that strictly mandate multi-step search. Because, as we said earlier, simple factual recall doesn't benefit from case sampling as much as raw reasoning does. Exactly. So they incorporated these synthetic data sets generated by Claude and GPT, focusing heavily on multi-step arithmetic, common sense causal reasoning, and spatial reasoning.

14:21And the results. I mean, And the math worked brilliantly. Approach 1, that model, based on modifying the negative log likelihood, was wildly accurate. It really was. When they compared the theoretical performance predicted by the T-squared formulas against the actual empirical results of those 21 newly trained models, the error rate was only 2.8%. Yeah, 2.8 % is an incredible level of precision. To predict the downstream loss of an unbuilt architecture in an untested over-trading regime with that small of a margin of error completely validates the underlying equations. It's mind-blowing. But the real litmus test, like the actual so what for the listener and the industry, is how these models performed against the chinchilla standard when the playing field was truly leveled.

15:07This is the crux of the whole paper. They took a theoretically perfect chinchilla optimal model, and they compared it head-to-head against a smaller, heavily overtrained T-squared model. Okay. And to make it a fair fight, they fixed the total inference compute budget. Right. Meaning because the T-squared model is so much smaller, it can execute way more forward passes for the exact same amount of compute that the larger model uses for just a few passes. Exactly. When constrained by the same inference compute budget, the smaller overtrained model utilizing high K sampling systematically outperformed the larger Chinchilla Optimal model.

15:44Wow. In every single test. Every single one. That's amazing. It wasn't just a marginal victory in multi-step arithmetic. Across spatial reasoning, causal logic, and even factual recall, the smaller architecture that was saturated with data utilizing repeated sampling dominated the efficiency curve. It proves that if you constrain computed deployment, the optimal strategy is unequivocally to overtrain a smaller network. Always. Okay. I do have to push back on this, though, because everything we've discussed so far applies to base models. True. The raw pre-trained weights straight out of the cluster.

16:19But as you know, no one actually deploys raw base models to end users. We interact with models that have undergone supervised fine-tuning or RLHF to act like a helpful assistant, format responses, that kind of thing. Alignment. Right. So my question is, does this Massey overtraining advantage actually survive the real world? Does it survive the gradient updates of the fine-tuning phase? That is the critical variable. Because if the structural advantages of overtraining collapse the moment you apply a fine-tuning data set, then the entire T-squared framework is totally useless for production. Right.

16:54The researchers totally understood this, so they explicitly tested the post-training phase. So they took their overtrained checkpoints and ran them through standard supervised fine-tuning specifically on three tasks, right? Right. ARC-Easy, PsyQ, and Open Book QA, just to see if the representations helped. And the verdict. Yes. The data confirms that the T-squared scaling advantages absolutely survive the post-training pipeline. The best small, overtrained checkpoints maintain their systematic superiority over the Chinchilla standards, even after alignment. That's a huge relief for the viability of this whole approach.

17:26But I noticed the research points out a fascinating nuance. The margin of superiority is slightly subdued after fine-tuning. There is an architectural friction that happens during fine-tuning with these specific models. Yeah, if we connect this to the bigger picture, it actually makes perfect sense mechanically. Heavily overtrained models suffer from a kind of optimization stubbornness. Optimization stubbornness? I'm trying to picture the loss landscape here. Think about it this way. If a small model has been bombarded with 6 trillion tokens, its weights aren't just loosely settled. They are entrenched deep inside extremely sharp local minima.

18:05Its internal representations of language and logic are intensely reinforced. Oh, okay. Because the architecture is small relative to the massive amount of data, the parameter weights have been aggressively consolidated. Exactly. So when you introduce a supervised fine-tuning phase, which typically uses a much smaller learning rate and a very narrow data set to teach the model how to, say, format bullet points or adopt a polite tone, the gradient updates struggle to nudge those entrenched weights. The model fundamentally resists the updates because its base distribution is just so cemented, it takes significantly more computational force to realign a network that dense.

18:41Right. Previous research on catastrophic forgetting and fine-tuning dynamics has hinted at this for a while. Over-indexing on pre-training data makes alignment mathematically harder. It makes them stubborn. Very stubborn. But what this specific paper isolates is that despite this stiffness in the latent space, the structural advantages of the T-squared architecture still outweigh the fine-tuning penalty. So even with that friction, they still come out on top. Exactly. The performance floor of the overtrained model is so high that even a slightly inefficient fine-tuning phase yields a superior final product.

19:17Wow. So what does this all mean? We've traced the mathematics from NLL power laws to beta distributions, and we've mapped the physics of overtrained checkpoint surviving alignment. But distilling this all down, what is the ultimate takeaway for you, the listener? Well, it signals a complete restructuring of how labs will allocate compute moving forward. If the goal is to build systems capable of deep reasoning systems that verify, search, and iterate at test time, the old 20-to-1 chinchilla rule is dead. Just throw it out. Throw it out. The new baseline requires engineering significantly smaller architectures, pushing the pre-training data to the absolute limits of availability, and shifting the intelligence burden from single-pass memorization to multi-pass search.

20:03It totally redefines what we consider optimal. The future of AI isn't strictly about parameter bloat anymore. It's not about relying solely on these massive trillion-parameter leviathans that require dedicated power grids just for a single forward pass. No, it's about extreme density. It's about leveraging the artificial patience of search algorithms to extract complex reasoning from much smaller footprints. Which leaves us with a really fascinating lingering thought to consider as this whole thing plays out. If the frontier of reasoning is no longer gatekept by the necessity of deploying these massive energy guzzling behemoth models, if the new paradigm proves that world class logic can emerge from tiny, highly overtrained models that just think thousands of times in the background, what does that mean for our everyday devices?

20:51It changes everything. Right. The hardware limitations of the phone sitting in your pocket right now might suddenly be irrelevant. It might not need the VRAM to hold a massive brain. It just needs the architectural density and the computational patience to search the latent space until it finds the absolute perfect path forward to answer you. It strongly suggests the barrier to local, high-level frontier reasoning is about to drop significantly. The era of the deliberate, iterative assistant might just run entirely locally on your phone. That is an incredible thought to leave on. Thank you so much for joining us on this deep dive into the research.

21:28We'll catch you next time.

From the publisher

Researchers from the University of Wisconsin-Madison and Stanford University propose Train-to-Test (T2) scaling laws to optimize the development and deployment of Large Language Models. Traditional scaling methods like Chinchilla focus primarily on pretraining efficiency, whereas T2 scaling jointly considers model size, training duration, and the compute required for repeated sampling at test-time. The study reveals that when accounting for these inference costs, the most effective strategy shifts toward extreme overtraining, which involves training smaller models on significantly more data than previously recommended. Small, overtrained models often outperform larger counterparts because they allow for more inference samples within the same total compute budget. The authors demonstrate that these T2 scaling predictions remain accurate and beneficial even after models undergo post-training processes like fine-tuning. Ultimately, the work provides a new blueprint for practitioners to maximize performance by balancing training investments with modern test-time scaling strategies.

More from Best AI papers explained

All 475 episodes
Test-Time Scaling Makes Overtraining Compute-OptimalBest AI papers explained · 21 min
Listen in VO