In short
Prescriptive scaling in AI—using a “capability boundary” (an S-curve) to predict model performance from pre-training compute, and to diagnose whether gains come from real reasoning or from issues like post-training failure or benchmark contamination.
Guest backgrounds
No guest names are provided in the transcript; the discussion is between two hosts/research commentators.
Key claims
Scaling follows a sigmoid, not a straight line. Knowledge-task capability boundaries are stable (2023–2025 models get similar scores at equal FLOPs), while reasoning-task boundaries shift upward/left (better and cheaper math). Post-training “unlocks” reasoning capacity; it adds little for knowledge but is crucial for math/instruction following. Adaptive sampling reconstructs the boundary using ~20% of models. Benchmark contamination is flagged if models exceed the predicted boundary.
Notable examples
MMLU Pro/Jeopardy-style knowledge; hard math benchmarks (Math Level 5, Competition Level Math); ifeval; AME 2025 (reported as clean, no contamination evidence).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Shift from Alchemy to Predictability
0:45 to 4:59
Discuss the transition from guesswork to predictability in AI model training.
“The idea that intelligence is just a function of raw scale.”
Understanding the Capability Boundary
4:59 to 9:45
Analyze the importance of focusing on high-performing models to identify AI limits.
“This is where you're spending money, you're training the model, but the performance just isn't moving.”
Knowledge vs. Reasoning Tasks
9:45 to 12:02
Examine the differences in performance and scalability for knowledge and reasoning tasks.
“because the research makes a very sharp distinction between a base model and a post-trained model.”
Post-Training Insights
12:02 to 14:00
Unpack the significance of post-training adjustments in AI model performance.
“A way to figure out exactly what is broken.”
Evaluating AI Model Integrity
14:00 to 15:07
Learn about the importance of detecting contamination in AI model training.
“This is the big shadow hanging over all AI benchmarks right now.”
Assessing Progress in AI Reasoning
15:07 to 15:40
Discover how recent benchmarks indicate genuine improvements in AI reasoning.
“The boundary is truly shifting up because of that human ingenuity we talked about earlier.”
The Future of AI: Reasoning vs Scale
15:40 to 16:52
Explore the nuanced relationship between funding, knowledge, and reasoning in AI development.
“So let's wrap this up by looking forward.”
Shifting Paradigms in AI Development
16:52 to 17:33
Understand the potential changes in AI economics and technology based on reasoning improvements.
“We go from mainframes back to personal computers.”
Transcript
Automatic transcript. May contain errors.0:00You know, if you look at the tech industry right now, really zoom out. It looks less like a software industry and more like heavy industrial manufacturing. Oh, absolutely. It's the industrial revolution of compute. The scale we're seeing is just staggering. Right. We are talking about pouring concrete, securing massive water rights for cooling, negotiating directly with nuclear power plants. Yeah. And the price tag matches that scale. Exactly. We are seeing capital expenditure in the hundreds of billions of dollars. And there is this palpable anxiety in the air specifically among the people writing those checks for these deep dives in AI.
0:38Because they're betting the farm. They really are. They are betting on a very specific premise, which is that if you just keep making the models bigger, you know, more GPUs, more data, more electricity, they're guaranteed to get smarter. Right. The famous scaling laws. Yeah. The idea that intelligence is just a function of raw scale. But the trillion dollar question that keeps coming up, the thing everyone is whispering about is, is that actually a law of physics? Or have we just been lucky so far? Exactly. Are we approaching a wall where throwing more money at the problem just stops working? That is the perfect place to start.
1:12Because for the last few years, training these massive AI models has felt a bit like alchemy. Alchemy. That's a good word for it. Yeah, you throw ingredients into a huge pot data, architecture, compute, you stir it with 100 ,000 GPUs, and you just hope gold comes out. And sometimes it does. Sometimes it does, but sometimes it doesn't. And you don't really know why until it's done. Which is terrifying if you just spend a billion dollars on the compute. Right. But the mission of our deep dive today is to explore a fundamental shift away from that alchemy. We are moving toward civil engineering. Okay, civil engineering.
1:50Yeah, the goal is to turn this into a strict discipline where you can predict exactly what will happen before you pour a single yard of concrete. And the term being used for this in the research we're looking at today is prescriptive scaling. Prescriptive scaling. It's a bit of a dense term, I know. It is, but the concept is incredibly powerful for anyone paying attention to this space. Exactly. It's the ability to look at your pre-training compute budget, literally saying, okay, I have this many FLOPs available, and predict with incredibly high precision exactly what downstream test scores that model is going to get.
2:25Before you even start the training run. Before you spend a single dollar on the compute. That completely shifts the paradigm from, let's see what happens, to here is the blueprint. Right. But to get to that level of predictability, you need data. And not just a little bit. You need to understand the behavior of thousands and thousands of models. And that is where things usually get really messy. Yeah. If you look at public leaderboards right now, like the open LLM leaderboard, it is absolute chaos. It's annoying. You see huge models performing terribly. Then you see tiny models punching way above their weight class.
3:01If you just tried to draw a trend line through that data, it would look like a shotgun blast. Right. You've got different code bases, different data mixtures, totally different levels of confidence in the actual engineering teams running the training. So how do we find a rule in all that noise? That is the crucial point in the analysis we're unpacking today. If you want to find the laws of physics for AI, you cannot look at the average model. Because the average model is dragged down by bad experiments. Exactly. You have to filter out the noise and look at the limit. The research calls this the capability boundary.
3:34The capability boundary. I like that term. It really implies there is a hard physical ceiling. Think of it like human athletics. If you wanted to scientifically determine the absolute maximum speed a human being can run. You wouldn't go to a local park and time everyone you see. No, you wouldn't time the dog walkers or people jogging in jeans. If you averaged all that data, you'd conclude that humans have a top speed of about six miles an hour. Right. To understand the actual limit, you have to look at the outliers. You look at the Olympic gold medalists. You look at the 98th percentile. And when you do that in AI, when you take thousands of models and brutally filter out everything except the absolute best performers for every given budget, that chaotic shotgun blast completely disappears.
4:18And a shape emerges. A very clear, highly predictable mathematical shape emerges. And it tells us a lot about the actual future of machine intelligence. Okay. So let's visualize this shape for the listener. Because for a long time, the baseline assumption the whole scaling law assumption was that it was just a straight line on linear you double the compute you get a fixed percentage better forever just a line going up and to the right and the data says no not quite the shape isn't a straight line at all it's a sigmoid an s-curve an s-curve so it has a start a middle and an end let's walk through those three phases because I think they explain so much of what we actually experience when we use these models phase one is the emergence phase right this is the bottom tail of the S, the line is essentially flat.
5:03This is where you're spending money, you're training the model, but the performance just isn't moving. The model just doesn't have enough capacity yet, not enough parameters to grasp the concept. It's basically just guessing. You can double your compute in this phase and see zero improvement. Which has to be incredibly frustrating for an engineering team. Oh, it is. But then you hit the tipping point. Phase two, linear scaling. This is the steep middle part of the S. And this is the era we have been living in for the last few years. The model groks the task. It suddenly gets it. Yes. And suddenly, every time you double your compute, your error rate drops predictably.
5:42It feels like magic. This is the gold rush phase. The era of scale is all you need. Precisely. It feels like you could just scale forever, and it will keep getting smarter at that same rapid pace. But the S curve has a top. Phase three. Saturation. The diminishing returns. The curve flattens out again at the top. You approach the ceiling of what is mathematically possible for that specific architecture on that specific task. So you could double your compute again, spend another$100 million, and you might only squeeze out a fraction of a percentage point in improvement. Exactly. It's a harsh reality check for the bigger is always better crowd.
6:19It proves that for any specific task, you cannot just brute force your way to perfection endlessly. It brings actual physics back into the equation. There is a capability boundary. And once you are near it, throwing more money at the problem yields almost zero value. But here is where this gets really nuanced. And this is probably the most important insight of this entire deep dive. Okay, lay it on us. We aren't just looking at one generic intelligence curve. The analysis broke this down by the type of task the model is doing. Yes, the tail of two ceilings. They looked at knowledge tasks versus reasoning tasks.
6:54And the difference between the two is stark. Let's start with knowledge. So we're talking about things like MMLU Pro, trivia, facts, general understanding of the world, the Jeopardy champion skills. Right. When they analyzed the capability boundary for those knowledge tasks, they looked at data from models trained in 2023, 2024, and 2025. And what did they find? They found that the boundary is incredibly stable. Stable? As in it hasn't moved at all. Barely an inch, meaning that a model trained in 2025 is not really any more efficient at learning facts than a model from 2023. If you take a model from two years ago and a model from today and you give them the exact same compute budget, the exact same number of FLOPs, they will achieve roughly the exact same score on knowledge benchmarks.
7:40That seems really counterintuitive because as users, we feel like the newest models are getting so much smarter. But on a per dollar per compute basis, we aren't actually getting better at memorization. We are just building bigger buckets. Exactly. We're moving further up the existing curve because we're spending more money, but the curve itself hasn't shifted. So knowledge appears to be largely a function of raw capacity. If you want to store more of the internet, you simply need a bigger hard drive. We haven't found a magic algorithm that compresses general knowledge significantly better than we could a few years ago.
8:14So for knowledge, scale really is the primary factor. If you want a genius encyclopedia, you just have to pay the massive compute bill. Yes. But then there is reasoning. This is the exception. The math stuff. Complex logic. Specifically, they looked at hard math benchmarks, like math level 5, competition level math. And here, the boundary is moving. It is shifting upward and to the left on the graph. Which means getting drastically better and cheaper. A model trained today is significantly better at math than a model trained two years ago, even if they used the exact same amount of compute. That is the game changer right there.
8:51That tells me that while knowing things is expensive and static, thinking about things is becoming far more efficient. It proves that reasoning is algorithmic. It's a skill. It's not just a storage capacity issue. We are finding better data recipes, better training objectives, better ways to teach the model logic. Human ingenuity is actually shifting the laws of physics for reasoning tasks. So if you are a check writer at a big tech company listening to this, this totally changes your roadmap. It has to. It means you can't just rely on building the biggest server farm on Earth. You have to invest heavily in the science of teaching these models.
9:27It implies that the future might not just be one giant model to rule them all. We might see a massive divergence. Where knowledge is handled by massive, expensive models. But high-level reasoning can be done by much smaller, highly efficient models that have just been taught incredibly well. Now, speaking of teaching, we need to talk about the difference between the raw brain and the educated brain, because the research makes a very sharp distinction between a base model and a post-trained model. This is a huge area of confusion in the public discourse. Yeah. We often assume that the post-training things like reinforcement learning from human feedback, instruction tuning, we assume that's what makes the model smart.
10:08Because that's the version you and I talk to. We talk to a chat interface. We don't interface with the raw base model straight off the cluster. Right. But the analysis shows something absolutely fascinating. For those knowledge tests we just talked about, the trivia, the facts post-training does almost nothing for the actual score. Really? Almost nothing? The gap is tiny. The base model already knows the answer. It'd pick it up during pre-training by reading Wikipedia or whatever. So what is the post-training doing then? It just teaches it to be polite and format the answer as a nice billeted list.
10:39But it doesn't add any underlying intelligence or knowledge. The raw neural network already has the facts. It just has terrible social skills. Basically, yeah. But for reasoning, for math and following complex instructions, the story is completely inverted. The gap between base and post-trained is huge. Massive. On complex instruction following tasks like ifeval or hard math, The base model often sits at the very bottom of the performance chart. It fails completely. So you spend a billion dollars training this massive statistical engine, and out of the box, it literally cannot solve a math problem.
11:14It has the potential to solve it. The capacity is there. You paid for that capacity during pre-training, but it is locked. Locked. The research uses that specific term, right? Unlocking. Unlocking, yes. It's a great visualization. Pre-training buys you the ceiling. It buys you the maximum possible intelligence for that model size, but the model starts on the floor. And post-training is the elevator that takes it up to the ceiling. Exactly. It aligns the model to actually use the capacity it has. It pushes the model up from the bottom of the chart to the capability boundary. That explains perfectly why we see so many open source-based models get released that look amazing on paper.
11:53But when you actually try to use them for logic or coding, they feel totally broken. Because they haven't been unlocked. They were sports cars without a transmission. And this brings us right back to why this entire framework is called prescriptive staling. It gives engineers a diagnostic tool. A way to figure out exactly what is broken. Right. Let's say you train a model and it scores 40 % on a math test. Without this framework, you don't really know if that's a good result or a bad result. Maybe 40 % is the absolute limit of physics for a model of that size. Or maybe the training run was a complete failure.
12:26You wouldn't know. But with the S-curve. You just look at the curve. If the math says a model of your exact compute budget should be getting 70 % and you are sitting at 40%, you know for a fact you have a post-trading problem. You haven't fully unlocked the potential you paid for. Exactly. But conversely, if the curve says the maximum possible score for your size is 42%. Then you're actually doing great. You're effectively at the limit. And you can stop wasting time and money trying to optimize code that is already practically perfect. That efficiency is so crucial right now. And speaking of efficiency, the data touched on a very clever way to actually map these curves without burning a fortune, because training thousands of models just to draw a chart seems incredibly wasteful.
13:14It would be. But they used a technique called adaptive sampling. It's brilliant in its simplicity. How does it work? Well, because we now know the shape is a sigmoid. We know it's an S. Right. Since we know the shape, we don't need to empirically test every single point along the line. We just need to find the knees of the curve, the inflection points where it bends. So the algorithm can figure out where those points are likely to be? Yes. They found they could mathematically reconstruct the entire capability boundary by training and testing only about 20 % of the models you'd normally need to look at.
13:45It's like connecting the dots, but you only need a few key dots to see the picture because you already know what the picture is supposed to look like. Exactly. And once you have that picture, once you have that highly accurate curve, you have something else very valuable, a lie detector. A lie detector for contamination. This is the big shadow hanging over all AI benchmarks right now. Contamination. Did the model actually learn to do the math or did it just memorize the test questions because they were accidentally scraped into its training data? Did it learn the underlying concept or did it just see the answer key beforehand?
14:21So how does the S curve help spot that? If you plot a new model on our graph and it falls right on the sigmoid curve, it's behaving according to the laws of physics. It's likely genuine intelligence. But if you see a model that suddenly jumps way above the boundary... Like impossibly high, given its small compute budget, that is a massive red flag. It's the student who gets C's all year and then suddenly gets 100 % on the final exam. It's highly suspicious. Very. But here's the good news for the industry. When they applied this exact diagnostic to some of the recent, really high-profile math competitions...
14:57Like the AME 2025 benchmark. Right, the AME 2025. The data actually looked clean. It did. There was no clear evidence of contamination. The models aren't just cheating. They are genuinely reasoning better. The boundary is truly shifting up because of that human ingenuity we talked about earlier. Better algorithms, better post-training. That is extremely reassuring. It validates that civil engineering approach. We are officially moving from a phase of explosive, chaotic discovery to a phase of measurement, predictability, and genuine, measurable improvement. We are maturing as an industry. We are finally understanding the actual relationship between the inputs, the money, the data, the electricity, and the output, which is the intelligence itself.
15:39It's a huge step forward. So let's wrap this up by looking forward. We started this deep dive with that trillion dollar anxiety. Does throwing more money at the problem actually work? And the answer is highly nuanced. If you want knowledge, yes, you have to pay the toll. That boundary is static. But if you want reasoning... Then money alone won't save you. You need better recipes. You need better post-training. You need to unlock the potential. That is where the real competition is happening right now. It's not just who has the most GPUs. No, it's who understands the physics of reasoning the best.
16:17It makes me think about the long-term trajectory for you, the listener, and for all of us using this tech. We tend to assume the future is just bigger and bigger data centers consuming entire cities worth of power. Right. But if the reasoning boundary keeps moving up, if we keep getting fundamentally better at teaching model to actually think... You're wondering if we is approaching the end of the brute force era. Exactly. Could the most powerful AI of 2030 be a relatively small model? One that doesn't need a dedicated nuclear power plant simply because we finally learned how to teach it properly rather than just forcing it to memorize the entire Internet.
16:49That is the ultimate goal, isn't it? If we can decouple reasoning from massive scale, the entire economics of AI changes completely. We go from mainframes back to personal computers. From alchemy to engineering and maybe eventually to elegance. I really like that. Elegance over scale. That seems like the perfect place to leave it. We've unpacked the physics, the S-curves, and the entire future of how these systems are built. Thanks for helping us navigate the deep dive today. Always a pleasure to be here. And to everyone listening, next time you see a headline about a billion-dollar data center being built, just remember the S-curve.
17:26It's not just about how big the engine is. It's about whether they've figured out how to unlock the gears. We'll see you next time.
From the publisher
This paper details the development of capability boundaries to predict the downstream performance of language models based on pre-training compute budgets. Researchers move beyond standard scaling laws by using high-quantile pinball loss and monotone splines to track the upper envelope of achievable results rather than average trends. This methodology addresses benchmark saturation, where static tests lose their ability to distinguish between top-tier models as capabilities grow. By focusing on a 98th percentile frontier, the framework creates a stable, probabilistic estimate of what competitive training pipelines can reliably achieve. Ultimately, the work offers a principled way for practitioners to manage resource allocation and evaluate model performance relative to an evolving global capability ceiling.




