Deriving neural scaling laws from the statistics of natural language

15 Feb 2026 · 19 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Explains a research breakthrough claiming neural scaling laws can be derived from statistical properties of natural language, not from model architecture. It introduces two “magic numbers” from text—gamma (predictability via decay of conditional entropy) and beta (how quickly long-range word correlations fade)—and argues they determine a “prediction time horizon” that limits learning by how much signal is distinguishable from noise.

Guest backgrounds

No guest is named in the transcript; it’s a host-led discussion of a specific paper.

Key claims

Learning curves follow a power law; the scaling exponent alpha_d equals gamma/(2*beta). Models learn instantly once correlations are within the horizon; the bottleneck is data/noise, not model capacity.

Notable examples

“Tiny Stories” (gamma 0.34, beta 0.88) predicts exponent ~0.19 and matches training; “Wikitext” (gamma 0.27, beta 0.94) predicts ~0.14 and matches. “Scaling collapse” shows curves for different context lengths collapse after rescaling by gamma and beta.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Neural Scaling Laws

0:45 to 2:48

Learn about the predictable relationship between data size and AI performance.

“You know, you double the data, you get a very predictable drop in error.”

The Shift in Paradigm: Language over Code

2:48 to 4:45

Discover how the structure of language influences AI learning.

“For anyone listening who isn't, you know, training models in their basement, what does that actually look like in practice?”

Gamma: Predictability in Language

4:45 to 7:12

Understand how the predictability of text affects AI learning efficiency.

“So gamma is all about predictability, or if you want the really technical term, the decay of conditional entropy.”

Beta: The Invisible Strings of Language

7:12 to 9:37

Explore how the connections between words influence learning efficiency.

“And the central claim here is that these two numbers, which are properties of the book, not the reader, are basically the DNA of learning.”

The Prediction Time Horizon Explained

9:37 to 14:01

Learn how the prediction time horizon affects AI's ability to learn.

“That learning curve, that line on the graph we all obsess over, is actually just a map of the expanding prediction time horizon.”

Exploring Neural Scaling Laws

14:01 to 16:44

Learn about the relationship between prediction and data in AI models.

“It's basically proof that short-term prediction and long-term prediction are governed by the exact same mathematical law.”

The Limitations of Language Learning

16:47 to 18:28

Examine the constraints of learning within the structure of language.

“Gamma, how predictable things are, and beta, how far those connections stretch.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you hang around in the world of artificial intelligence long enough, or I mean even if you just scroll through Twitter these days, you are going to hear one phrase repeated repeated. It's almost like a mantra. Oh, yeah. Scale is all you need. It's basically the industry's biggest open secret. The hammer that makes every single problem look like a nail. And honestly, looking at the last few years, it is really hard to argue with it. The recipe seems so simple. You get a bigger computer. You build a bigger neural network. And most importantly, you just feed it more text. All the text. The entire internet, if you can get your hands on it.

0:36And what happens is, well, it's predictable. It's almost like a law of physics. the AI just gets smarter. Right. The error rate goes down. It follows what we call a power law. And it's incredibly consistent. You know, you double the data, you get a very predictable drop in error. But here's the thing that's always kind of bugged me. And I know it bugs a lot of people in this space. We know that it happens. We know if we spend, say,$100 million on compute, we get this much smarter. But until, well, until very recently, nobody could really tell you why the numbers are what they are. It's like we've been watching apples fall from trees for a decade, but nobody had figured out gravity yet.

1:15That's a perfect analogy. We have been engineering these massive systems based on pure observation. It's just brute force. We plot a line on a graph and we say, well, the line is going up, so I guess we throw more coal in the furnace. But the actual mechanism, the why of it all. Total black box. We've been operating in the dark. Well, that changes today. And that's why I'm so excited for this deep dive. We're looking at a massive breakthrough that essentially cracks the code. And the twist, and I love this part, is that we aren't looking at the AI anymore. We're not analyzing the silica. We're looking at the language itself.

1:52Exactly. That is the whole paradigm shift. It turns out the intelligence of a model, or I should say how fast it can learn, isn't just about the code. It's dictated by the mathematical shape of the text it's reading. The geometry of language. I mean, it sounds like something straight out of a sci-fi novel. It does, but it's really just pure statistics when you get down to it. So today we are going to unpack this. We're going to look at two magic numbers hidden in human language that basically determine the speed limit of AI. We're going to talk about something called the prediction time horizon, which kind of explains why models sometimes feel like they're hallucinating.

2:28And by the end of this, you'll understand why chat GPT learns the way it does, not because of how it's built, but because of how we write. A fascinating pivot. We're really moving from computer science to something more like the physics of information. Okay, so let's start with the status quo, just to really drive home why this is such a big deal. We mentioned neural scaling laws. For anyone listening who isn't, you know, training models in their basement, what does that actually look like in practice? So historically, if you wanted to know how, let's say, learnable a data set was, You had to do it the hard way.

3:02You take your model, you feed it the data, you run the training, which costs a fortune in electricity and hardware, and then you measure the error rate. So you have to build the plane and fly it just to see if it's going to crash. Precisely. And you do this for what we call the data-limited regime. That's where your model is huge, basically infinite capacity. But the real bottleneck is just the amount of text you have. You want to know, if I double my text, how much does my error drop? And that drop usually follows a straight line if you plot it on a special log graph. Exactly. But you only get that line after you've spent all the money.

3:37It's totally reactive. You're just measuring the result of the experiment. The analogy that I kept seeing, which I thought was just brilliant, was the car analogy. Oh, yes. So imagine you want to know how fast a race car can go. The old way is you put a driver in it, you hit the gas, and you measure its top speed on the track. Brute force testing. Exactly. But this new approach, this is like putting the car in a wind tunnel. You measure the aerodynamics, you measure the friction of the tires, and you use physics to calculate the top speed without ever turning the key. And in this metaphor, the aerodynamics is just the statistics of the text.

4:12That's it. The breakthrough is realizing that the learning curve of an AI is dictated by these statistical properties of the data. If you can measure the shape of the language, how the words connect to each other statistically, you can predict the learning curve of the AI before it reads a single sentence. That is a very bold claim. I can look at the stack of books and tell you exactly how smart an AI will get from reading it. So let's get into the mechanics. The research identifies two specific properties, or two magic numbers. Let's take them one by one. The first is called gamma. Gamma. What are we actually measuring here?

4:49So gamma is all about predictability, or if you want the really technical term, the decay of conditional entropy. Let's stick with predictability for a second. So this is about how surprising the next word is. Yeah. Think of it like a guessing game. If I give you just one word, say, and I ask you to guess the next word, it's incredibly hard, right? Impossible. It could be anything. It could be cat, president, end. The entropy or the surprise is very, very high. Okay. But if you give me a whole sentence. Right. If I say the quick brown fox jumps over there. Bumps a dog. Easy. Exactly. As the context grows, the surprise just plummets.

5:25The more you know about the past, the easier it is to predict the future. Which feels super intuitive. I mean, we all do that when we listen to someone speak. We do. But what gamma measures is how fast that surprise drops. Does the language get predictable right away? Or does it kind of stay mysterious for a long time? Gamma is the exponent, the number, that describes the steepness of that curve. So a high gamma means the text gets predictable really quickly. Yes. It means the patterns are right there on the surface. Got it. So gamma is the speed of predictability. Now the second number is beta.

5:59This one felt a little more, I don't know, abstract to me. The paper talks about power law decay of correlations. It sounds scary, I know, but it's actually quite beautiful. Beta measures the invisible strings that tie words together across time. Invisible strings. Think about reading a novel. If I introduce a character named Captain Ahab on page one, that influences the words I might use on page 50. I might use words like whale or shift or ocean. Right. You're probably not going to use spaceship or microchip if you're 50 pages deep into Moby Dick. Exactly. There's a correlation between word one and word 5000.

6:35But, and this is the key, that connection gets weaker the further apart they are. The connection between word one and word five is strong. The connection between word one and word 5000 is very weak. And beta measures how fast those connections fade away. So it's kind of the rate at which the past becomes irrelevant. In a sense, yeah. A high beta means the correlations die out really quickly. Words only care about their immediate neighbors. But a low beta means there are these long lingering connections that can stretch across entire chapters. I see. Okay, so gamma is how much does context help me guess?

7:09And beta is how far back does that context even matter? That's a great way to put it. And the central claim here is that these two numbers, which are properties of the book, not the reader, are basically the DNA of learning. They dictate the speed limit. But how? This is the part where my brain started to hurt a little. We have these stats about the text. How does a neural network actually lose them? We usually just say the AI learns patterns, but that feels a bit hand wavy now. It is hand wavy. And to understand this, we have to talk about something called the prediction time horizon. Which definitely sounds like something from a Christopher Nolan movie.

7:42It does, but it's the crucial insight. We need to stop thinking of AI as learning patterns and start thinking of it as fighting noise. Fighting noise. Okay, unpack that for me. Go back to those invisible strings, the connection between word one and word 100. We know there's a connection, but it's very, very weak. Right, because it's so far away. And because it's weak, it is mathematically very hard to tell it apart from just random chance from noise. If you only read a small amount of text, a statistical fluke looks exactly the same as a real pattern. So if I only read one book and I see the word king followed by burger three times, I might just think king always predicts burger.

8:24Exactly. That's noise. It's a false correlation. To know for sure that King predicts crown, 100 words later, you need to see it happen thousands and thousands of times. You need a massive amount of data to clear the fog. I like that, the fog analogy. So imagine the AI is standing in this thick fog. It can see the words right in front of its face perfectly. But words further back in the sentence, they're blurry. And words from a paragraph ago, they are totally invisible. They're hidden behind this prediction time horizon. And the only way to push that fog back. is to add more data. Adding more data extends the horizon.

8:59It allows the model to see correlations that were previously completely indistinguishable from noise. This implies something really interesting about the AI itself. It's almost saying the AI is a perfect learner. That is the kicker. This whole theory assumes that if a correlation is visible, if it's inside that horizon, the model learns it instantly and perfectly. The bottleneck isn't the model's brain. The bottleneck is the horizon. The model cannot learn what it cannot see. That completely flips the script. We usually think, oh, the model isn't smart enough yet. You're saying the models is fine.

9:34The data is just statistically too noisy until you have billions of tokens. Precisely. That learning curve, that line on the graph we all obsess over, is actually just a map of the expanding prediction time horizon. And guess what determines how fast that horizon expands. It has to be beta, doesn't it? The invisible strings. It's beta because beta tells us how weak those long range signals are. If beta is high, the signals fade incredibly fast. And that means you need an exponential amount of extra data just to see a little bit further into the past. So if a language has these really long, complex, subtle connections, it's harder to learn.

10:09Well, it's actually the opposite. If the correlations fade slowly, So a low beta. They're easier to spot from a distance. If they fade quickly, a high beta, they disappear into the noise almost immediately. Ah, I see. So a high beta is actually bad for learning efficiency? In terms of data efficiency, yes. Yeah. If the signal dies out that fast, you need a mountain of data just to recover it. Okay, so we have the theory. We've got the fog. We've got the horizon. But now for the put your money where your mouth is moment. The researchers actually came up with a formula. they claim they can predict the exact slope of that learning curve using only gamma and beta.

10:46It's the golden formula of this whole discovery. The scaling exponent, which we can call alpha d, is equal to gamma divided by two times beta. So alpha d gamma, two beta. That's the one. Simple division. It seems almost too simple. You take the predictability, you divide it by twice the correlation decay, and boom, you have the learning speed of a billion dollar AI. In theory, yes. But did it work? Because I can write a formula on a napkin, but reality is usually a lot messier. And this is where the work really shines. They didn't just theorize. They tested it. They took two very different types of text.

11:20First, they looked at a data set called tiny stories. I've heard of this. It's just what it sounds like, right? Stories written with the complexity of a children's book. Exactly. Tom has a ball. The ball is red. Very simple vocabulary, simple grammar. So they ran the stats on this text. They found a gamma of 0.34 and a beta of 0.88. Okay, so fairly predictable, and the correlations don't die out instantly. And then they plug those numbers right into the formula. So 0.34 divided by 2 times 0.88, that gives you a predicted exponent of roughly 0.19. Meaning that for every doubling of data, the error rate should drop by a very specific amount determined by that 0.19 slope.

12:00And when they actually trained a real AI on tidy stories, the loss curve matched the prediction perfectly. The slope was 0.19. Wow. That's incredible. It's like predicting the weather and getting the temperature right to the decimal point. But then they made it harder. They switched to Wikitext. So, you know, Wikipedia articles. Which is a totally different beast. You've got complex sentences, technical jargon. In 1942, the geopolitical situation was? A completely different geometry. Gamma was lower-bedged 0.27, so it's harder to guess the next word. And beta was a bit higher, 0.94. The signal-to-noise ratio is just tougher.

12:37So the formula would predict a different learning speed for Wikipedia. It predicted an exponent of about 0.14, a shallower learning curve. Which makes perfect sense. Wikipedia should be harder to learn than a toddler story. You need more data to get the same gain in intelligence. Exactly. And again, when they ran the training, the actual results lined up perfectly with the prediction. So just by scanning the text without ever training a single neural network, we knew exactly how hard it would be to learn. We knew the physics of the book. We knew the speed limit. There was one visual in the findings that you flagged for me as the mic drop moment, the scaling collapse.

13:13Can you explain what we're looking at there? Oh, this is for the data nerds, but it's so powerful. So usually when you analyze an AI, you look at its performance in different contexts. How well does it predict the next word given 10 words of history? Or 100 words? Or 1 ,000 words? And I'm assuming those are all different. It's easier to predict with more history. Right. So if you plot them on a graph, you get this messy spaghetti of different lines. But the theory says, wait, these are not different mechanisms. It's all the same mechanism, the prediction time horizon. It's all just clearing the fog.

13:45Exactly. So they took all those messy separate lines and they rescaled the axes of the graph using gamma and beta. They applied the magic numbers to the raw plot. And what happened? All the lines just collapsed onto a single beautiful perfect curve. Wow. It's basically proof that short-term prediction and long-term prediction are governed by the exact same mathematical law. It's all just a function of, do I have enough data to see this far back? That's the kind of unity physicists dream about, finding the one equation that explains the apple falling in the moon orbiting. It's a huge validation of the theory.

14:21If the variables were wrong, the lines wouldn't overlap. But the collapse is incredibly tight. I want to zoom out a bit here. We've been talking about language models, transformers. Does this apply to everything? If I build an AI to predict the stock market or, I don't know, to fold proteins, does this loss still hold? That is the billion-dollar question. The suggestion is that this math applies to what's called a universality class of deep learning models. Universality class? That sounds very grand. It's a physics term. It means that lots of different systems, even if they look different on the micro level, behave the same way on the macro level.

14:58They think transformers and probably most modern deep learning architectures fall into this class. But not everything. Not everything. They specifically point out that older, shallow networks don't follow this law. They get overwhelmed. They suffer from what's called the curse of dimensionality. So the old AIs were just too dumb to even see the horizon. Essentially, yes. They couldn't handle the complexity. But modern LLMs are in what they call a fast learning regime. They're so good at pattern recognition that the only thing holding them back is the noise in the data. They're maxing out the statistics of the language.

15:34This raises a really interesting, maybe slightly philosophical point. If the learning speed is dictated by gamma and beta, and gamma and beta are properties of, say, the English language. I think I know where you're going with this. Does that mean there's a fundamental speed limit to learning English? It implies exactly that. It implies that you cannot learn faster than the statistics that the language allow. It doesn't matter how smart your engineer is or if you have a quantum computer. The very structure of our data dictates the maximum pace of improvement. We are limited by the book, not by the reader.

16:08Precisely. Unless we fundamentally change the language itself, unless we start writing in a way that has a higher gamma or a lower beta, we are basically locked into these exponents. That is humbling. We always talk about AI as this infinite exponential curve, a rocket ship to godlike intelligence. But this suggests the curve has a shape that was written by us centuries ago when we started forming sentences. We're analyzing the physics of information now. We're not just coding anymore. We are discovering the natural laws of communication. So if we sum this all up, scale is all you need is true.

16:42But now we finally know why. It works because scale fights noise. And we know that the battle against that noise is determined by two numbers. Gamma, how predictable things are, and beta, how far those connections stretch. And the interplay between those two numbers gives us the prediction time horizon. The fog. The fog. And the learning curve is just us watching that fog clear step by step. It really changes how you look at a data set. I used to just see more text. Now I see a structure. I see density, connectivity. It makes you wonder about other types of data, too, like code. Or DNA. Or music.

17:15They must have their own gammas and betas. Oh, absolutely. I'd bet that code has a very different beta than poetry. The constraints are so much tighter, the correlations are stricter. Which might explain why these models get so good at coding so quickly. Maybe the physics of code is just more learnable than the physics of natural language. Here's a thought to leave you with them. If the learning curve is dictated by the text, and we've basically proven that human language has a specific efficiency limit, have we already hit the point of diminishing returns? That is the provocative question, isn't it?

17:48If the exponent is fixed by the language, then simply adding more data becomes exponentially more expensive for smaller and smaller gains. We've been asking how much data do we have, but maybe the right question is how much information is actually hidden in that data. Quality over quantity, but defined by math. We might be running out of the easy wins, but at least now, thanks to this insight, we finally have the map. It's a whole new way to see the world. And a massive thank you for guiding us through the geometry of language today. And to our listeners, next time you read a sentence, think about those invisible strings connecting the words.

18:24You're processing gamma and beta in real time. Thanks for listening to the Deep Dive. We'll see you next time.

From the publisher

This paper introduces the first theory capable of quantitatively predicting neural scaling law exponents for large language models based solely on the statistical properties of natural language. The researchers identify two primary drivers of performance: the decay of next-token conditional entropy as context length increases and the weakening of pairwise token correlations over time. By combining these metrics, they derive a first-principles formula that accurately forecasts how test loss improves with larger training datasets without requiring synthetic data or free parameters. Their theoretical predictions show a remarkable match with experimental results from GPT-2 and LLaMA-style models trained on the TinyStories and WikiText benchmarks. Ultimately, the study suggests that a model's learning efficiency is fundamentally governed by a data-dependent prediction horizon, where more data progressively unlocks the ability to utilize longer-range linguistic patterns.


More from Best AI papers explained

All 475 episodes
Deriving neural scaling laws from the statistics of natural languageBest AI papers explained · 19 min
Listen in VO