In short
How reasoning emerges from the standard AI pipeline (pretraining + supervised fine-tuning + reinforcement learning), using chess and math as controlled testbeds; includes scaling laws, compute budgeting, and failure modes of RL.
Guests
No guest names or backgrounds are provided in the transcript.
Key claims
Pretraining sets the performance ceiling and RL learning speed; RL can “surface” correct moves buried in the probability tail, but can also “amplify” incorrect moves (“wrong mode amplification”), improving pass@1 while harming pass@16. RL broadens short-horizon search but struggles with deep multi-step planning.
Notable examples
Chess models (5M–1B params) pretrained on 54B human move tokens; RL rewards only exact full solution lines (binary 1/0). Metrics: pass@1 vs pass@16; phenomena: ground-truth amplification, tail discovery, wrong mode amplification (balloon-squeezing). Transfer test: ALMO2 (1B params) on up to 200B math tokens shows similar scaling and RL behavior; RL improves candidate breadth but not long-horizon depth.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI's Learning Mechanisms
0:46 to 2:07
Exploring the complexities of AI learning, focusing on pre-training and reinforcement learning.
“And then you have reinforcement learning, or RL, where the model learns through trial, error, and outcome-based feedback.”
Chess as a Controlled Environment
2:08 to 3:40
Using chess to analyze the reasoning processes of AI models in a controlled setting.
“We're going to shrink things down to a highly controlled, incredibly specific laboratory.”
The Limitations of Pre-training
3:41 to 4:52
Discussing how pre-training prepares AI, but doesn't develop problem-solving skills.
“Researchers took AI models ranging from 5 million to 1 billion parameters.”
The Role of Reinforcement Learning
4:53 to 6:56
Examining how reinforcement learning builds on pre-training to enhance AI reasoning.
“It builds an intuition for what a normal chess game looks like, but it absolutely does not teach problem solving.”
The Impact of Reasoning Traces
6:57 to 8:13
Analyzing how reasoning traces affect AI performance across different metrics.
“A single mistake anywhere in the chain yields a zero reward.”
Budgeting for AI Training
8:14 to 10:33
Understanding the financial challenges of balancing pre-training and reinforcement learning.
“Yeah, it improves pass at 1, pass at 4, pass at 16.”
Probabilities and Reinforcement Learning
10:34 to 12:16
Delving into how reinforcement learning adjusts the probability of moves in AI decision-making.
“The compute optimal frontier actually shifts based on the scale of the model you are building.”
Ground Truth vs. Wrong Mode Amplification
12:17 to 14:00
Exploring the dual effects of reinforcement learning on AI's problem-solving capabilities.
“And the answer depends entirely on the difficulty of the puzzle it is facing.”
Understanding Wrong Mode Amplification
14:00 to 16:30
Explore the concept of wrong mode amplification in AI reasoning.
“a phenomenon known as wrong mode amplification.”
Applying Findings to Language Models
16:30 to 18:04
Learn how findings from chess AI apply to language models in math tasks.
“Okay, but if you're listening to this and wondering why a chess experiment matters for the chatbot you use at work every day, this is where the math jumps out of the 64 squares and into human language.”
Show all 12 chapters
Limitations of Reinforcement Learning
18:04 to 19:13
Discover the depth limitations of reinforcement learning in AI reasoning.
“By meticulously analyzing the chain of thought generated during these math tasks, we can see that RL makes the model search much broader.”
Balancing AI Training Methods
19:13 to 21:48
Understand the balance needed in AI training between pre-training and RL.
“It's exactly like a chess player who is an absolute genius at short-term tactics, you know, spotting immediate traps and capturing pieces, but has absolutely no long-term grand strategy for how to actually win the game.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Right now, you know, the AI industry is pouring just billions of dollars into building a machine that can actually reason. Oh, absolutely. Billions. But here is the multi-billion dollar secret. They're basically throwing darts in the dark. Yeah, that's a good way to put it. I mean, the exact mechanism of how an AI learns to think, rather than just like parrot information back to you, is one of the biggest mysteries in computer science today. And today, we're going to solve it. Well, we'll try. But yeah, it is the absolute definition of diagnostic muddy waters. Right. We know the standard training pipeline has two main phases.
0:37First, you have pre-training, where the AI reads vast amounts of data to learn the basics of language and logic. It's reading the whole internet basics. Exactly. And then you have reinforcement learning, or RL, where the model learns through trial, error, and outcome-based feedback. But the interaction between those two phases, it's a complete black box. And for you listening, whether you're working in tech or you're prepping for a strategy meeting, or you just want to know how the AI on your phone is getting so smart so fast, this is the core puzzle. Yeah, it really is. How do you actually build a machine that thinks?
1:14Because pre-training corpora, those massive data sets, these models read initially, they're so unimaginably vast and messy. Highly uncontrolled. Right. Uncontrolled. So it becomes almost impossible to look at a smart AI and figure out, hey, did this specific brilliant behavior come from something it read during pre-training? Or did it learn it later during RL practice? Exactly. It's incredibly frustrating for researchers. I can imagine. If you want to run systematic tests on these massive frontier models, sweeping through different combinations of compute to see what works best, it's prohibitively expensive.
1:49Because they're just too big. Right. You can't just run 100 different versions of a trillion parameter model to see what happens. I mean, the server time alone would bankrupt a small country. Okay, so let's unpack this. Today's mission is to uncover the quantitative laws that bridge that gap between pre-training and reinforcement learning. Yes. And to do that, we aren't going to look at massive, messy language models right away. We're going to shrink things down to a highly controlled, incredibly specific laboratory. Which is so smart. We are diving into the game of chess. We're going to see exactly what happens inside an AI's brain as it shifts from just memorizing human behavior to truly learning how to reason out a problem.
2:31Right. Because to understand how a massive language model learns, we have to scale things down to an environment where every single move, you know, every single shift in mathematical probability can be perfectly measured. Which language doesn't allow. No, not at all. The problem with natural language is that it is infinitely complex and ambiguous. If I ask an AI to write a beautiful poem about the ocean, how do you verify with mathematical certainty that the poem is correct? You can't. You can't. It's totally subjective. Right. You can't write a mathematical formula for good vibes. But chess is completely different.
3:03Chess is a very compact, rigid vocabulary. Exactly. The entire game, every possible scenario on the board can be expressed with a vocabulary of just 81 tokens. Yeah, just 81. And those tokens cover the pieces, the specific coordinates on the board, and special flags like check or pawn promotion. That's amazing. And most importantly, move quality in chess can be verified exactly. By using a traditional chess engine, you can determine if a move perfectly solves a tactical puzzle or if it completely blunders the game. So a move is either correct or it isn't. Exactly. There's no debate. Okay. So in this controlled laboratory, the setup looks like this.
3:43Researchers took AI models ranging from 5 million to 1 billion parameters. Which is a huge range, yeah. Right. And they pre-trained them on 54 billion tokens of human chess games. We are talking about actual move sequences from blitz and rapid matches played by humans online. Millions and millions of games fed into the machine. And at this stage, the model is essentially doing what a language model does when it reads the Internet. It's performing next token prediction. Okay. So it sees a sequence of moves and it tries to predict the next one based on the patterns it absorbed from those 54 billion tokens.
4:20Right. It is absorbing the rhythm of human chess. Wait, but if it's just reading old games, it isn't actually reasoning. It's basically just a novice chess player sitting in a massive library reading thousands of transcripts of old Grandmaster matches. That's a great way to put it. Sure, they might memorize some popular opening moves, but if I give it a complex novel puzzle it has never seen before, it should fail, right? Mimicry isn't reasoning. It's just copying. And that is the crucial limitation of pre-training. And it's exactly why the next step in the pipeline is required. Okay. Pre-training gives the AI prior knowledge.
4:54It builds an intuition for what a normal chess game looks like, but it absolutely does not teach problem solving. To elicit actual reasoning capabilities beyond direct imitation, you need outcome-based feedback. You have to move into supervised fine-tuning, or SFT, followed by the grueling arena of reinforcement learning. But before we throw the AI into that brutal trial and error RL environment, it needs a mechanism to process its thoughts, right? It can't just blurt out the first move that comes to its digital mind. No, it needs to pause and think. Right. Historically, famously with systems like AlphaGo, the AI used an external search algorithm.
5:33Oh, right. I remember that. Yeah, it would pause, run a separate program like Monte Carlo Tree Search, which essentially simulates thousands of possible future games to see which one mathematically ends in a win most often and then pick a move. Okay, but modern large language models don't work like that. No, they don't. They generate text left to right, one token at a time. So, to make the chess model behave like a modern language model, it is trained to generate a chain of thought using synthetic reasoning traces. Instead of outsourcing the thinking to a separate simulation program, the model writes out its own internal reasoning using those 81 chess tokens.
6:08It literally maps out a decision tree thinking out loud. That is wild. It proposes a move, predicts the opponent's most likely reply, explores a few different branches of reality, and only after mapping all that out does it finally commit to an official move. Right. It is forced to deliberate in its own language of chess notation. And once it has been fine-tuned to generate these reasoning traces, it is thrown into the reinforcement learning environment. And the RL setup here is utterly unforgiving. Oh, completely brutal. The model is given complex chess puzzles to solve. It only gets a reward of one if every single executed move matches the exact ground truth solution line.
6:47Which is a very high bar. Meaning, if a puzzle requires a brilliant five-move combination, and the AI gets the first four moves perfectly right but messes up the fifth move... It gets a zero. Total failure. A single mistake anywhere in the chain yields a zero reward. Verifiable binary rewards. No partial credit. So here's where it gets really interesting for me. Does forcing the model to think out loud and map out all these branches actually make it better? Or does it just slow it down? That's the big question. Because I can easily imagine a scenario where forcing the AI to map out every single possibility just introduces way more chances for it to hallucinate and get confused.
7:26And it's a valid concern, but the data provides a very clear answer to that. Okay, what is it? If you train the model without these reasoning prices, if you just force it to spit out the final answer immediately, it actually does improve its ability to get the answer right on the first try. Oh, really? Yeah, we call that metric pass at one, your absolute best single guess. But there is a catch. A massive catch. Without reasoning traces, it completely fails to improve on a metric called pass at 16. Okay, so if you give the AI 16 different attempts to solve the puzzle, it doesn't do much better than if you only gave it one attempt.
8:02Exactly. Because without a chain of thought to guide it, its guesses lack useful diversity. Ah, I see. It's just blindly throwing darts at the exact same spot on the board over and over. But training the model with reasoning traces improves all metrics. Oh, wow. Yeah, it improves pass at 1, pass at 4, pass at 16. The model isn't just memorizing a static answer. It is fundamentally learning the mechanics of how to search a problem space. So if we know that RL makes the model actively explore and search, but it requires a solid foundation of pre-training to have good guesses to choose from, that creates a massive financial headache for an AI lab.
8:40Right, absolutely. You have a fixed budget, hundreds of millions of dollars of supercomputer time. Every dollar spent on pre-training reading is a dollar stolen from RL practice. So how do you divide the money? Well, by running 36 different combinations of pre-training and RL across those models, ranging from 5 million to 1 billion parameters, a fascinating joint scaling law emerges. Okay. First, the pre-training loss, which is basically a mathematical measure of how well the AI absorbed the initial data, strongly predicts the post-RL performance ceiling. Right. A better reader sets the stage to become a better practitioner.
9:13Yeah. Furthermore, the slope of the RL learning curve, how fast the AI improves during that trial and error phase, improves roughly linearly with the number of pre-training tokens. So more initial reading directly equals faster subsequent learning. Exactly. Let me try an analogy here. Go for it. If I'm studying for a massive, terrifying medical board exam, reading the giant textbooks is my pre-training phase. That gives me the foundation of human anatomy. Sure. Early on, if I haven't read the books, taking practice tests, which is the RL phase, is completely useless. Right. You wouldn't know anything.
9:49I'm just guessing whether the answer is lupus or a common cold. I get a zero, but I don't know why. Yeah. I need a strong pre-training base. Yes. But once I'm an advanced student, once I really know the core material, I shouldn't just keep re-reading the same textbook. I should spend a much bigger percentage of my time just grinding through complex practice cases to sharpen my clinical reasoning. If we connect this to the bigger picture, your analogy perfectly illustrates why reinforcement learning is deeply initialization dependent. You can't just skip the heavy pre-training phase, throw an untrained model into an RL puzzle environment, and expect trial and error to create genius out of thin air.
10:27Makes total sense. The model needs the prior knowledge to even have a chance at guessing the first step correctly. And that reality totally changes the math on the budget. The compute optimal frontier actually shifts based on the scale of the model you are building. It does. For smaller models, or if you have a lower total compute budget, RL is highly limited by its initialization. Like the freshman medical student taking the boards. Right. So you have to spend the vast majority of your compute budget on pre-training. But as the total compute budget grows, the optimal allocation shifts dramatically toward RL.
11:01Looking at the actual data, for a 50 million parameter model, the ideal RL compute share is only about 20%. You spend 80 % of your money on pre-training, 20 % on RL. Yes. But if you scale up to a 680 million parameter model, the ideal RL share rises to 28%. Because as the model gets larger and the budget increases, you start to hit diminishing marginal returns from just feeding it more raw textbook data. Right. Once the foundation is rock solid, the absolute best way to get smarter is to increase the proportion of time spent actively practicing and reasoning through those synthetic traces. Okay.
11:38So we know the budgeting math, and we know RL makes the model score higher overall. But for you listening, I want to take you under the hood of the neural network. Let's do it. What is actually shifting inside the model's brain at the level of individual chess moves, like when the RO reward signal dings and says, good job, you get a one, is the AI suddenly discovering new brilliance. Or is it just getting more confident in what it already knew from reading those 54 billion tokens? Exactly. This raises an important question, and it goes to the heart of what AI actually is. Does RL teach entirely new skills or just sharpen old ones?
12:15By tracking the exact probability the model assigned to every single legal chess move before and after RL, researchers can see exactly what happens. And the answer depends entirely on the difficulty of the puzzle it is facing. Okay, break down the easy puzzles first. On easy puzzles, we see a phenomenon called ground truth amplification. After pre-training and fine-tuning, the model likely already favored the correct move. Like it had a hunch. Yeah. Maybe it gave the correct move a 40 % probability and the next best move a 20 % probability. RL comes in, rewards the correct move, and just crystallizes this preference.
12:53Suddenly that 40 % shoots up to 95%. It becomes highly, highly confident in what it already knew. So it's rewarding the obvious. But what about the hard puzzles? The ones where it didn't know the answer originally, where the correct move wasn't its first instinct? That brings us to something called tail discovery. Tail discovery. Okay. On hard puzzles, the correct move was often nearly absent in the initial model's brain. It might have given the brilliant correct move less than a 5 % probability of being played. Wow. Barely a blip. Exactly. It was buried deep in the mathematical tail of the probability distribution.
13:28But through the rigorous trial and error of RL, the model actually surfaces this buried correct move. That's crazy. It's not just finding a needle in a haystack. The RL signal is dragging that needle to the very top of the pile so the AI can't ignore it. It promotes it to the top choice. That is incredible. RL really can discover hidden brilliance. It can pull out a genius maneuver that the pre-training phase barely even registered. It can, however. Uh-oh. Yeah, there is a very dark side to reinforcement learning on hard puzzles, a phenomenon known as wrong mode amplification. Oh boy, what is wrong mode amplification?
14:07On hard puzzles, while RL is desperately trying to surface that buried correct move, it also frequently amplifies highly incorrect moves. Wait, really? Yeah. It takes a wrong move that had a decent probability and makes the model incredibly confident that the wrong move is actually the brilliant solution. Wow. So RL essentially gives the AI the confidence of a guy on the Internet who skimmed one Wikipedia article and is now absolutely certain he's an expert. Exactly that. It becomes confidently incorrect. Yeah. But wait, if RL is punishing wrong moves with a zero reward, meaning the model gets zero positive feedback for a failure, why on earth is it amplifying those incorrect moves?
14:48Shouldn't a zero reward squash them down to zero probability? You'd naturally assume that, but probability mass inside a neural network is incredibly messy. How so? An AI model's policy is a probability distribution across all possible tokens, right? And that distribution must always sum to 100%. Okay, follow you. When the RL algorithm tries to boost the tiny, say, 2 % signal of a correct move buried in the tail, it fundamentally scrambles the surrounding mathematical distribution. Oh, see. It alters the weights of the neural network to make that specific rare token more likely. But because tokens share mathematical representations deep in the model's layers, boosting that one rare move inadvertently drags up competing mathematically similar, completely incorrect moves.
15:34It's like squeezing a balloon. If you pinch one tiny spot and pull it outward to make a new shape, the rest of the balloon contorts in really weird, unpredictable ways. You can't just isolate one single probability without pulling and stretching the neighbors. That balloon imagery is a perfect way to visualize it. And this dynamic perfectly explains a frustrating phenomenon seen across all of AI development. Which is? It's why RL drastically improves pass at one, your best single guess, but completely fails to consistently improve pass at 16. Because if you ask for 16 different guesses, that squeezed, scrambled balloon of probability mass means the model is now generating highly confident, but totally incorrect, variations for its backups.
16:17Exactly. It lost the healthy, open-minded diversity it had before the balloon got squeezed. It is surfacing the brilliant move for its very first guess, but for guesses 2 through 16, it is mostly churning out highly amplified garbage. That is absolutely mind-blowing. Okay, but if you're listening to this and wondering why a chess experiment matters for the chatbot you use at work every day, this is where the math jumps out of the 64 squares and into human language. Right. Chess is a perfectly closed loop. It's an 8x8 board with absolute rules. Do these exact same mechanical laws, the scaling, the balloon squeezing, the tail discovery, do they apply to the messy, infinite world of human language and logic?
16:57To answer that critical question, researchers applied the exact same framework to a completely different domain. Which was? A one billion parameter language model known as ALMO2 trained on up to 200 billion tokens of math text. 200 billion tokens of math. That is an absurd amount of data. We're talking about basically every digitized math, textbook, serum, and step-by-step calculus proof available, right? Pretty much. Math is closer to pure language, it requires long-form explanation, but it still has verifiable answers. What happened when they tested this? The exact same predictive pattern emerged.
17:32Really? The laws hold. Longer pre-trained language models reached a higher post-RL performance ceiling. And just like in the chess laboratory, their reward curves grew significantly faster during RL. The sheer amount of pre-training tokens directly dictated the slope of the RL improvement in the math domain. The more textbooks it read, the faster it learned from practice. So the laws hold up outside the 64 squares. They do. And what's fascinating here is that the structural limitations of the reasoning choices also transferred over. Okay, tell me about that. By meticulously analyzing the chain of thought generated during these math tasks, we can see that RL makes the model search much broader.
18:13Meaning it considers more candidate options or formulas before deciding what to do next. Yes. It generates a wider variety of options, and the quality of those initial options improves drastically. However, RL fundamentally struggles to increase the depth of the search. Depth. Okay. In the chess environment, the model rarely recovered continuations that required more than five sequential moves. And in the math environment, we see a similar bottleneck. The AI struggles to chain together long sequences of logical jumps. So RL is making the AI more open-minded to different immediate options, and it's picking much better short-term moves.
18:48but it's not actually helping it look further down the timeline. Exactly. RL improves candidate generation incredibly fast. It is world class at saying here are three solid ideas for what to do in the next step of this proof. Right. But deep long horizon search planning 10 or 15 complex steps ahead to prove a novel scientific theorem remains a massive bottleneck for current AI training paradigms. Interesting. The model naturally tends to broaden its immediate search rather than plunge deeply into a single extended logical chain. It's exactly like a chess player who is an absolute genius at short-term tactics, you know, spotting immediate traps and capturing pieces, but has absolutely no long-term grand strategy for how to actually win the game.
19:32And that tells the industry that just throwing more RL compute at the problem might not automatically yield AIs capable of writing a flawless 10 ,000-line software program. Or discovering a new scientific theorem. Right. If those tasks require massive uninterrupted search depth, there is a structural limit to how far ahead it can see. So bringing it all together for you listening, we've discovered some incredible mechanics today that govern the future of AI. We know that pre-training, the raw reading of data, sets the ultimate ceiling and the learning speed for an AI. You simply cannot skip it.
20:05But reinforcement learning acts as the crucial mechanism to pull hidden reasoning capabilities to the surface. And because of that dynamic, we see a clear shift in compute strategy. AI labs are moving from an 80-20 pre-training to RL split in smaller models toward a much heavier RL focus as budgets and model sizes grow. Once the foundation is there, practice makes perfect. Exactly. But we've also seen that RL is at a confident amplifier, for better or worse. It can surface very brilliant through tail discovery, dragging that needle to the top of the pile. But it can just as easily scramble the brain, amplifying wrong modes and making the AI confidently incorrect, like squeezing a balloon.
20:44So the next time you see a tech company announce a massive leap in AI reasoning, know that you are witnessing this exact delicate mathematical balancing act between textbook memorization and trial and error practice. It is a delicate balance and one that the industry is still actively fine-tuning every single day. Which leaves us with one final lingering thought to ponder. In chess and in math, we have a perfect objective verifier. A chess engine or a calculator can tell the RL system with 100 % certainty if it succeeded or failed. Right. That binary one or zero is the guardrail keeping the AI on track.
21:21But we know that RL naturally amplifies wrong modes alongside correct ones, even on hard problems where the objective truth is known. Right, the scrambled balloon. So what happens when we unleash this exact same training method on complex real-world problems? I'm talking about ethics or law or art. Areas where there is no absolute ground truth to verify the answer. Yeah. If the verifier itself is subjective or entirely flawed, what exactly is the AI amplifying? That is the real multi-billion dollar question.
From the publisher
Researchers utilized chess as a controlled testbed to investigate how pretraining choices influence the effectiveness of reinforcement learning (RL) in large language models. By systematically scaling models from 5M to 1B parameters, the study established a joint scaling law where a model's pretraining loss accurately predicts its subsequent RL performance. The findings reveal that extended pretraining not only provides a better starting point but also increases the speed at which a model improves during RL training. Mechanistic analysis showed that while RL amplifies correct moves on simple tasks, it can also surface previously hidden solutions on difficult problems. Furthermore, the authors demonstrated that these predictive patterns transfer to the math domain, suggesting the results are applicable to broader reasoning tasks. Ultimately, the study suggests that as total compute budgets grow, a larger share of resources should be allocated to the RL phase.




