In short
Compositional generalization in language-model reasoning—why supervised fine-tuning (SFT) can make models rigid, and how reinforcement learning (RL) “untangles” reasoning into reusable skill and routing modules.
Guest backgrounds
No guests are identified in the transcript; it’s a two-host discussion (“Welcome to the Deep Dive…”) with no named experts.
Key claims
Perfect step-by-step training traces can cause “SFT hidden support,” entangling specific skills with specific routing choices. RL trial-and-error decomposes these fused templates, improving generalization to longer, unseen compositions.
Notable examples
Synthetic string-transformation tasks (DEP3 trained with SFT; tested up to 8 steps) where RL improves accuracy by +40.3%. RL on compound traces beats RL on isolated one-step skills by 37.8%. Real-world analysis on “Quinn III” models: RL-trained models show more diverse short skill sequences (“engram” diversity) than SFT-only models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Quest for Genuine Reasoning
0:34 to 1:24
Discussion on the transition from memorization to genuine reasoning in AI.
“Today, we are looking at why perfection makes AI stupid and why that chaotic plunder is literally the only way to build the tools you are probably using every single day.”
Understanding Hierarchical Latent Selection
1:24 to 2:28
Breakdown of the hierarchical latent selection model and its components.
“Okay, but before we can even understand how an AI learns to reason, we have to figure out what a thought actually looks like to a machine.”
Skills vs. Routing in AI
2:28 to 4:00
Explanation of the roles of skills and routing mechanisms in AI reasoning.
“So if the AI is looking at a sentence, a skill might be applying a grammatical rewrite rule.”
Supervised Fine-Tuning Explained
4:00 to 6:00
Introduction to supervised fine-tuning and its impact on AI learning.
“Let's look at a concrete algebra example to see how the AI actually snaps these pieces together.”
The Flaw of Perfect Examples
6:00 to 7:12
Analysis of how perfect training examples can hinder AI flexibility.
“It never realizes that the skill and the router are separate, reusable pieces.”
Reinforcement Learning as a Solution
7:12 to 8:38
Discussion on how reinforcement learning helps to untangle AI reasoning.
“Going back to the kitchen, like if I only ever learned to chop an onion while making spaghetti sauce.”
Experimental Evidence of Learning
8:38 to 11:15
Exploration of experiments demonstrating the effectiveness of RL over SFT.
“How do researchers actually prove the AI is untangling these concepts under the hood?”
Combining SFT and RL Strategies
11:15 to 14:01
Insights into how to optimally combine SFT and RL for best results.
“That brings us to a huge practical dilemma for the people actually building these models.”
Exploring Decomposition Theory in AI Models
14:01 to 16:48
Learn how decomposition theory holds up in state-of-the-art language models through empirical data and analysis of problem-solving skills.
“But I want to take this out of the synthetic lab.”
Understanding Compositional Generalization
16:48 to 17:08
Discover the paradigm shift in AI reasoning as models dynamically combine skills for complex problem-solving.
“Taking a step back, this completely changes how I view these systems.”
Show all 11 chapters
The Future of AI Self-Exploration
17:08 to 18:18
Examine the potential future where AI autonomously creates its own learning paths by identifying and exploring unknown skills.
“The machine isn't reciting a magic incantation.”
Transcript
Automatic transcript. May contain errors.0:00If you take the most advanced reasoning AI in the world, right? and you feed it the absolute perfect step-by-step answer to a complex math problem 10 ,000 times, you will actually break its brain. It will become almost completely unable to solve anything new. Yeah, it is a massive paradox. You know, you give a machine a flawless blueprint and it gets incredibly rigid. But if you throw that same AI into a, well, a chaotic trial and error blender where it has to stumble, guess, fail. Suddenly learns how to truly think. Exactly. It's wild. Welcome to the Deep Dive. Today, we are looking at why perfection makes AI stupid and why that chaotic plunder is literally the only way to build the tools you are probably using every single day.
0:43We are talking directly to you, the listener. Maybe you're prepping for a tech meeting or, you know, maybe you were just insanely curious about how these massive large language models actually go from reciting memorized trivia to solving multi-step coding problems they have literally never seen before. Because that jump, right, from memorization to genuine reasoning, that is the holy grail of artificial intelligence. We are dissecting the one-two punch that makes this possible. So that's supervised fine-tuning or SFT and reinforcement learning, RL. Right, the big buzzwords. Yeah. But our goal today is to actually uncover the hidden mechanisms underneath them, specifically a concept called compositional generalization.
1:23That is the engine that allows an AI to break down rigid, memorized answers into flexible, reusable brainpower. Okay, but before we can even understand how an AI learns to reason, we have to figure out what a thought actually looks like to a machine. Because, I mean, it's obviously not a little voice in its head. No, definitely not. It operates on a fundamentally different architecture. To grasp this, researchers use a theoretical framework known as the hierarchical latent selection model. Okay, hierarchical latent selection. That is a mouthful. Let's break that down, starting with latent. Yeah.
1:55Because usually that implies something is hidden underneath the surface, right? That is the perfect place to start. It is totally hidden from us, the users. When you ask an AI a question, you see the text screaming across your screen. But before it types a single word, the AI is making a cascade of invisible, discrete choices. Like rapid-fire decisions in the background. Right. We call these choices atomic modules. And the hierarchical part of the name means there is a strict chain of command for these modules. Basically, there is a boss and there is a worker. So we are talking about two different categories of building blocks for every single thought.
2:30Yeah. The first category is the worker. We call these skills. A skill is a local operation. It's the actual execution of a task. So if the AI is looking at a sentence, a skill might be applying a grammatical rewrite rule. Okay, and in math. In math, a skill is performing a specific arithmetic step, like adding two numbers or isolating a variable in an equation. It's the doing. And if the skill is the worker doing the heavy lifting, what is the boss doing? The boss handles the routing mechanisms. If skills are the labor, routing is the logistics. The logistics, meaning what exactly? Meaning routing mechanisms decide what intermediate information to use next.
3:10They look at the board and determine, you know, do we carry this previous result forward? Do we pull in a new piece of information or do we branch off and try a completely different path based on that last calculation? OK, let's unpack this. Let me see if I can map this to the physical world. Let's try cooking. The skills, the workers are the actual physical actions. Chopping an onion, whisking an egg, searing a piece of meat. Sure. But if you only have skills, you just end up with a pile of chopped raw onions on the counter. Right. You need a recipe to connect those isolated actions. Exactly.
3:45So the routing is the recipe. The routing is the logic that says, look at the state of the pan. The oil is hot. Take the output of the chopping step and route it into the pan for the sauteing step. That's a great analogy. The routing connects the outputs of previous skills to the inputs of the next ones. Let's look at a concrete algebra example to see how the AI actually snaps these pieces together. Boy, you're on me. Imagine you give the AI this word problem. The sum of one number and twice another is 11. The first is 2 greater than the second. Oh man, classic high school math trauma right there.
4:17I know, I know. But underneath the surface, the latent model kicks in. First, it selects a routing mechanism, which we can label translate to a system. This router takes the natural language and maps it into two mathematical equations. X plus 2Y equals 11 and X equals Y plus 2. Okay, so it's set up the board. Exactly. Now the router has set the stage and it needs a worker. It calls up a skill. In this case, the substitute skill. So the substitute skill takes the value of X from the second equation and physically plugs it into the first. Yep. The entire thought process is just a long, literal sequence of snapping together these invisible routing and skill modules.
4:56I am stuck on something here, though. If an AI's reasoning is just made up of these atomic Lego blocks the skills and the routing, who is handing the AI the instruction manual? Because we need to talk about supervised fine-tuning. You mentioned at the start that giving the AI perfect blueprints actually breaks its brain. It does. So supervised fine-tuning, or SFT, is the first phase of training. We feed the model thousands and thousands of canonical golden traces. Golden traces meaning, like, perfect answers. Flawless, step-by-step demonstrations of how to solve a problem. The SFT phase supplies all the raw atomic materials.
5:32The model sees every possible skill and every possible routing mechanism in action. Sounds ideal. It does, but this method has a fatal flaw. It leaves those modules statistically entangled. Entangled, like they get tied in a knot. More like they get inextricably glued together. In these golden training examples, a specific skill and a specific router almost always co-occur in perfect harmony. Because they are always presented together, the AI just memorizes the fused block. Oh, I see. Yeah. It never realizes that the skill and the router are separate, reusable pieces. It learns the entire sequence as one giant, inflexible chunk.
6:08I have to push back on this, though. If SFT is giving the model the absolute perfect, golden answer, Shouldn't a perfect example create a perfect student? I mean, if I practice a piano scale perfectly 10 ,000 times, I am going to be incredible at that scale. Why does this make the AI fail? Because of a phenomenon called SFT hidden support. The SFT model only ever sees the visible subset of perfectly arranged traces. It develops a massive blind spot for variation. A blind spot. Yeah, let's use your piano analogy. You learn the sheet music perfectly. You can play that one song flawlessly. But if someone asks you to play jazz, you know, to improvise...
6:47It completely frees. Exactly. You freeze because you have never seen what happens when the notes are played out of order. Okay, I think I see it. If the AI only ever sees the substitute math skill used immediately after the translate router, it assumes those two things are permanently fused together. It doesn't know it can use the substitute skill in, say, geometry or calculus. The AI completely lacks the evidence to prove to itself that the modules can exist independently. It just relies on the template. Right. Going back to the kitchen, like if I only ever learned to chop an onion while making spaghetti sauce.
7:19If you asked me to make salsa, I'd say I can't chop an onion. There's no boiling water for pasta. I've completely entangled the onion chopping with the spaghetti making. And because SFT leaves these modules tangled up in rigid templates like that, we need a mechanism to forcefully break them apart so they can be recombined. Enter reinforcement learning. The great decomposer. The great decomposer. I like that. So how does RL actually break these templates apart? RL fundamentally reshapes the AI's reasoning through massive scale trial and error. Instead of showing the AI the perfect answer, RL gives the AI a problem and forces it to try multiple different rollouts or trajectories.
7:58It only gets a reward if the final answer at the very end is correct. So it's wandering through a maze blindfolded. Hoping it bumps into the cheese, exactly. But the way it learns from bumping into the walls is what matters. By forcing the model to try endless variations, RL exposes it to entirely new contexts. And how does that help untangle the pieces? The final reward signal acts like a scalpel. It isolates specific local events, for example. It notices when a specific skill happened to work in a brand new context. Over millions of iterations, RL decomposes those rigid SFT templates into truly independent, reusable atomic modules.
8:37Okay, but we must have hard data showing this happening, right? How do researchers actually prove the AI is untangling these concepts under the hood? They ran a highly controlled synthetic experiment using string transformations. Think of it as a complex, multi-step code-breaking task. You take a sequence of letters and apply a series of convoluted rules to transform them. Sounds like a logic puzzle. Yeah, exactly. So they trained one set of models strictly using SFT on compound traces of DEP3. Def 3 meaning the problems required exactly three steps to solve. Yes. They handed the SFT model perfect three-step blueprints.
9:11Then they evaluated these models on unseen, longer compositions. They gave them problems requiring up to eight steps to solve. Oh, testing if they can actually play jazz and improvise beyond the three-step sheet music? Precisely. And the SFT-only model absolutely crashed. It fell apart because it had just memorized the three-step templates. The moment a fourth step was required, it didn't know how to route the information. Wow. But the model trained with reinforcement learning maintained incredibly high accuracy. It showed an average gain of plus 40.3 % over the SFT model on those unseen, deeper problems.
9:48Just massive. Because the RL model was forced to stumble around and try different things to get the reward, it accidentally figured out how the underlying pieces worked. Let me try to visualize this. Here's where it gets really interesting. SFT is like taking a kid and handing them a fully glued together Lego castle, right? Right. They know exactly what a castle looks like, but they can't build anything else. They only know the final shape. But RL is forcing them to smash the castle against the floor, look at the individual bricks scattered everywhere, and suddenly the kid realizes, oh, this blue block and this red block, I can snap these together to build a spaceship.
10:21It's the breaking apart that creates the flexibility. The data strongly supports that visualization. Researchers actually tried an alternate experiment to test that exact idea. Oh, really? Yeah. They asked, what if we skip the castle entirely? What if we just hand the AI isolated Lego bricks and teach it the individual skills one by one? Like giving it isolated math problems that only require one single step. Right. And they found that running RL training on compound traces, letting the AI play with the full multi-step structures yields a massive 37.8 % performance gain compared to running RL only on isolated single-step atomic modules.
11:00Wow. So you can't just hand the AI a single Lego brick and say, learn this. No, the AI needs to see the bricks embedded in a structure to learn how they connect. It has to learn the interfaces. If you just give it isolated skills, it never learns the routing mechanisms required to spring them together. That brings us to a huge practical dilemma for the people actually building these models. If SFT provides the glued-together castle and RL teaches the AI how to pull the briffs apart and build spaceships, how should engineers actually feed the training data into these two phases? There are two massive, totally counterintuitive findings regarding how to pair this data.
11:35First, it's about what researchers call combinational exposure. Let's say during the RL phase, you remove a specific atomic skill entirely. You hide the chopping skill. You cannot fix the model later by just injecting that skill back in isolation. Meaning showing it an isolated chopped onion doesn't help it learn to cook. Exactly. If you try to train it on depth one tasks of that missing skill, the model fails. The skill must be re-injected inside a compositional trace. Atoms are useless without context. The model needs to see how the chopping skill connects to the heat of the pan and the oil.
12:08Okay, that makes sense based on the Lego bricks needing to be in a structure. What is the second finding? The second finding is about disjoint data, and it completely upends how we think about the relationship between SFT data and RL data. It turns out the absolute worst reasoning performance happens when your SFT data set completely overlaps with your RL data set. Hold on, wait, wait, wait. You are saying that to make the AI better at combining things, we literally want the second phase of training the RL phase to use completely different combinations than the first phase. Yes. That violates basically every rule of human learning.
12:43Usually, repetition and consistency are the keys to mastering. I know it sounds totally contradictory, but look at the division of labor between the two algorithms. SFT's job is just to make sure all the raw atomic modules exist somewhere in the AI's brain. It is just stocking the pantry with ingredients. RL's job is exploration. It is the chef trying to invent a new menu. Oh, I see. If the chef just cooks the exact same recipes that are already in the pantry's instruction manual. The AI doesn't learn any new interfaces. It just reinforces those same glued together castles. By intentionally forcing the RL phase to explore novel compositions, combinations that were entirely outside the SFT support data, you force the AI to create new local witnesses.
13:26Local witnesses. Let's unpack that term. What is a local witness? A local witness is created when the AI successfully proves to itself that a specific skill and a specific router can plug into each other safely in a brand new environment. Without disjoint data, the AI never has to seek out new local witnesses. It just gets lazy and relies on memory. Ensuring there is zero overlap between the SFT combinations and the RL combinations forces the model to actually understand the underlying logic. So you essentially have to pull the rug out from under the AI during phase two to force it to actually think.
14:01Basically, yeah. But I want to take this out of the synthetic lab. It's one thing to prove this with fake code-breaking string experiments. Does this decomposition theory hold up in state-of-the-art, open-source models out in the wild? It absolutely does. Empirical data collected from looking under the hood of the Quinn III family of models provides real-world proof. Researchers gathered 643 highly advanced math problems. We're talking grueling, competition-level math from data sets like A and Math 500. So that breaks humans. Yeah, exactly. They fed these problems into two different models to compare them.
14:35One SFT-only model and one RL-trained model. Yep. They used QEN3-4B, the base SFT model, and QEN3-4B thinking, which had survived the RL crucible. But they didn't just check the final answers. They dug into the actual methodology of the text generation. They built an atomic skill library, and then they algorithmically mapped every single sentence the AI generated back to those specific skills. Wait, how do you map a paragraph of text back to invisible skills? They looked for markers in the language that indicated a specific mathematical operation was happening. Once they tagged all the text with these skill labels, they analyzed short engrams.
15:13Engrams. Yeah, an enneagram in this context is a sequence of two or three atomic skills used in a row. For example, skill A followed by skill B followed by skill C. They wanted to see if the models were using the same enneagram chunks over and over or if they were mixing it up. Ah, looking to see who was playing the sheet music and who was playing jazz. What did the Enneagram data show? The RL model showed vastly more diversity in these short combinations. The SFT model, on the other hand, heavily favored long, repetitive sequences. It was leaning entirely on rigid trace templates it had memorized during training.
15:47Unbelievable. But the RL model was combining and recombining short skill sequences in a huge variety of ways, dynamically adapting to whatever the specific math problem demanded. So what does this all mean? It's like the S.F.T. model was a politician reciting a long, rigid, pre-written monologue. Oh, that's a perfect way to put it. No matter what question the debate moderator asks, the politician just falls back on the same five-minute stump speech because that's the template they memorized. They don't actually answer the prompt. Right. But the R.L. model is having a dynamic, fluid conversation.
16:21It is pulling short phrases, adapting its tone, and snapping ideas together based on the actual person standing in front of them. It proves that the RL model isn't just generating longer text to create the illusion of thinking. It has genuinely abstracted those step-level skills, and it has the confidence to plug them into novel neighboring operations across entirely different, unseen problems. It has achieved real compositional generalization. Taking a step back, this completely changes how I view these systems. For you, the listener, the next time you boot up a tool like ChatGPT or Gemini, and you watch it successfully reason through some bizarre multi-step logic puzzle that you just invented off the top of your head, you now know the secret mechanism at play.
17:07It really is a paradigm shift. Yeah. The machine isn't reciting a magic incantation. Under the hood, it is dynamically snapping together atomic skills and routing mechanisms in real time. It is utilizing pieces that it only learned to detangle from one another through the intense exploratory trials of reinforcement learning. And understanding that mechanism raises an incredibly thrilling question for the future of this technology. Because right now, human engineers are the ones carefully curating all of that training data. We decide what goes into the SFT phase, and we curate the novel disjoint combinations for the RL phase to explore.
17:42But we now know that RL thrives on exploring unseen combinations. So what happens when we remove the humans from the curation process? Imagine an AI system designed to actively track its own local interfaces. The AI continuously monitors itself, logging which atomic skills it has never tried plugging together before. It automatically identifies its own blind spots and then dynamically generates its own custom RL curriculum to explore those exact unknown interfaces. We could be looking at a future where an AI autonomously builds its own obstacle courses to guide its own path to super reasoning.
18:17An AI that knows what it doesn't know and builds the maze required to teach itself? That is equal parts terrifying and completely fascinating. Thank you so much for joining us on this deep dive today. Keep questioning the hidden mechanics of the technology around you. Because as we learned today, perfection is a trap. And sometimes you have to smash the Lego castle to learn how to build a spaceship. Catch you next time.
From the publisher
This paper studies how post-training pipelines transform large language models into effective reasoners through compositional generalization. The authors propose a hierarchical latent selection model that separates reasoning into atomic skills, such as local operations, and routing mechanisms that dictate how information is composed. Their theory suggests that supervised fine-tuning (SFT) provides the necessary raw materials, while reinforcement learning (RL) identifies and decomposes these elements into reusable modules. Controlled experiments validate that RL enables models to solve novel tasks by recombining learned atoms in ways not seen during training. Ultimately, the study concludes that SFT should focus on broad module coverage while RL should target genuinely new compositions to maximize out-of-distribution performance.




