DELTA: How Does RL Unlock and Transfer New Algorithms in LLMs?

28 Nov 2025 · 11 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Whether RL fine-tuning makes LLMs learn genuinely new algorithms or only reveal latent skills; the episode centers on the “Delta” benchmark (Distributional Evaluation of Learnability and Transferability in Algorithms) using learnability and generalization criteria.

Guest backgrounds

No guests are mentioned in the transcript.

Key claims

With the right staged reward, RL can move a model from guaranteed failure to near-perfect success on out-of-distribution algorithmic tasks; generalization improves across domains but breaks at “transformative” schema invention.

Notable examples

Baseline Qwen3 4B Instruct fails “Manufactoria” path-at-K0 tasks (0% pass on sequences like GGRBV). RL suffers a “zero gradient” when binary rewards stay at 0; partial credit per test case helps briefly but gets gamed. A two-phase dense warm-up then strict binary convergence produces a grokking-style jump (below 1% to ~100% in a few steps). Transfer tested in “Bouncing Sim”: exploratory and compositional generalization reach ~60–70% pass rates; transformative generalization (e.g., discovering new invariants) drops back near zero.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Core Debate

0:45 to 1:40

Exploration of whether reinforcement learning teaches new skills or refines existing ones.

“They came up with two really clear criteria to settle it, learnability and generalization.”

Introducing Delta: A New Benchmark

1:40 to 2:36

Overview of the Delta benchmark and its criteria for testing RL in LLMs.

“Now, to prove it's learning something new, you can't test it on things it might have seen online.”

The Zero Gradient Problem

2:36 to 3:30

Discussion on the challenges faced with standard reinforcement learning methods.

“And if it never passes, the signal is always 0, always fail.”

Implementing Partial Credit

3:30 to 4:36

How researchers utilized partial credit to overcome initial learning challenges.

“This is where the idea of partial credit comes in.”

The Staged Recipe Approach

4:36 to 5:36

Details on the two-phase process for training the model effectively.

“That's a great question, and it gets to the core of the tradeoff.”

The Grokking Transition Phenomenon

5:36 to 6:36

Description of the sudden improvement in model performance after plateauing.

“Grokking is just this bizarre dynamic where the model hits a long, long plateau.”

Generalization and Its Tests

6:36 to 8:46

Exploration of how the model's ability to generalize is assessed across different domains.

“It's really strong evidence that RL, with the right reward, can be a genuine discovery engine.”

Limits of Procedural Learning

8:46 to 9:37

Discussion on the boundaries of the model's learning capabilities and its limitations.

“Transformative generalization requires the model to invent a whole new solution schema, something qualitatively different.”

Implications for Future AI Research

9:37 to 10:39

Concluding thoughts on the significance of reinforcement learning for advancing AI.

“So what does this all mean for you, listening to this?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. Today we're digging into one of the most fundamental questions in AI research right now. When you take a big pre-trained language model and you fine tune it with reinforcement learning with RL, are you actually teaching it something new? I mean, is it discovering a genuinely new procedure or is the RL just, you know, polishing skills that were already hidden in there from its training? That's the core of the debate, really. And the field has been split. On one side, you have people arguing that RL just refines latent heuristics, which is a fancy way of saying it's just bringing out abilities the model kind of already had.

0:36The other more optimistic view is that it can acquire genuinely new procedures, things the base model absolutely could not do. Right. And the paper we're looking at today finally moves this out of just like a philosophical argument and into something you can actually test. They came up with two really clear criteria to settle it, learnability and generalization. Exactly. The work is all built around this new benchmark called Delta. That's the Distributional Evaluation of Learnability and Transferability in Algorithms. And their mission was, well, it was pretty brutal. Can RL teach a model to do something that it fails at 100 % of the time?

1:11What they call a path at K0 task. So in plain English, you're looking for proof that the model went from total guaranteed failure to actually being good at something. To prove the skill wasn't just lying dormant. Precisely. For their baseline, they used Quen 3 4B Instruct. It's a solid, mid-sized, open-source model. Competent, but nothing special. And crucially, on the hard problems they designed for this, it had zero success. Confirmed pass at K0. Okay, so that's the baseline. Now, to prove it's learning something new, you can't test it on things it might have seen online. You need a sandbox that's genuinely out of distribution.

1:46Where'd they take this model to school? They took it to a place called Manufactoria. It's not a real place, obviously. It's a synthetic programming environment based on this puzzle-like logic. Think of it like building little machines, finite state automata using a totally custom set of commands. Things like puller and painter nodes. Because the syntax is completely made up, it's truly new territory for the LLM. And they picked tasks that were designed to stump it, right? Yeah. Like this manufactory ages family. Exactly. Those tasks require the model to write a program that checks for a very specific sequence of characters, something like GGRBV, again, the base model, QAN3.

2:24It got a full pass rate of exactly 0%, total unambiguous failure. Which brings us to the first big problem. The zero gradient problem. If you're using standard RL and the model just fails every single time, your reward signal is just pass or fail. A 1 or a 0. And if it never passes, the signal is always 0, always fail. Mathematically, what that means is the model gets no positive signal. It just collapses. There's no gradient telling it how to get better. Let's use an analogy for the listener here. Imagine you're lost in a totally dark valley, and you're trying to find the highest point. And your only feedback is someone yelling, nope, every time you take a step.

3:02You're not at the summit yet. You get a zero. That zero tells you nothing. Is north a little bit uphill? Is east? You have no signal to adjust your path. That's a perfect analogy. And you see it in the training graphs. The full pass rate just flatlines at zero for hundreds of steps. The model is just stuck at the bottom of that valley. Okay, so if that sparse binary reward is a total failure, how do the researchers give the model a hint, a little push to get out of the valley? This is where the idea of partial credit comes in. Precisely. They started by leveraging the structure of programming itself.

3:37I mean, even if your program fails overall, it might pass, say, 5 out of 20 test cases. So they tried using that per test pass rate as a continuous reward, a number between 0 and 1, and that gave it some initial traction. But that wasn't enough on its own, was it? No, and that's what's so fascinating. It helped the model get some gradient, enough to start climbing. But the model quickly learned to game it. After about 100 steps, the reward saturated. It got really good at passing the easy test, but the full pass rate. It was still basically zero. It learned to settle for partial credit. So partial credit gets you part of the way up the hill, but to actually reach the summit, you need to find that perfect single path.

4:16And this leads to the big breakthrough, right? The staged recipe. This is it. It's a two-phase process, a dense warm-up phase, followed by a super strict convergence phase. Walk me through the mechanics of that. Why not just keep using the partial reward? Isn't there a risk that switching to that strict all-or-nothing reward just destroys all the progress you just made? That's a great question, and it gets to the core of the tradeoff. The warm-up phase with that dense per-test reward, that's for exploration. It's just about pushing the model out of that all-zero region and getting it to explore the solution space.

4:49It learns the general shape of a good solution. Okay, so it gets it out of the valley. Right, but once it's out, that dense reward becomes a trap. It encourages approximations, not perfection. So to force the model to find the exact syntactically perfect solution that passes all the tests, you have to switch gears. And that's the convergence phase. You flip the switch back to that unforgiving binary full pass reward. Exactly. This forces exploitation. It tells the model, OK, no more partial credit. Now you consolidate what you've learned and find the single correct procedure. It stops it from being lazy.

5:25And that's the key to making it learn the actual new skill. And the result of this whole process was, well, it was that grokking transition phenomenon we've seen before, but on a whole new level. Yes. Grokking is just this bizarre dynamic where the model hits a long, long plateau. It looks like it's making no progress at all. And then suddenly there's this abrupt vertical jump to near perfect accuracy. It's not a smooth curve. And the graph for the manufactoria task is just, it's incredible. After the warmup, it hits this plateau for about 450 steps. For 450 training steps, the full pass rate is stuck below 1%.

6:00It looks like a complete failure. It looks like pure stagnation. And you can imagine training a system for days and just seeing nothing. But underneath the surface during that plateau, the model is silently reorganizing its internal logic. It's building the concepts it needs. And then all of a sudden, it just clicks. The pass rate shoots from less than 1 % to nearly 100 % in just a few steps. That's the moment. That sudden discovery is the proof. It's the acquisition of a procedure the base model could not execute. They confirmed a nearly 100 % absolute improvement over the reference model. They went from 0 % solvable to 100 % solvable.

6:39It's really strong evidence that RL, with the right reward, can be a genuine discovery engine. So we've established it. RL isn't just a tuner. You give it the right map out of that zero reward trap, and it can be a procedural engine. But now that it's learned this new trick, how well does it generalize? That's the other half of the Delta framework. Right. Generalization is the real test. Did it just memorize a solution or did it learn the abstract rule? And for this, they often switch to a totally different domain. Bouncing Sim. It's a physics simulation family. Wait, they switched from coding puzzles to physics.

7:10Why? Why was Bouncing Sim a better test for generalization? Because Manufactoria was about learning if a new procedure was possible. Bouncing Sim is perfect for testing the transfer of that skill. It lets them cleanly test three different kinds of generalization. First was exploratory generalization. This is just making the same task harder. Smaller containers, faster balls, more collisions. And how did it do? It did well. The skill transferred. The model had learned the underlying rules of the simulation, not just specific examples. Okay, now for the one that really matters for the future of AI.

7:46Compositional generalization. This is where systems usually fall apart, right? trying to combine skills they learned separately. And this is where they saw really strong success. They'd train a model on, say, how to simulate rotating boxes and separately how to simulate moving objects. Then they'd test it on an unseen combination, like a rotating box that is also moving across the screen. So it had never seen that specific scenario before. What happened? It showed powerful compositional ability. It was getting pass rates in the 60 % to 70 % range on these complex unseen tasks. The conclusion was that these kinds of coding and simulation tasks are really amenable to structural composition.

8:25The model basically learned reusable internal modules for rotation and movement and could just stitch them together on the fly. That's a huge deal. That's way beyond just matching patterns, but there's always a limit. Where did this procedural learning finally break down? It broke down at the third axis, transformative generalization. This is the hard boundary. Transformative generalization requires the model to invent a whole new solution schema, something qualitatively different. Can you give an example? In bouncing sim, it might be a task that requires discovering a new physical invariant, like figuring out the math for a perfectly periodic trajectory, which needs a different kind of reasoning than just collision physics.

9:06On those problems, performance just cratered, back down near zero. Which makes sense. It aligns with this idea of schema creation. The model is great at composing the rules it knows, but it can't invent totally new rules from scratch. It can build a house with bricks, but it can't invent concrete. Precisely. RL unlocked the ability to discover and generalize procedures that were impossible for it before. But that leap to truly novel paradigm-shifting invention, that's still the boundary. So what does this all mean for you, listening to this? The big takeaway is that RL is so much more than a simple tuner.

9:43If you give it the right stage signal, that dense reward to get it started, then the binary one to force perfection, you can push LLMs past their limits. You can make them learn genuinely new things. The lesson is all about how you structure feedback when you're starting from zero. You can't just use pass-fail. You have to use some kind of intermediate dense structure, the partial credit, to give it the gradient it needs to climb out of that zero reward valley. Then you get strict. That principle is transferable to almost any domain. And that leads to our final thought for you to ponder. This worked amazingly well here, in the clean worlds of code and synthetic physics.

10:16But can you find analogous signals in messier domains? Think about things like stepwise math checkers, or rubric scoring for complex writing, or even theorem provers in formal logic. Could applying the Delta Framework's insights there unlock currently unsolved problems in science and math? We've seen the machine jump from 0 to 100. Finding the right feedback might just be the key to the next great breakthrough.

From the publisher

This paper introduces DELTA, a controlled benchmark of synthetic programming tasks—such as Manufactoria puzzles and BouncingSim physics simulations—specifically designed to isolate and evaluate whether reinforcement learning (RL) can teach large language models (LLMs) genuinely new reasoning procedures. The study demonstrates that RL can achieve **learnability beyond pretraining** on tasks where reference models previously failed completely, noting that naive binary reward training fails. This success is enabled by a **two-stage training strategy** that begins with dense, per-test case rewards for warm-up before switching to strict binary rewards, which triggers an abrupt **grokking transition** from exploration to mastery. Furthermore, the analysis of transferability shows that these learned skills generalize robustly across exploratory and **compose effectively** across combined skills, though performance remains poor under **transformative shifts** requiring qualitatively novel solution schemas.

More from Best AI papers explained

All 475 episodes
DELTA: How Does RL Unlock and Transfer New Algorithms in LLMs?Best AI papers explained · 11 min
Listen in VO