LLMs Can Learn to Reason Via Off-Policy RL

19 Mar 2026 · 20 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that current LLM reasoning training wastes massive compute because trainer and inference engines fall out of sync, making on-policy RL updates (e.g., PPO/GRPO) mathematically invalid. It presents OAPL (Optimal Advantage-Based Policy Optimization with Lagged Inference Policy) as an off-policy method that embraces lag instead of forcing synchronization.

Guest backgrounds

No guest names or bios are provided in the transcript.

Key claims

Off-policy lag breaks PPO/GRPO assumptions; industry uses importance sampling, clipping, slowing inference, or discarding sequences to cope, causing variance, instability, and wasted compute. OAPL uses KL regularization (“bungee cord”) plus a squared regression objective to grade lagged trajectories without importance-sampling ratios, preventing entropy collapse and enabling stable learning even with extreme lag.

Notable examples

Benchmarks on AME25, HMMT25, Brumo25; coding comparison to DeepCoder. OAPL matches/outperforms while using ~200k vs 650k samples and scales pass@K strongly under up to 400 gradient-update lag.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Structural Bottleneck in LLMs

0:45 to 1:49

Understanding the forced synchronization issue in traditional AI learning.

“Because for a long time, the whole industry has operated under this very rigid mathematical assumption about how an AI has to learn.”

Real-World Analogy: The Driving Instructor

1:49 to 3:45

A driving instructor analogy illustrates the lag problem in AI training.

“I mean, reinforcement learning, or RL, is basically how we teach an AI to make good decisions, right?”

Current Workarounds: Importance Sampling

3:45 to 5:43

Examining importance sampling as a temporary solution to training delays.

“To put this in a real-world context for you listening, imagine a driving instructor and a student driver.”

The Cost of Inefficiency in AI Training

5:43 to 8:05

Exploring the vast inefficiencies caused by discarded data in AI training.

“Yeah, I mean, the primary workaround the industry uses is a technique called importance sampling.”

Introducing OAPL: A Paradigm Shift

8:05 to 8:56

OAPL embraces off-policy learning to improve LLM training.

“So if throwing away this expensive data is bankrupting labs, someone had to realize that forcing this perfect synchronization is a losing battle.”

How OAPL Handles Stale Data

8:56 to 11:05

OAPL's unique approach to processing delayed data without crashing the model.

“But how do you train effectively if the data is stale?”

Preventing Entropy Collapse with OAPL

11:05 to 13:19

How OAPL maintains diverse problem-solving approaches in AI models.

“Isn't the student literally just practicing bad habits?”

Benchmarking OAPL Against Traditional Models

13:19 to 14:05

Comparing OAPL's efficiency and performance in code generation tasks.

“But, you know, math is really only half the story for these frontier models.”

Off-Policy Learning and Its Efficiency

14:05 to 16:42

Discover the advantages of off-policy learning in AI, highlighting reduced compute time and increased learning efficiency.

“OAPL achieved the exact same reasoning power using just 200 ,000 samples.”

Shifting the AI Infrastructure Paradigm

16:42 to 17:18

Learn how OAPL can democratize AI development by lowering costs and removing the need for high-speed infrastructure.

“Right now, building a frontier reasoning model is a game mostly reserved for massive tech giants with limitless budgets who can afford perfectly synced high-speed infrastructure.”
Show all 13 chapters

Evolving AI Learning Strategies

17:18 to 18:06

Explore how OAPL challenges traditional reinforcement learning methods and enhances AI's ability to learn from asynchronous data.

“It could democratize the creation of hyper-smart AI, allowing smaller research labs, open-source communities, or startups to compete without needing a billion-dollar supercomputer perfectly synced to the millisecond.”

The Insight on AI Learning

18:06 to 18:34

Understand the transformative insight that AI learns better when not micromanaged, allowing for diverse experiences.

“to a simulator-trained expert exploring every possible route, saving millions of dollars in compute along the way.”

Future Implications of OAPL

18:34 to 19:28

Contemplate the future of AI learning from historical data without real-time constraints, reshaping interactions with the world.

“But, you know, this raises an important question, a truly provocative thought about the future.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Right now, there are massive AI companies basically throwing millions of dollars of perfectly good logical reasoning into the digital trash can. Yeah, it's actually painful to watch if you know how expensive that compute is. Right. And why are they doing it? Well, because the brain generating the answer was literally just a few seconds out of sync with the brain grading the test. Exactly. It sounds completely absurd. But, well, this is the hidden mechanical bottleneck in how large language models, or LLMs, are trained to actually think. So today we're doing a deep dive into this structural flaw in current AI training methods.

0:37And more importantly, we're looking at a radical new algorithm called OAPL that, frankly, completely throws out the established rulebook. It really does. Because for a long time, the whole industry has operated under this very rigid mathematical assumption about how an AI has to learn. And the mission for our deep drag today is to explore exactly why that forced synchronization is failing. Yeah. And how OAPL solves this by just, you know, embracing the very lag that engineers have spent years desperately trying to fix. So whether you're a machine learning engineer or just someone insanely curious about where AI is heading, you are going to walk away with a really clear understanding of the mechanics driving the next generation of reasoning models.

1:18It's going to be a fun one. But you might be listening to this and thinking, wait, didn't models like DeepSeek R1 already solve this whole AI efficiency thing? I mean, we just saw these massive breakthroughs doing incredible things with way less compute. Oh, they absolutely pushed the boundaries for sure. But even DeepSeq was forced to use what we call on-policy reinforcement learning, specifically an algorithm called GRPO. Right. So they were still basically bumping their heads against this exact same structural ceiling. Exactly. I mean, reinforcement learning, or RL, is basically how we teach an AI to make good decisions, right?

1:55You give it a complex math problem, it tries to solve it step by step, and you give it a reward if it gets the final answer right. Simple enough. Right. But the absolute core theory of these on-policy algorithms like PPO or GRPO is that the data used to update the AI's neural network, it must be generated by that exact same perfectly up-to-date network. OK, let's unpack this. Meaning, if the AI figures out a brilliant new way to solve a calculus problem, you have to reward the exact version of the AI that just solved it, not like a slightly older version from five minutes ago. Precisely. The policy, which is just the AI's current state of knowledge that generates the action, has to be the exact same policy that gets updated by the reward.

2:39But in the real world of massive AI infrastructure, maintaining that perfect sync is, well, it's physically and computationally nearly impossible. Oh, it's a nightmare. Because modern setups essentially split the AI into two completely different physical engines. Right. You have the trainer, which is this heavy duty engine doing all the complex math to update the model's weights. that's usually running on something like hugging face. Yeah, exactly. And then you have the inference engine. And that one is just a highly optimized, super lightweight engine designed to spit out text as fast as possible, like VLLM.

3:11And that separation is where the entire system starts to kind of fracture. Yeah, it breaks down. You have a trainer and a generator. And even if both of those engines start with the exact same weights, they get out of sync almost immediately. Because the generator is just moving too fast. Way too fast. It's churning out thousands of potential reasoning paths. By the time the trainer calculates the gradient update and tries to send the new weights over, the generator has already moved on. It's lagging behind. Right. The generator is suddenly multiple steps behind the trainer. The whole pipeline is fundamentally asynchronous.

3:44Okay. To put this in a real-world context for you listening, imagine a driving instructor and a student driver. Oh, I like this analogy. So the instructor is the trainer, right, evaluating performance and calculating updates to the student's driving habits. And the student is the inference engine actually steering the car and generating the actions. But they are communicating over a radio that has, say, a 10-second delay. The student makes a sharp left turn, but the instructor doesn't see it until 10 seconds later. Yeah. And by the time the instructor says, you turn too sharply, adjust your steering, the student is already on the highway doing something completely different.

4:23Exactly. The instructor is evaluating the student based on totally outdated maneuvers. You're close, but it's actually worse than that. It's not just a 10 second delay in normal traffic. The student is driving at 200 miles an hour. Oh, wow. OK. Yeah. So in those 10 seconds, they aren't just down the street. They are in a completely different city. And in classic policy gradient theory, that delay literally breaks the math. Because the algorithm is making an assumption that isn't true anymore. Exactly. The algorithms for PPO and GRPO assume the probability distribution of actions perfectly matches the current network weights.

4:58If the data is delayed, the gradient update points in the totally wrong direction. So it could actually hurt the AI. It can literally tear the neural network's internal logic apart. It causes the AI to unlearn good behaviors or, in worse cases, suffer a complete mathematical collapse. Just because the data became off policy. Right. Because the student driving the car right now is not the exact same student the instructor is critiquing from 10 seconds ago. So if this delay is an inherent physical flaw in how we build these massive systems, how have the smartest minds in AI been dealing with it?

5:34Because we clearly have working reasoning models today. Well, they do work, but it involves a lot of mathematical duct tape. Duct tape, nice. Yeah, I mean, the primary workaround the industry uses is a technique called importance sampling. Let's break that down. How does importance sampling actually work mechanically? It's essentially a ratio. The algorithm looks at the probability of an action under the trainer's current state and divides it by the probability of that same action under the generator's lagged state. So it's trying to mathematically estimate how much the policy has drifted. Exactly.

6:07And then it applies that ratio as a weight to the loss function. It's basically the algorithm saying, OK, this data is a bit stale. The student driver was a few miles back when they made this move. So let's discount the value of this lesson by 20 % before we update the model. But I'm guessing that creates its own set of problems. I mean, if you're constantly guessing how stale the data is and applying these arbitrary weights, doesn't that make the learning process incredibly chaotic? What's fascinating here is that it introduces massive variance. If the ratio gets too large, meaning the generator and trainer are too far out of sync, the update just explodes and ruins the model.

6:46So how do they stop that from happening? They use clipping operations. They literally chop off the mathematical update if it crosses a certain threshold just to make sure the changes don't get too large. Or they force the inference engine to slow down, right? Right. They'll explicitly force the lightning-fast inference engine to wait for the trainer to catch up, which just completely destroys your hardware efficiency. And in some extreme cases, if the algorithm determines that a generated sequence is just too off policy, like if the lag is just too great, they just throw the entire sequence away.

7:16Yes. They literally just discard it entirely. And that is the staggering waste we mentioned at the start. You have a massive AI model running on thousands of incredibly expensive GPUs, and it spends precious compute time generating this brilliant 50-step chain of logic to solve a difficult math problem. A very expensive chain of logic, yeah. Right. But then the system throws that perfect answer in the trash simply because the generation engine was a few steps behind the trainer. That's millions of dollars of compute power just being highly inefficient. It's a massive resource drain. Generating those long chains of reasoning is computationally the most expensive part of the process.

7:55Every time you add one of these importance weights or clip a gradient or throw away a sequence, you deviate heavily from the foundational theories of policy optimization. You're starving the model of the data just spent thousands of dollars generating. So if throwing away this expensive data is bankrupting labs, someone had to realize that forcing this perfect synchronization is a losing battle. Did the creators of OAPL just decide to stop fighting the delay? That is exactly the paradigm shift here. OAPL, which stands for Optimal Advantage-Based Policy Optimization with Lagged Inference Policy, completely flips the script.

8:32It's a bit of a mouthful, but okay. Yeah. The acronym is DENT, but the core philosophy is revolutionary for post-training language models. Instead of applying mathematical band-aids to GRPO to force the two systems to stay synced, OAPL completely embraces the off-policy nature of the data. So it just accepts that they are out of sync. It assumes from day one that the generator and the trainer are out of sync, and it treats that as a feature, not a bug. But how do you train effectively if the data is stale? Like you said earlier, PPO and GRPO mathematically collapse when the data is delayed. So how does OAPL process that stale data without destroying the neural network?

9:08OAPL fundamentally changes the mathematical objective. It treats this mismatch as a KL regularized RL problem, and it solves it using a squared regression objective. Well, so those are two very heavy technical terms. Yeah. KL regularization and squared regression. Break down how those actually work inside the AI. Okay. Let's start with KL regularization. Sure. Imagine KL regularization like a heavy bungee cord attached to the AI's current policy. Okay, a bungee cord. Right. So as the AI explores and tries to grab a new reward, like, say, finding a better way to code a Python script, it wants to move its internal weights to reinforce that new behavior.

9:47But the bungee cord physically stops it from snapping too far away from its original starting point. Oh, I see. So it allows for learning, but prevents the model from forgetting everything else it already knew just to chase one specific reward. Exactly. Keeps the learning stable. Okay, so the bungee cord is the stability, but what about the squared regression objective? How does that actually handle the lag? Think back to our driving instructor. Instead of yelling at the student live over a delayed radio, the instructor looks at the simulator tape a week later. The squared regression objective is the instructor's grading rubric.

10:21The instructor says, all right, I know you were driving with last week's rulebook, so I'm going to grade your turn based on what you knew then, not what I know now. Here's where it gets really interesting. Yeah. The squared regression calculates exactly how much tension should be on that bungee cord, even when the AI is completely out of sight. It explicitly anchors the training policy to the lagged inference policy without needing to calculate those messy, important sampling ratios. It adjusts the grading rubric to match the historical timestamp. You got it. Hold on, I have to challenge this, though, because letting the inference engine run wild without talking to the trainer sounds pretty dangerous.

11:00If the student driver is practicing on a simulator for a whole week without any feedback from the instructor, Isn't the student literally just practicing bad habits? It's a fair point. Right. Like, how does the trainer not just learn absolute garbage from reviewing stale data that is full of mistakes? That is the exact skepticism the industry had. But the proof is absolutely in the empirical benchmark data. They tested OAPL against the standard GRPO baseline with important sampling on some of the most grueling math competitions available. Like AME25, HMMT25, Brumo25. Exactly. These are complex, multi-step logical reasoning challenges where a single bad habit in the chain of thought ruins the entire answer.

11:43And how did it actually perform? OAPL didn't just match the GRPO baseline. It outperformed it across the board. It converged to a higher accuracy and remained incredibly stable throughout the entire training process. Wow. But the reason why it outperformed GRPO is the most crucial part. OAPL prevents something called entropy collapse. Entropy collapse. Let's dig into that. What does that actually look like in AI's behavior? Well, in information theory, entropy measures randomness or diversity. During AI training, if sequence entropy collapses, it means the AI has narrowed in on a single, highly specific, and extremely fragile way of thinking.

12:18So it loses its creativity, essentially. Right. Imagine an AI learns to solve math problems only by using complex algebra, even when drawing a simple geometry graph would get to the answer in two seconds. It completely forgets the geometry approach because the algebra path got reinforced early on and its thinking just became rigid. It gets tunnel vision. Exactly. And GRPO struggles massively with this. Because of its rigid clipping mechanisms, you know, chopping off updates to force the model to stay on policy, it frequently suffers from entropy collapse, forcing the AI into that tunnel vision.

12:54But OAPL avoids this. Yes, because it sinks in frequently and uses that bungee cord, the KL regularization, to anchor to the historical inference policy, it maintains a high level of sequence entropy. Meaning the AI keeps a diverse range of problem-solving approaches alive in its network. Right. It doesn't just learn a way to solve the math problem. It keeps its mind open to multiple ways to solve it. That makes it a significantly more robust reasoner. It's not just memorizing a path. It's genuinely exploring logic. But, you know, math is really only half the story for these frontier models. The ultimate test of reasoning is code generation.

13:30Oh, absolutely. And looking at the benchmark data on the live code bench results, this is where the efficiency argument really crystallizes for me. They compared OAPL to a highly capable, publicly available reasoning model called DeepCoder. Right, and DeepCoder was trained using the old way GRPO with all the duct tape heuristics, clipping, and filtering we talked about. And the comparison is just stark. OAPL matched or slightly outperformed DeepCoder across the board on coding tasks. But the input requirement is what totally changes the game. DeepCoder needed roughly 650 ,000 training samples to achieve that level of intelligence.

14:05Yep. OAPL achieved the exact same reasoning power using just 200 ,000 samples. Yeah. It reached the same destination with roughly three times fewer generations. Three times. It is the pure power of off-policy efficiency. By not discarding sequences, and by using the squared regression objective to grade the historical tape accurately, you extract significantly more learning signal from every single attempt the AI makes. That is massive. Three times less compute time required, simply because you aren't throwing away the delayed data. And we really need to emphasize just how extreme the lag was during this coding test.

14:41They allowed the policy lag to reach 400 gradient updates. 400 updates. Just to be clear for everyone, that means the heavy-duty trainer updated the AI's core brain 400 separate times before it finally bothered to sync those new weights back over to the lightweight generator making the actual code. It's an astronomical disconnect. I mean, in standard RL for language models using GRPO, if you had a lag of 400 updates, the mathematical variants would explode and the training would immediately collapse into pure noise. So you'd just break. Completely. But OEPL handled that 400 update delay perfectly stably.

15:14without needing a single importance sampling ratio. Furthermore, this stable high-entropy learning led to incredible test-time compute scaling. Specifically measured by a metric called PAS at K, right? Exactly. Wait, let's pause on PAS at K because I see this metric everywhere in modern AI research now. I know PAS at 1 means giving the AI exactly one try to get the right answer, but how does OAPL change the game when we scale that up to PAS at 64 or PAS at 256? Well, test time compute is basically the new frontier right now. The philosophy is, if you give a model more time to think, generate multiple drafts, and try different logical paths, it should eventually figure out the right answer.

15:55Makes sense. But a model that has suffered from entropy collapse can't do that. If it only knows one way to solve a problem, giving it 64 tries won't help. It will just try the exact same wrong path 64 times. Oh, wow. Yeah, because it has that tunnel vision. Right. Because OAPL prevents entropy collapse and maintains diverse problem-solving strategies, its pass at K-scaling is incredibly strong. As you increase the number of attempts from 1 to 64 or even 256, its ability to explore new logic trees and eventually find the right answer keeps growing significantly. It doesn't just hit a wall. So what does this all mean?

16:32This completely shifts how we think about the infrastructure of AI. Training these massive reasoning models requires astronomical costs and entire warehouses of GPUs. Billions of dollars. Yeah. If developers can train models that are just as smart or smarter using three times less data entirely asynchronously without needing expensive real-time synchronization networks connecting all those computers, I mean, it changes the entire economic landscape of the industry. It is the ultimate bottleneck breaker. Right now, building a frontier reasoning model is a game mostly reserved for massive tech giants with limitless budgets who can afford perfectly synced high-speed infrastructure.

17:08OAPL proves that you can build highly asynchronous, messy, distributed systems and still get state-of-the-art reasoning. This dramatically lowers the barrier to entry. It could democratize the creation of hyper-smart AI, allowing smaller research labs, open-source communities, or startups to compete without needing a billion-dollar supercomputer perfectly synced to the millisecond. If we connect this to the bigger picture, the entire AI industry essentially accepted as gospel that on-policy learning was the only way to do post-training for language models. Everyone was just iterating on the duct tape, you know, tweaking the clipping algorithms, adjusting the important sampling ratios.

17:48Treating the symptom, not the disease. Exactly. OAPL stepped back, looked at classical older reinforcement learning theories regarding off-policy learning, and realized that the fundamental assumption was just flawed. Abandoning that forced synchronization actually solved a cutting-edge problem. It really is a fascinating evolution. We go from a micromanaged student driver constantly interrupted by a delayed radio to a simulator-trained expert exploring every possible route, saving millions of dollars in compute along the way. If you take nothing else away from this, remember this. For years, AI developers thought the only way to teach an AI to reason was through live, real-time micromanagement.

18:28OAPL proves that AI actually learns better, faster, and significantly cheaper when you leave it alone, let it generate a massive amount of diverse experience, and simply review the game tape later with the right grading rubric. That is a brilliant distillation. But, you know, this raises an important question, a truly provocative thought about the future. Okay, I got me. If OAPL proves that an AI can learn stably and efficiently from highly asynchronous, severely lagged data without needing perfect alignment, What does that mean for how AI interacts with the real world? Oh, wow. Right. Could future AI models constantly learn in the background from vast troves of offline historical human data?

19:07If they don't need real time on policy syncing, could they essentially evolve their reasoning skills continuously, absorbing the entire Internet's history of problem solving without ever needing a formal pause training phase in a lab? An AI that just continuously reviews the historical tape of humanity and gets smarter every single day. completely untethered from a live trainer. That is a wild paradigm shifting thought to leave off on. Thank you so much for joining us on this deep dive. Whether you're dealing with a frustrating delay in your own projects or a 400 gradient update lag in a supercomputer, remember that sometimes the best way to move forward is to stop fighting the disconnect and rethink the rules entirely.

19:45Keep your curiosity alive and we'll catch you next time.

From the publisher

Researchers have introduced OAPL, a new reinforcement learning algorithm designed to improve how Large Language Models (LLMs) learn complex reasoning for math and coding. Traditional methods often struggle when the training policy and the inference engine are out of sync, a common issue in large-scale, asynchronous computing. Instead of trying to force these mismatched systems to align, OAPL embraces this discrepancy by using a squared regression objective that functions effectively even with significant policy lag. This approach eliminates the need for complex importance sampling or heuristics that can destabilize training. Empirical results show that OAPL outperforms existing methods like GRPO on competitive benchmarks while using significantly fewer computational resources. Furthermore, the model maintains higher sequence entropy, which prevents the performance collapse often seen in other post-training techniques.

More from Best AI papers explained

All 475 episodes
LLMs Can Learn to Reason Via Off-Policy RLBest AI papers explained · 20 min
Listen in VO