LLMs Can Learn to Reason Via Off-Policy RL

27 Feb 2026 · 20 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Off-policy RL for large language model (LLM) reasoning, proposing OAPL (Optimal Advantage Based Policy Optimization with Lagged Inference) to fix training instability caused by “on-policy” synchronization lag.

Guest backgrounds

No guests mentioned; only two hosts discussing the paper and experiments.

Key claims

On-policy assumptions (PPO/GRPO) require inference data from the exact current model; distributed lag makes data stale and causes exploding importance-sampling ratios, destabilizing training. OAPL replaces ratio-based updates with a regression-style objective using “optimal advantage,” stabilized by KL regularization toward the lagged inference policy. This enables “lazy synchronization,” allowing 50–400 gradient steps of inference lag without collapse.

Notable examples

Benchmarks AIM 2025 and Harvard-MIT math tournament; smoother training curves vs GRPO. Coding: OAPL matches DeepCoder while using ~200k samples vs ~650k for DeepCoder. Entropy collapse mitigation shown via pass@K scaling up to K=256, widening the gap over GRPO.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The AI Revolution's Surface vs. Reality

0:45 to 2:11

Discussing the superficial allure of AI and the underlying training issues.

“Because there's a massive problem down there.”

Understanding OAPL and Its Importance

2:11 to 4:05

Introduction to OAPL as a solution to synchronization issues in AI training.

“To understand why OAPL is such a massive breakthrough, we kind of have to understand what is currently broken.”

The Problems with Current Algorithms

4:05 to 6:01

Explaining the limitations of existing algorithms like PPO and GRPO in AI training.

“Because you mentioned distributed systems.”

Shifting the Paradigm with OAPL

6:01 to 9:01

Detailing how OAPL redefines learning processes for AI models.

“It's like trying to learn how to play golf.”

Operational Benefits of OAPL

9:01 to 10:37

Exploring how OAPL changes data generation efficiency in AI training.

“It anchors the learning process so the model doesn't just spin out of control.”

OAPL vs GRPO Performance

10:37 to 11:23

Comparing OAPL's performance against traditional GRPO methods in AI training.

“But does it actually produce a smart AI?”

The Cost Implications of Data Efficiency

11:23 to 13:21

Discussing the financial impact of improved training algorithms in AI.

“But honestly, looking at the charts, it's the way it won that really matters.”

Understanding Entropy Collapse

13:21 to 14:00

Explaining the concept of entropy in AI training and its implications.

“We've been throwing away the nutrients because our stomach, the algorithm just couldn't digest them.”

Entropy and Exploration in AI Training

14:00 to 16:00

Learn about the importance of high entropy during AI training and how it affects solution exploration.

“High entropy means the probability is spread out over many options.”

OAPL vs. GRPO: A Comparative Analysis

16:00 to 18:00

Discover how OAPL outperforms GRPO in reasoning capabilities, showcasing diverse strategies.

“Meaning OAPL benefited significantly more from having those extra attempts.”
Show all 11 chapters

The Human Parallel: Embracing Lag for Better Learning

18:00 to 20:00

Understand how the concepts from AI training apply to human learning and collaboration.

“When you can get the exact same performance with three times less data, the market is going to move there.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. So right now, I think it's safe to say the entire tech world is just absolutely obsessed with the shiny part of the AI revolution. Completely. You know, you look at models like DeepSeek R1 or the latest stuff from the big labs and you just marvel at the reasoning. The inner monologue. Exactly. We love seeing that. The model pauses. It thinks, maybe corrects itself, and then boom, it gives you this brilliant answer. It feels like magic to you as the user. It really does feel like magic on the surface. But if you actually lift the hood and look at the engine room, it is incredibly messy in there.

0:37And that is exactly what we are doing today. Our mission for this deep dive is to ignore that shiny paint job and really climb down into the grease and the gears. Because there's a massive problem down there. Right. We've got this snack of new papers and benchmark data, and they all point to this huge expensive bottleneck in how these thinking models are actually trained. And honestly, almost nobody is talking about it. It's the classic iceberg problem. You know, we see the clever answers above the water. But underneath the infrastructure required to train these things is hitting a serious wall.

1:10We're burning millions of dollars in compute time? Millions. Easily. Because we are trying to force these massive supercomputers to act like perfectly synchronized telepathic twins. When in reality, they're constantly drifting apart. Drifting apart, which is a very polite way to put it. From what I've been reading in these papers, it sounds more like a terrible long distance relationship. Yeah, that's a good analogy. But today we're actually talking about a fix for that relationship. A totally new paradigm called OAPL. OAPL, which stands for Optimal Advantage Based Policy Optimization with Lagged Inference.

1:47A total mouthful as always in AI. It is, but the promise here is just huge. This new approach claims to fix that synchronization bottleneck, stabilize the training of these reasoning models, and this is the part that really caught my eye, make the whole training process three times more data efficient. Which, in the world of billion-dollar training runs, is literally the difference between a profitable product and an absolute money pit. Exactly. So let's unpack this for you. To understand why OAPL is such a massive breakthrough, we kind of have to understand what is currently broken. You mentioned those telepathic twins.

2:23Who are the twins in this scenario? So in the post-training phase of an AI, specifically when we're doing reinforcement learning or RL, which is how we actually teach these models to reason, you effectively split the AI system into two totally distinct roles. Okay. You have the trainer and you have the inference engine. All right, let's define those. The trainer is like the boss. Think of the trainer as the brain that's currently undergoing surgery. It is the policy, the neural network itself, that is actually being updated. Got it. It computes the gradients. It does all the heavy calculus. And it changes its own synaptic weights to get smarter.

2:58It's the one that's actually learning. And the inference engine. The inference engine is essentially the student doing the homework. Its only job is to generate text. It spits out code or solves math problems or writes essays. We call these rollouts. Right. And in a massive training run, this isn't just one chip doing this. It's distributed across hundreds or even thousands of GPUs, just churning out enough data for the trainer to learn from. Okay, so the inference engine does the practice problems, like thousands of them, and then the trainer looks at the answers to decide how to update the brain.

3:31That's the basic intuition, yeah. Now, traditionally, algorithms like PPO, Proximal Policy Optimization, or even the newer GRPO, which is what DeepSeq uses, they rely on a really strict rule. It's called the on-policy assumption. I've seen this term a lot in the sources. On policy. It basically means the student and the teacher have to be on the exact same page, right? Strictly on the same page. It means the data used to learn must come from the exact version of the model currently being trained. So if the trainer is on version 5.0, the homework must be generated by version 5.0. No exceptions.

4:04But here is where reality interferes, right? Because you mentioned distributed systems. If I have a thousand GPUs generating text and just one central brain updating weights. You get lag. Yeah. It is completely unavoidable. The trainer updates its weights. It becomes version 5.1. But the inference engine is physically located on a different rack, maybe? Or it's just finishing up a massive batch of homework using version 5.0. So by the time that data gets back to the trainer, it's stale. It's technically off policy data now. It's like expired milk. Exactly. And it gets even more subtle than that.

4:40Even if the weights are identical, say the trainer is using a different software kernel, like a hugging face implementation, and the inference engine is using something like VLLM just to make it run faster. Oh, so the software itself is different. Right. So the floating point math might differ just slightly. The exact probabilities of the words won't match perfectly. So you have a mismatch. The student is handing in homework based on yesterday's textbook or maybe just a slightly different edition of the textbook. But why is that such a catastrophe? I mean, humans learn from old textbooks all the time.

5:11Humans are adaptable. Mathematical formulas are not. The standard algorithms like GRPO handle this mismatch by using a technique called importance sampling. Ah, yes, the Band-Aid approach I've heard about. It is a very sophisticated band-aid, but yes, important sampling tries to mathematically translate the old homework so the new model can understand it. How does it do that? It calculates a ratio. It basically asks, how likely is this data under the new model versus the old model? So it's just a correction factor? Right. If the models are really close, the ratio is near 1.0, everything is fine.

5:46But if the student is far behind or if the thinking process took a completely different path, that ratio can absolutely explode. It can become huge. And when you plug a huge number into a delicate gradient descent update. You grow up the model. The training just destabilizes. The variance becomes too high. It's like trying to learn how to play golf. And on one swing, your coach just screams at you through a megaphone at maximum volume. You don't learn anything. You just panic. You just panic. Yeah. So engineers spend their lives fighting this. How? What do they do? They add things like clipping where they artificially cap the ratio.

6:21They'll say, if the ratio is higher than 1.2, just pretend it's 1.2. Which feels like cheating the math. It is. Or they filter out data that looks too weird. Which sounds incredibly wasteful. You are generating all this data, paying for the electricity, paying for the GPUs, and then you're just throwing it away or handicapping it because your math can't handle a little bit of lag. Precisely. You are fighting your own infrastructure. You're pausing your supercomputer to wait for synchronization, or you're tossing out totally valid experiences because they don't fit that strict on policy rule. It's super inefficient and it's unstable.

6:56Okay, so this is the setup. We have a rigid system that demands perfect synchronization in a highly asynchronous world. So enter OAPL. What is the fundamental philosophical shift here? OAPL basically stops fighting. It asks a very simple, somewhat arrogant question. It asks, is on-policy learning actually necessary for reasoning models? And the answer it comes back with? A resounding no. So it embraces the lag. But how do you do that? You can't just feed bad math into a neural network and hope for the best right. No, you change the math entirely. Instead of using those volatile importance sampling ratios, OEPL treats the whole problem as a squared regression task.

7:36Okay, let's unpack that for a second. Because regression usually means fitting a line to data right. How does that apply to training a superintelligence? Think of it this way. In the old method, you're constantly asking, how much better is this new step compared to the old one? And then scaling it by that dangerous exploding ratio. OAPL changes the objective. It looks at the lag data, the old homework, and calculates something called the optimal advantage. Optimal advantage. Basically, it asks, what was the best possible outcome from that situation, regardless of which version of the model actually generated it?

8:08It estimates the value of the action using the rollouts from the lagged engine directly. And then what? Then it simply tries to minimize the difference between the trainer's current policy and that optimal outcome. It's a direct regression loss. Yeah. It's just trying to close the gap between what I would do now and what the best outcome was. Hold on. So it just doesn't care that the data came from an old model. Oh, it cares. Yes. But it handles it via something called KL regularization. This is the real secret sauce of OAPL. KL divergence is a mathematical way of measuring how different two probability distributions are.

8:44Right. So OAPL regularizes the training policy towards the inference policy. In plain English for the rest of us. It effectively says, hey, learn from this experience, but don't drift too far away from the logic that originally created it. Ah, I see. This keeps the math stable without needing those exploding ratios. It anchors the learning process so the model doesn't just spin out of control. So by swapping out that volatile ratio math for this cleaner regression math, they just eliminate the instability entirely. Correct. There's no clipping, no deleting stale tokens, just clean math. And because the math is so stable, it unlocks this massive operational benefit called lazy synchronization.

9:25Lazy synchronization. Now, lazy is usually a pretty big criticism in tech, right? Lazy loading, lazy code. But here, I'm assuming it's a feature. Oh, it's a superpower. hour. Because OAPL doesn't panic over stale data anymore, they don't have to sync the trainer in the inference engine constantly. So they can just let it run. Exactly. In the experiments in the papers they showed, they could let the inference engine lag behind for 50, 100, even 400 gradient steps. Wait, 400 steps. In the old GRPO world, wouldn't that just absolutely destroy the model? Oh yeah. In the GRPO world, if you lag by even a few steps without heavy artificial clipping, the model collapses.

10:01At 400 steps it would probably be generating absolute gibberish or just getting stuck in an infinite loop. But OAPL handled it. OAPL handled 400 steps of lag and it still kept learning. That is great. I mean that's the engine room fix right there. You can just let your GPUs run wild generating data at their own pace and the trainer just absorbs it whenever it's ready. It totally unblocks the pipeline. It does. It turns a synchronous bottleneck into an asynchronous flow. It allows you to scale up to thousands of GPUs without worrying that one slow chip in the corner is going to hold up the entire billion dollar training run.

10:36That makes total sense from an infrastructure standpoint. But does it actually produce a smart AI? I mean, we can make a factory very efficient at producing garbage. Just because the math is clean doesn't mean the actual reasoning is good. And that is the ultimate test. It's easy to be efficient if you aren't doing a good job. But they ran OAPL against the absolute heavyweights. They compared it directly to GRPO, which, remember, is the current state of the art used by DeepSeek and others. And they tested it on benchmarks like AIM 2025 and the Harvard-MIT math tournament. And these are serious Olympiad-level math problems.

11:09This isn't just high school algebra. Oh, these are extremely hard problems. Problems that require multiple steps of deep reasoning, backtracking intuition. And OAPL outperformed GRPO on almost all of them. Wow. But honestly, looking at the charts, it's the way it won that really matters. How so? If you look at a training curve for GRPO, it looks like a seismograph during a major earthquake. It just bounces up and down. Because of that panic you mentioned earlier. Right. It gets a batch of data. It learns a bit. Then it gets confused by the lag or the clipping. The performance drops sharply. Then it slowly recovers.

11:44It's incredibly volatile. But OAPL. OAPL's curve is just completely smooth. It just steadily, relentlessly climbs up. It's the difference between a panicked student cramming for a test all night and a student who is just calmly mastering the material over a semester. I love that analogy. And then there's the coding benchmark. Because I saw a stat in the source material here that honestly stopped me in my tracks regarding the efficiency. You're talking about the live code bench test. Yes. Yeah. Break that down for us. So they compared an OAPL-trained model against a model called DeepCoder. Now, DeepCoder is a very solid model, but it was trained using GRPO with all the usual hacks, the clipping, filtering out long sequences, all of that.

12:25And what was the result? OAPL completely matched DeepCoder's performance. They ended up at the exact same level of coding capability. Here's the kicker. DeepCoder required about 650 ,000 samples to get there. Okay. Okay. OAPL used approximately$200 ,000. $200 ,000. So it got to the exact same destination using less than a third of the gas. Exactly. Less than a third. And think about the cost implications of that. Yeah. If you're a startup or even a big lab like OpenAI or Anthropic, compute is your biggest expense by far. Oh, easily. If you can cut your data generation costs by over 66 % just by changing the underlying mathematical algorithm, that is literally the difference between commercial viability and bankruptcy.

13:08It's wild. It really implies that a lot of this quote unquote data hunger we always talk about in AI might just be algorithm inefficiency. Like we don't necessarily need more data. We just need to stop wasting the data we already have. That's exactly it. We've been throwing away the nutrients because our stomach, the algorithm just couldn't digest them. OAPL changes the entire metabolism of the model. Okay, I want to pivot a bit to a concept that appeared in the discussion of OAPL that sounds a little more abstract that seems really critical to how this all works. Sure. Entropy collapse. It sounds like the end of the universe, but in AI training, it's a very specific failure mode, right?

13:45Break this down for us. So entropy in the context of a language model is basically a measure of uncertainty. Or to put it in more human terms, you can think of it as curiosity. Curiosity, I like that. Think of it as the model's willingness to explore different options. High entropy means the probability is spread out over many options. It's considering many different ways to solve a problem. And low entropy. Low entropy means it is 99 % sure that this one specific word or this one specific method is the only way to go. And generally, we want high entropy during training. We want it to stay high for a long time.

14:20We want the model to actively explore the solution space to try different reasoning paths. Right. The huge problem with GRPO and these on-policy methods is that they suffer from entropy collapse very quickly. Meaning they find one trick that works and they just spam it. Exactly. They get a reward for a specific type of answer and they immediately lock onto it. They completely stop exploring. They become rigid. They become dogmatic. So how does OAPL prevent this dogmatism? Ironically, the lag actually helps here. Yeah. Because the inference policy, the student's slightly behind the trainer, and because of that KL regularization we talked about.

14:56The anchor. Right, the anchor. The model is anchored to a slightly older, potentially much more diverse set of behaviors. Ah, so it forces the model to keep its options open. It prevents the model from rushing too quickly into a single narrow valley of thought. It maintains that exploration. And we know for a fact this worked because of a metric called pass at K. Pass at K. This is a pretty standard reasoning metric, right? Yes, very standard. It's very simple to understand. Pass at 1 just means, did you get the answer right on your very first try? Okay. Pass at 100 means, if I let you try 100 times, did you get it right at least once?

15:32So it measures the model's ability to brainstorm correct answers. Correct. Now, if a model has suffered from entropy collapse, If it only knows one way to solve the problem, then giving it 100 tries doesn't help at all. Because it's just going to repeat the same wrong answer 100 times. Exactly. Or the same right answer, but it can't find any alternative routes. But with OAPL, when they scaled K from 1 all the way up to 256 attempts, the performance gap between it and GRPO actually widened. Meaning OAPL benefited significantly more from having those extra attempts. Yes. Which proves it didn't just memorize a single path to a solution.

16:09It actually learned a highly diverse set of reasoning strategies. It actually settles a major debate in the field right now, doesn't it? Because you hear skeptics say reinforcement learning doesn't actually teach reasoning. It just sharpens the model's confidence in what it already knew from pre-training. Like giving it a pep talk rather than a lecture. Yeah, just hyping up the model to give the answer it already had. Exactly. But if OAPL improves so drastically on pass at K, it implies it has actually acquired new reasoning capabilities. It has filled its toolbox with totally new tools, not just polished the one hammer it already had.

16:45That distinction is so crucial. It's the difference between a confident parrot and an actual problem solver. And it's all because we stopped trying to force the student to be perfectly synchronized with the teacher. We actually allowed for that divergence. It's an incredibly compelling narrative. We moved from fighting the lag with all these messy mathematical hacks to actually using the lag to stabilize the whole system. It really challenges this deep-seated intuition we have that faster feedback is always better. You know, in control theory or even in management, we usually assume that the tighter the feedback loop, the better the performance.

17:19You want instant correction. But here a looser loop, a 400-step delay, resulted in a much more robust intelligence. It suggests that the future of training these massive reasoning models, when we call this system two thinkers, is going to be fully asynchronous. Right. We're going to see pipelines that are much more relaxed about time, but much stricter about the true mathematical value of the data. So if you are a researcher listening to this or just someone building with LLMs, the real takeaway is that on policy might just be a legacy constraint that we can finally ditch. I genuinely think on policy learning for post-training is on its way out.

17:55The efficiency gains of off-policy methods like OAPL are just too massive to ignore. When you can get the exact same performance with three times less data, the market is going to move there. You simply can't afford not to. It's fascinating how the solution to a cutting-edge AI problem ends by being something like, hey, just relax a little bit. It's often the case, honestly. Complexity is usually the enemy. Simplicity wins out. Before we wrap up, I really want to touch on that human parallel you hinted at earlier. because we usually end with a takeaway for you, the listener. And this whole lag concept feels incredibly applicable beyond just banks of GPUs.

18:32I've been thinking about this a lot, actually, while reviewing this material. We live in this world of Slack and Teams and instant notifications. We just obsess over real-time synchronization. Right. If you don't reply in five minutes, you're totally out of the loop. We treat lag as a fundamental bug in our organizations. We basically want everyone on policy all the time. That is so true. But look at what happened here with the AI. The student model performed significantly better when it was allowed to drift, when it was allowed to process a large batch of experience entirely on its own terms before finally syncing up with the teacher.

19:05It maintained high entropy. It stayed curious. Maybe constant, immediate helicopter parent-style feedback actually causes entropy collapse in humans too. Oh, wow. If you are corrected the very second you make a mistake, you never actually explore the solution space. You just learn to please the teacher. You learn the one exact path that avoids the buzzer. You become completely rigid. You optimize for the immediate reward rather than actual deep understanding. So maybe the lesson here is to consciously embrace the lag. You know, sleep on it. Let your team run off policy for a while. You might just find that the variance goes down and the actual reasoning goes up.

19:47I really love that. We are always told to fail fast, but maybe we should actually be failing slowly and in batches. Or at the very least learning asynchronously. That is OAPL in a nutshell. Clean math, lazy sync, and much better reasoning. It is a massive step forward for the efficiency of AI and honestly a pretty good reminder for the rest of us. Thank you so much for breaking this all down. It was a pleasure. And thank you for listening. Go find some efficiency in your own leg, and we'll see you in the next Deep Dive.

From the publisher

This research introduces Optimal Advantage-based Policy Optimization with Lagged Inference policy (OAPL), a novel reinforcement learning algorithm designed to improve Large Language Model (LLM) reasoning. Traditional methods like GRPO often struggle with "off-policy" data caused by technical mismatches between training and inference engines. OAPL embraces these discrepancies by using a squared regression objective and KL-regularization, allowing the model to learn effectively even when data is significantly outdated. Empirical tests show that OAPL outperforms existing benchmarks in competition mathematics and matches top-tier coding models while using three times fewer training samples. Furthermore, the algorithm prevents entropy collapse, ensuring that the model maintains diverse and scalable problem-solving capabilities during test-time. Ultimately, the authors demonstrate that fully asynchronous, off-policy training is a more stable and efficient path for advancing machine reasoning.

More from Best AI papers explained

All 475 episodes
LLMs Can Learn to Reason Via Off-Policy RLBest AI papers explained · 20 min
Listen in VO