In short
The episode argues that PPO (Proximal Policy Optimization) is poorly suited for LLM reinforcement learning because its trust-region mechanism uses token-level ratio clipping, which over-penalizes rare “reasoning” tokens and under-penalizes harmful updates. It introduces DPPO (Divergence Proximal Policy Optimization) to replace ratio clipping with a distribution-level divergence constraint (binary approximation of KL/total-variation), improving stability and logic learning.
Guest backgrounds
No guest names or bios are provided in the transcript.
Key claims
PPO’s ratio clipping blocks rare but correct reasoning steps (e.g., logarithm-related tokens) while allowing large probability drops for common tokens (e.g., “the”). PPO clipping audits show frequent clipping of math symbols (1, 4, +, =, vector notation) and logic connectives (thus, since, however, therefore). DPPO reduces training/inference mismatch and prevents “panic” from negative samples (over-aggressive punishment of previously likely tokens), avoiding model collapse.
Notable examples
“therefore” token probability spikes causing clipping; rare token 10^-5 to 10^-3 yielding a 100x ratio; common token 0.99 to 0.80 not clipped; PPO clipping “therefore” and math symbols; DPPO outperforming baselines on AIME 2024/2025 and working even without R3.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Reinforcement Learning
0:46 to 2:12
Explains the differences between pre-training and reinforcement learning in AI.
“If you download a training library today, PPO is likely the first thing you'll see.”
The Flaw in PPO
2:13 to 5:08
Details how PPO's safety mechanisms hinder AI learning logic and reasoning.
“It all comes down to something called the trust region, right?”
Exploring DPPO as a Solution
5:09 to 7:32
Introduces DPPO as a new framework designed to improve AI learning capabilities.
“But then the model gets some feedback and realizes that word is actually wrong in this context.”
Technical Innovations of DPPO
7:33 to 10:09
Discusses the mathematical innovations in DPPO and how they enhance AI training.
“DPPO, Divergence Proximal Policy Optimization.”
Stability in Learning Models
10:10 to 12:28
Details how DPPO addresses model stability and prevents catastrophic failures.
“But there's another layer to this stability thing, right?”
Real-World Impact of DPPO
12:29 to 14:00
Highlights the performance improvements seen in AI models using DPPO.
“It allows for failure without total catastrophe.”
Unlocking AI Reasoning with DPPO
14:00 to 15:21
Learn how DPPO enhances AI models by improving logical reasoning capabilities.
“By fixing the over-clipping of those rare tokens, the model can explore reasoning paths much, much faster.”
The Billion Dollar Question
15:21 to 15:55
What cognitive capabilities in AI have been suppressed due to blunt training tools?
“These are the building blocks of thought.”
Transcript
Automatic transcript. May contain errors.0:00Okay, let's unpack this. We talk a lot about AI learning. And you usually have this mental image of a robot sitting in a digital library just memorizing the internet. And sure, that's part of it. That's how they learn to predict the next word. Right. That's the pre-training. That's just the model learning grammar and facts. But that's not really where the reasoning happens. Exactly. The real magic, the stuff that makes an AI helpful or, you know, smart or able to act like a doctor or a coder, that happens in what you could call the gym. That's reinforcement learning. It really is like a gym. You take that book smart model and you force it to solve problems.
0:37If it gets the answer right, you give it a digital treat, a positive reward. And if it gets it wrong, a penalty. You're shaping its behavior. Precisely. And for the longest time, the undisputed king of this gem, the standard algorithm everyone used, was something called PPO, proximal policy optimization. PPO is the default. It's the de facto standard. It is. If you download a training library today, PPO is likely the first thing you'll see. It's been the backbone of the industry for years. But here is where it gets really interesting. It turns out we've been training our language models using an algorithm that wasn't designed for language at all.
1:13Not at all. PPO was built for robotics. Robotics. That's the crux of it. PPO was designed to teach a digital stick figure how to run across a screen. It works great for continuous physical movements. But then, as an industry, we basically just copy-pasted that logic into the world of large language models. And that copy-paste job has apparently caused this massive invisible problem. Because today we are diving into how PPO, specifically the safety mechanism inside it, is actually sabotaging the AI's ability to learn logic. It's actively stopping the model from learning how to reason. It's a classic case of a safety feature becoming a roadblock.
1:52So we put training wheels on the AI to keep it stable. But it turns out those training wheels are square. And today we're going to look at the fix, a new framework called DPPO, or Divergence Proximal Policy Optimization, which finally gives us round wheels. I love that. Square training wheels. So let's get into the mechanics of this failure. Why is PPO breaking our AI brains? It all comes down to something called the trust region, right? Right. The trust region. This is maybe the most important concept in reinforcement learning stability. Just imagine you're learning to play tennis. You have a swing that works okay.
2:29You can get the ball over the net. Now, you want to improve. Okay, so I want to tweak my form a little. Exactly. But if you change your swing completely overnight, like you start holding the racket upside down or something, you're just going to play terrible tennis. Your performance will collapse. Right. I'll forget everything I already knew. So the trust region is just that safety zone. It says you can change your swing, but don't stray too far from what worked yesterday. You make small incremental updates so you don't break the model. And PPO enforces this safety zone using something called ratio clipping.
3:01Yes. And this is really the villain of our story, ratio clipping. Okay, explain this to me. How does ratio clipping actually work? It's surprisingly simple math, which is probably why engineers liked it so much. PPO looks at a specific word, a token, and it compares how likely the model was to say that word before the update versus after the update. So it calculates a ratio, new probability divided by old probability. That's all it is. So if you used to say hello 10 % of the time, and now you say it 20 % of the time, the ratio is 2. Simple enough. And PPO sets a limit. So if that ratio goes above, say, 1.2, it clips the update.
3:39It says, whoa, slow down. That's too big of a change. And it just ignores the rest to keep things stable. That sounds reasonable on paper. Keep it stable. Don't let the model go wild. So where's the flaw? The flaw is that language isn't a robot arm. A robot arm's movement is continuous. But language is discrete. It's sparse. An LLM has a vocabulary of, what, 100 ,000 words? Most of those words at any given moment have a near zero chance of being chosen. You're dealing with tiny, tiny numbers. So tiny. And this is where the aha moment happens if you walk through the math. Let's say the model is solving a hard math problem.
4:12And there's a very specific step, maybe it involves a logarithm, that it has almost never used before. The probability is like 10 to the minus 5. Okay, 1 in 100 ,000, basically zero. It's not even on the model's radar. Basically zero. Now, during training, the model gets a reward and it realizes, hey, logarithms are actually the key here. So it bumps the probability up to 10 to the minus 3, 1 in 1 ,000. It's still a very small number. It's not like it's screening logarithm. It's just dipping its toe in. Exactly. In terms of absolute change, it's a tiny shift. But look at the ratio. 10 to the minus 3 divided by 10 to the minus 5.
4:49That's 100. It's 100x increase. Right. PPOC is 100x spike and it just freaks out. It thinks the model is exploding. So it aggressively clips this update. It just slams the brakes on this learning moment. So even though the change was tiny and safe and actually a really smart move, PPO blocked it because the ratio looked scary. Correct. It punishes the model for discovering new, rare ideas. Now let's flip it. Take a common token. Something the model is 99 % sure of. Like the word the. Probability is 0.99. Okay, super confident. But then the model gets some feedback and realizes that word is actually wrong in this context.
5:27So it drops the probability to 0.80. That feels like a huge drop. You're going from almost certain to just pretty sure you lost almost 20 % of your confidence. It is a massive shift in the model's thinking. But do the math again. 0.8 divided by 0.99, it's about 0.8. That ratio is very close to 1. PPO looks at that and says, eh, looks fine. No clipping needed. Wow. OK, so it's completely backwards. It punishes the model for exploring brand new concepts, but it lets the model just destroy its foundational knowledge without even blinking. It over penalizes exploration and under penalizes instability.
6:01It creates this awful learning dynamic where the model is fighting its own safety rails just to learn something new. And this isn't just theoretical math, right? Because when they actually looked under the hood at what PPO was clipping, that's the part that really floored me. It wasn't just clipping random junk. No, it was a noise. This is the smoking gun. They audited the training runs to see exactly which tokens were hitting this PPO speed limit. And the list is just, it's like a who's who of intelligence. It really is. The tokens that were getting clipped the most were numbers. One, four. They were math symbols.
6:36The plus sign, the equal sign, vector notation. So the model literally tries to do math and PPO tells it to stop. And it gets worse. It was clipping logic words. Words like thus, since, however, therefore. Therefore. That's literally the word you use when you have a breakthrough. A plus B equals C therefore. Exactly. And think about why. When a model is struggling to learn how to reason, those logic words are rare. It doesn't know how to use them yet. So they start with a very low probability. The moment the model tries to use one, the moment it tries to reason, the probability spikes, the ratio blows up, and PPO clips it.
7:13That is just, it feels tragic. We are actively preventing the model from learning logic because the algorithm thinks a sudden insight looks like instability. It explains so much about why training reasoning models has been so inefficient for so long. We've been forcing them to learn with one hand tied behind their back. We're telling the AI to shut up every time it has a good idea. So enter the solution. DPPO, Divergence Proximal Policy Optimization. How does this finally fix the square wheels? So DPPO flips the script entirely. Instead of looking at a flimsy ratio of a single word, it looks at the divergence of the entire probability distribution.
7:50Divergence. Okay. Unpack that for us. PPO is asking a really noisy, unreliable question. Did the probability of this specific word change too much? DPPO asks a much better one. Did the overall meaning and the distribution of all my choices change too much? So it's looking at the big picture, not did my pinky finger move an inch, but did my entire tennis swing actually break? That's a great way to put it. It uses a mathematical concept called total variation or KL divergence. It's a principled way to measure the distance between two probability distributions. It doesn't care if a rare word jumped 100x as long as the total probability mass didn't shift in a dangerous way.
8:30That sounds much smarter, but I'm sensing a but here. If you have to check the entire distribution, and an LLM has a vocabulary of what? 100 ,000 words. Isn't that mathematically way too heavy? It would be. Calculating the full divergence for every single update would be memory prohibitive. Way too slow. You'd run out of GPU memory instantly. So how do they get around that? They use this really brilliant simplification. They call it a binary approximation. Binary zeros and ones. Right. Instead of calculating the shift for every single word in the dictionary, they just collapse the entire world into two buckets.
9:05Bucket one is the token the model actually chose. And bucket two. Is everything else. That's it. Just the thing I did versus all the things I didn't do. That's it. And it turns out that simple binary split captures about 99 % of the necessary signal. It lets the algorithm detect if the model is making a dangerous move without needing to track every single synonym in the dictionary. They also tried a top K approach, right? Like looking at the top 20 words or something. They did. They thought maybe they needed more detail. But surprisingly, the simple binary approach worked just as well. It's elegant, it's fast, and it completely solves that rare token problem.
9:41So now if I have that rare therefore token that jumps from 10 to the minus 5 to 10 to the minus 3, DPPO looks at the binary split and says what? It says the total mass that moved is tiny. This is safe. Go ahead and learn. And if I drop my confidence on a common word from 0.99 to 0.80, DPPO sees a huge chunk of probability mass moving from the first bucket into the everything else bucket and says, whoa, that's a major shift in behavior. We need to check this. It restores the correct safety rails. It's amazing how just fixing the math aligns the incentives properly. But there's another layer to this stability thing, right?
10:16Something called the training inference mismatch. Now, this sounds a bit deeper in the weeds, but I think it's important because it explains why models sometimes just go crazy. This is a deep technical issue, but it's crucial. When you train these huge models, you're doing it on massive clusters of GPUs using lower precision math floating point numbers that aren't perfectly precise. We do this to save memory and to go faster. Just tiny rounding errors. Tiny rounding errors. But over billions and billions of calculations, they add up. It means the model that is learning, the one calculating the gradients, isn't exactly the same as the model that's generating the data.
10:54There is a mismatch. So the map doesn't quite match the territory. Exactly. And without a strong trust region, the model starts learning from these errors. It starts chasing ghosts. And this leads to what we call model collapse. The model just starts gibbering or repeating itself or its performance just panks. And they did a sanity test on this, right? They did. They took a standard data set of math problems, the math data set, and they just let it train. The standard methods, like the ones that came before PPO, they all collapsed. The mismatch just grew and grew until the modder broke. And DPPO.
11:27Rock solid. It maintained a low mismatch, stable training the whole way through. But the most fascinating part of this experiment was why models collapse. They actually isolated the bad updates. The bad updates. This sounds like a detective story. Who killed the model? They found that it wasn't the positive updates. When the model got a question right, they caused the problem. It was the negative samples, the wrong answers. The punishment. Specifically, it's when the model tries to aggressively penalize a token it previously thought was good. So imagine the model is 99 % sure the answer is 4, but it turns out to be wrong.
12:01The model panics and tries to smash that probability down to zero instantly. A total panic reaction. I was wrong. Burn it all down. That's exactly what it is. And that panic reaction creates this huge divergence that destabilizes the entire system. DPPO's mask effectively blocks these specific bad updates. It says, OK, you were wrong, but let's not, you know, lobotomize ourselves to fix it. Let's just adjust slowly. So it stops the model from overreacting to its own mistakes. it acts like a shock absorber. Precisely. It allows for failure without total catastrophe. Let's talk results. Because all this theory is great, but does it actually make the AI smarter?
12:42The benchmarks are pretty undeniable. They tested this on AIM 2024 and AIM 2025, and these are very hard math competitions. We're talking high school Olympiad level stuff. About adieu. DBPO outperforms GRPO, which is the current, you know, Steve the Art PPO variant used by the big labs. And it does so pretty significantly. It learns faster, it gets better rewards and fewer steps, and it reaches a higher final score. There was a detail about something called R3 that seemed like a big deal. Oh, the rollout router replay. Yes, R3. It sounds like a droid from Star Wars. It's a heavy-duty technique you usually need for these mixture of experts models.
13:16Those are the models that have different little experts inside them for different tasks. And usually, keeping them stable is a nightmare, so you need this R3 technique to keep everything in sync. It's like expensive scaffolding. And with DPPO. They found that DPPO is so naturally stable, it beats the baseline models even without using R3. Wait, so you can just strip out this heavy, complex stabilization technique and the algorithm itself is robust enough to handle it. Yes. And when you add R3 back in, DPPO gets even better, but it doesn't need it to survive. That's a huge testament to how broken the old way was.
13:49We were adding layers of complexity just to banish the fact that our core mechanism was flawed. We were putting duct tape on square wheels, and now we just have round wheels. That's the perfect way to put it. By fixing the over-clipping of those rare tokens, the model can explore reasoning paths much, much faster. It stops fighting the algorithm, and it just starts learning the math. It's so funny. We always worry about AI becoming too powerful, but in this specific case, we were actively holding it back from being logical. We were. We moved from a heuristic designed for robots ratio clipping to a mathematically principled approach for language divergence constraints.
14:25And the difference is just, it's night and day. So what does this all mean for you listening? We aren't all training LLMs in our basement. Why should we care about DPPO? It matters because the reasoning barrier is the current frontier of AI. Everyone is waiting for models that can plan and code reliably and do research. DPPO shows that one of the biggest bottlenecks wasn't a lack of data or lack of chips. It was just blunt math. We were using a hammer to do brain surgery. Exactly. And now that we have a scalpel, we can expect the next generation of models to be much more efficient. They'll learn logic faster because we finally stopped punishing them for it.
15:02The safety rails we put on AI training were actually hindering its ability to learn logic. By fixing the rails, we don't just get stability, we actually unlock the model's ability to reason. And that's the key takeaway. Stability and reasoning shouldn't be a trade-off. With the right math, you get both. I want to leave everyone with a final thought here. We saw that standard PPO was clipping words like therefore and because and since. These are the building blocks of thought. They're the connective tissue of logic. Right. So if our standard algorithms have been actively suppressing these reasoning tokens for years, what other cognitive capabilities have we been accidentally clipping out of our AI models simply because our training tools were too blunt?
15:45That is the billion dollar question, isn't it? We might find that AI is a lot more creative or nuanced than we ever gave credit for once we stopped telling it to shut up every time it has a new idea.
From the publisher
This research paper introduces Divergence Proximal Policy Optimization (DPPO), a novel reinforcement learning framework designed to improve the fine-tuning of Large Language Models (LLMs). The authors identify a structural flaw in the standard PPO algorithm, noting that its ratio-clipping mechanism incorrectly penalizes rare tokens while failing to stop destabilizing shifts in common tokens. To resolve this, DPPO replaces heuristic clipping with a principled trust region constraint based on direct estimates of policy divergence. Because calculating exact divergence is memory-intensive, the study proposes Binary and Top-K approximations to maintain efficiency without sacrificing performance. Theoretical proofs and empirical tests demonstrate that DPPO achieves significantly better training stability and efficiency than existing methods like GRPO. Ultimately, the framework provides a more robust foundation for aligning LLMs with complex reasoning tasks and human preferences.




