Principled RL for diffusion LLMs emerges from sequence level perspective

11 Dec 2025 · 12 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Why standard token-level RL (e.g., GRPO/PTO) fails for diffusion LLMs, and how ESPO (ELBO-based sequence-level policy optimization) fixes it by treating the whole sequence as one action and using a sequence-level ELBO proxy.

Guest backgrounds

No guests are mentioned in the transcript.

Key claims

Diffusion LLMs generate by iterative denoising over the whole sequence, so token-factorized RL probabilities are ill-defined. ESPO uses the sequence ELBO as a rigorous variational lower bound. Stability requires normalizing the log-ratio by sequence length L, and using a KL estimator K2 (quadratic) instead of K3 (exponential) or K1 (zero gradient).

Notable examples

Sudoku (global rule satisfaction; up to 60+ point gains; avoids garbled rule-violating grids), Countdown (20–40 point gains), plus smaller but consistent gains on math/coding. Computationally practical: ~47% FLOPs increase even with 4x Monte Carlo samples.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Shift to Diffusion LLMs

0:45 to 2:24

Understanding the transition from autoregressive models to DLLMs.

“It, like, unlocked massive performance gains.”

Challenges with RL in DLLMs

2:24 to 4:36

Exploring the difficulties faced in applying standard RL methods to DLLMs.

“They need to be able to calculate the probability of the whole sequence by multiplying the probabilities of each individual token.”

The ESPO Framework: A Paradigm Shift

4:36 to 6:34

Introducing the ESPO framework for optimizing DLLMs using sequence-level decisions.

“It's only mathematically rigorous for the entire sequence.”

Stabilization Challenges and Solutions

6:34 to 8:16

Discussing the instability issues with RL and the innovative stabilization techniques used.

“Their stabilization technique was mathematically simple but absolutely critical.”

Performance Gains and Validation

8:16 to 10:26

Examining the performance improvements and benchmarks achieved with ESPO.

“The problem is, in this specific context, the K1 estimator's gradient is zero.”

Cost Efficiency of ESPO

10:26 to 11:47

Analyzing the computational efficiency and practicality of the ESPO framework.

“The biggest computational cost in training DLLMs is the generation phase, all those denoising steps.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. If you've been following the world of AI, you know the field is just. It's undergoing this massive architectural shift. For years, everything, and I mean everything, revolved around autoregressive models, AR models, the classic left-to-right token-by-token generation. But now there's a new player, Diffusion Large Language Models, or DLLMs, and they are fundamentally different, a whole new frontier. So if you're the learner today, our mission is to cut right through that complexity. We're going to dive into why reinforcement learning, RL, just completely broke what applied to these new models.

0:35And then we'll unpack a really brilliant new framework called ESPO. That's ELBO-based sequence-level policy optimization. And this thing didn't just fix the problem. It, like, unlocked massive performance gains. The appeal of DLLMs is just enormous, and that's why this technical hurdle really needed to be cleared. Unlike the old AR models, DLLMs, they generate language like a sculptor. They start with a block of noise and iteratively clean it up, refining the whole thing at once. Not a writer putting down one word after another. So they completely bypass that rigid, linear, left-to-right dependency.

1:07What does that actually buy them, you know, in practice? Well, our sources point to some huge advantages. Because they operate on the whole sequence, they have all this potential for handling extremely long contexts. They're also sort of naturally built for multimodal tasks, images, and text together. and they show a lot of promise for faster, more parallel inference, so more flexible, potentially faster, and just a better conceptual fit for really complex problems. Which brings us to refinement, because we all know RL is the gold standard for taking a powerful model and making it genuinely useful, especially for reasoning.

1:43We see this with AR models all the time. RL fine-tuning gives you substantial gains on reasoning tasks, things like math, coding, planning. We absolutely need that for DLLMs. Exactly. But that was the brick wall researchers hit. Trying to apply the standard RL methods that worked so well before, things like GRPO or PTO to DLLMs, it just didn't work. The tool for the old architecture completely broke the new one. Okay, let's unpack that. That fundamental mismatch. For those of us in the AR world, we know RL works by reinforcing choices, token by token. So you can ask, what was the probability of that specific token.

2:19That's the key. Standard RL algorithms, they fundamentally assume a token level factorization. They need to be able to calculate the probability of the whole sequence by multiplying the probabilities of each individual token. And that calculation gives them the importance ratio, which is how they decide how to update the policy, right? Exactly. But the DLLM architecture, it just laughs at that request. It has no idea what you're asking for. None. Because it doesn't generate the next token based on the previous one. It generates the whole sequence at once through these iterative denoising steps.

2:54So that standard conditional probability that RL needs is either ill-defined or hard to compute. The whole concept just doesn't apply. So the algorithm is trying to put a square peg in a round hole. What did researchers try first to kind of bridge that gap? Well, the early attempts had to rely on heuristic approximations, basically workarounds that sacrificed accuracy just to get a number. The first ones used what's called a mean field surrogate. It basically guessed the likelihood of a token without any context from the others. Oh, that sounds like a terrible idea. It was wildly inaccurate. It's like trying to grade an essay by just looking at the word count.

3:32You miss the entire structure, the coherence, everything. Which is the whole point of these hard reasoning problems. If you're writing code or solving a puzzle, every single token is deeply connected to every other token. Precisely. So then came some more sophisticated attempts, things like UniGRPO. They tried to use the model's own math by using the token-level contribution of the ELBO. Okay, let's pause right there. ELBO, Evidence Lower Bound. It's a terrifying acronym if you're not a statistician. In simple terms, what is it? Think of it like this. With these models, calculating the true probability of a sequence is basically impossible, intractable.

4:07The ELBO is a mathematically guaranteed way to calculate a lower bound on that probability. It's the best, most rigorous proxy you can get. Got it. So it's our best guess for the sequence likelihood. Yes. But these previous attempts, they tried to chop that proxy up into token-level pieces. And while that was a bit better aligned with DLLM nature, the core problem was still there. An individual tokens piece of the ELBO, it has no formal probabilistic interpretation. It's only mathematically rigorous for the entire sequence. And that's the big realization. The issue wasn't finding a better token-level proxy.

4:46It was that the token-level decomposition itself fundamentally does not fit for DLLMs. We were caging a holistic model in a fragmented token-level world. And that is the ESPO paradigm shift. Instead of trying to force the model to fit the old algorithm, ESPO adapts the algorithm to respect the model. Its main insight is, you know, actually pretty straightforward. It treats the generation of the entire sequence as a single atomic action. An action. One single action. Instead of asking about token one, then token two, they just look at the final output, the solved Sudoku, the complete block of code.

5:19And that is the single thing that gets judged by the RL reward. So that solves the action space problem immediately. but you still need to calculate the probability of that one big action. That's where the ELBO comes back in, but used correctly this time. Exactly. Since the true log likelihood is still intractable, ESPO uses the sequence level ELBO as its proxy. And by doing this, they're using the ELBO exactly where it's mathematically sound, as a rigorous variational lower bound for the whole thing. It finally validates its use. Okay, so we have the right action space, the whole sequence, and the right proxy, the sequence ELBO.

5:54But the sources say a massive pitfall immediately popped up, an instability challenge. Oh, yeah. It relates directly to the sequence length, which we'll call L. When you calculate that importance ratio, you're looking at the difference between the ELBO of the new policy and the old one. And that raw difference, it scales linearly with L. Wait, so the longer the sequence, the more the math just breaks? It completely breaks. Imagine a sequence of 512 tokens or, you know, 496 tokens. That difference becomes huge. Then you have to exponentiate it for the ratio and you get numbers that are astronomically large or infinitesimally small.

6:30Exploding gradients. The classic nightmare. Textbook. So how did they fix it? Their stabilization technique was mathematically simple but absolutely critical. They just normalized the log ratio by the sequence length, L. They just divide by L? They just divide by L. It transforms that massive, unstable cumulative score into a stable per token scale. It's like turning a team's total points into an average score per player. It keeps the ratios manageable no matter how long the sequence is. That's a genius move. It means the method doesn't just fall apart when you scale up to huge contexts. Right.

7:04But does it actually work in practice? Let's get into the benchmarks. The ablation studies on the Sudoku benchmark tell the whole story. They tested four different approaches. Well, the two mean field approaches, the really simple ones, they just flatlined, learned nothing. And what about the red curve, the token level plus ELBO approach, the one that tried to break the ELBO into pieces? That one, you can see it right in the chart. It suffers from high instability and eventual collapse. It's the empirical proof that breaking the ELBO's integrity just doesn't work. The only one that worked fast, stable learning, highest reward was ESAFO, the blue curve.

7:39So that validates the sequence level decision. But this is where it gets really interesting because they hit another stabilization problem with the KL divergence constraint. Right. The KL constraint is like the safety brake. It stops the new policy from just, you know, forgetting everything it learned during pre-training. But the safety brake was causing the car to crash. Pretty much. Standard RL often uses what's called a K3 estimator for that KL term, but K3 has an exponential term in it, and that brought back the exact same instability problem. The ablations show K3 either stagnates at a low level or just has severe instability and gradient spikes.

8:14So K3 explodes. Why not use K1, the simpler one? Great question. The problem is, in this specific context, the K1 estimator's gradient is zero. Zero. It provides no signal at all. So your safety brake is on the floor, but it's not connected to anything. The policy just drifts off into junk territory. So K3 is too wild. K1 does nothing. They needed a Goldilocks solution. And that's the brilliance of the K2 estimator. It's a simple quadratic function of the ELBO difference. Because it's a polynomial, it avoids that exponential term entirely. No more explosions. It's clean. It's bounded. And it works.

8:50That one choice, moving from the exponential K3 to the quadratic K2, gave them superior stability and consistent gradient norms. It was the final piece of the puzzle for robust training. So a master class in systematically solving one problem after another. When they put it all together, what did it mean for performance? The gains were dramatic, especially on tasks that demand holistic consistency. This is the ultimate proof of their main idea. Let's talk about those planning tasks like countdown and Sudoku. Exactly. On Countdown, which is a numerical reasoning game, ESPO beat the old baselines by 20 to 40 absolute points.

9:26That is a huge margin. A massive jump. But Sudoku, that's where you really see the power. The gains there were up to over 60 points. And that's because a Sudoku grid has to be perfect. Every single number has to obey the global rules of every row, column, and block. All at once. You saw some of the outputs in the source material, right? What was the actual difference? The older methods would get parts right, but they'd often produce these garbled outputs that just violated the rules. They failed the global check. ESBO, because it optimizes the whole grid as one action, was the only method that reliably produced a fully solved, globally consistent grid.

10:03What about on the more standard knowledge-based tasks, math and coding? There, the gains were more modest, but still reliably positive. It consistently beat the other DLL-MRL methods. It shows this isn't just a niche fix for puzzles. It's a generally better framework. OK, we have to talk about cost because all this mathematical rigor sounds expensive. You'd think so, but the good news is it's not. The biggest computational cost in training DLLMs is the generation phase, all those denoising steps. And that cost is fixed. So the extra work EastBO does is in the policy update phase. Right. And that's much less intensive.

10:38Their analysis showed that even when they quadrupled the number of Monte Carlo samples to estimate the ELBO, the total cost, the FLOPs, only went up by about 47%. It's computationally efficient and totally practical. So if we step back, ESPO is a clean, principled, stable, and efficient fix. It's a solution that respects the DLLM architecture instead of fighting it. That's it perfectly. It establishes sequence-level optimization as the principled and empirically effective paradigm for RL in these models. It validates the entire architecture. So the key takeaway for you, the learner, is this. The choice of action space token versus sequence, and the stability of those gradient estimators, especially that simple K2 quadratic function.

11:21These weren't minor details. They were the fundamental decisions that made the difference between total failure and huge success, which leaves us with a final thought. Now that we have a robust sequence-level way to optimize, what does this mean for the future? Think about complex, globally constrained problems. AI agents that need to execute a perfect multi-step plan. How might ESPO's perspective of accelerate the development of truly reliable long-form AI. Something to think about until the next deep dive.

From the publisher

This paper establishes sequence-level optimization as the superior paradigm for fine-tuning diffusion LLMs. They introduce a new machine learning framework called ELBO-based Sequence-level Policy Optimization (ESPO), designed to address the fundamental mismatch when applying Reinforcement Learning (RL) to non-autoregressive diffusion Large Language Models (dLLMs). Traditional RL methods rely on token-level conditional probabilities, which dLLMs lack due to their holistic, non-autoregressive generation process. ESPO resolves this by treating the entire sequence generation as a single action and utilizing the Evidence Lower Bound (ELBO) as a tractable, sequence-level likelihood proxy for optimization. Through comprehensive experiments on tasks like mathematical reasoning and planning, the authors demonstrate that ESPO consistently and significantly outperforms prior token-level RL baselines by enabling stable and principled large-scale training. The results establish sequence-level optimization as the superior paradigm for fine-tuning dLLMs.

More from Best AI papers explained

All 475 episodes
Principled RL for diffusion LLMs emerges from sequence level perspectiveBest AI papers explained · 12 min
Listen in VO