In short
How to stably combine reinforcement learning with LLMs by justifying a token-level surrogate objective for sequence-level rewards, and preventing failure modes when the key approximation breaks.
Guests
No guest names or bios are provided in the transcript; it’s a single “Deep Dive” discussion.
Key claims
Directly optimizing sequence-level reward is unstable due to high-variance sequence importance-sampling products (variance and scale). The token-level objective Jada-Peta is a first-order approximation to the true sequence reward only if target and rollout policies stay nearly identical.
Notable examples
Two threats to policy closeness: training-inference discrepancy (different kernels; worse in MoE due to router differences) and policy staleness (off-policy mini-batch updates). Fixes: important sampling correction (removing it collapses even on-policy), clipping (PPO-like leash; prevents drift), and routing replay for MoE (R2 fixes staleness; R3 fixes both discrepancy and staleness). Results: on-policy Mini-RL works best with IS correction; off-policy requires clipping and routing replay or training collapses. Stability predicts final peak performance across base models (e.g., “WN3-Max,” “DSEC-R1,” “GPT-AUS-120B”).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Core Challenge
0:45 to 1:42
Discussion on the challenges of combining RL with LLMs and the reward structure.
“You're trying to figure out the best move for, say, token number five based on a score you won't even see until token 100.”
Theoretical Foundations of Token-Level Objectives
1:42 to 4:36
Exploration of Jada-Peta as a first-order approximation and its implications.
“We're going to break down that central link.”
Factors Threatening Policy Closeness
4:36 to 7:50
Identification of training inference discrepancy and policy staleness as core threats.
“This entire thing is only mathematically valid if those two policies, the one we're optimizing and the one that sampled the data, stay close.”
Strategies for Stabilizing Training
7:50 to 12:00
Discussion of IS correction, clipping, and routing replay as solutions.
“The paper lays out these practical recipes to pull the policies back into alignment.”
The Importance of First-Order Approximation
12:00 to 13:16
Analysis of how maintaining the first-order approximation ensures stability.
“If you omitted either of them, training collapsed.”
The Paradigm-Shifting Conclusion
13:16 to 14:02
Emphasis on the significance of stability in RL processes for model performance.
“It seems like this paper provides a recipe for success that's just rooted completely in theoretical validity.”
Shifting Focus in RL Research
14:02 to 14:32
Learn how the stability of RL processes influences performance and research directions.
“It's saying that the stability of the RL process itself dictates the final performance ceiling.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we're cutting straight through the complexity to get at a really core puzzle in modern AI alignment. how do you stably connect the way language models work with how reinforcement learning works? Right. Because the second you try to combine reinforcement learning-RL with a large language model, you hit this huge wall, this conflict. The massive one. An LLM generates a response, you know, one token at a time. But the reward, the score that tells it good job or bad job, that only shows up at the very, very end. It's a sequence level reward, just one number for the entire output.
0:37But all our standard RL algorithms, things like policy gradients, they're built for token level optimization. It's the ultimate mismatch, really. You're trying to figure out the best move for, say, token number five based on a score you won't even see until token 100. So this isn't just some engineering problem. No, not at all. It's a fundamental question. Is the thing we're optimizing, the token level objective, actually getting us closer to what we care about, which is the final sequence reward? And that mismatch is exactly what this new research from the Quinn team at Alibaba is all about. They tackle it head on.
1:10Their work gives us a really rigorous theoretical reason for why this whole approach can even work. Exactly. They justify using that token level objective, which they call Jada-Peta, as a legitimate first order approximation of the true sequence reward. First order approximation. So that's our foundation. It is. But we have to justify it. And to do that, you have to understand the conditions where it holds up because if those conditions fail, the training doesn't just get a little worse. It fails often violently. Okay. So that's our mission for this deep dive. We're going to break down that central link.
1:44What makes this approximation work? And what are the two critical conditions that threaten it in the real world? And then how do these practical techniques, things like important sampling, clipping, and something really interesting called routing replay, how do they actually work to keep everything stable? Okay, let's unpack this. Let's do it. So let's start with the math. Why can't we just, you know, optimize the true sequence objective directly? Why is that so unstable? It really boils down to two things. Variance and scale. When you try to calculate the gradient for that true sequence objective, you have to use a sequence-level importance sampling weight.
2:21And that's a ratio, right? It's a ratio, yeah. It's the probability of the entire sequence under the policy you're trying to learn, Our target policy divided by the probability of that same sequence under the policy that actually generated it. Our rollout policy. The probability of the entire sequence. The entire sequence. And since a sequence can be hundreds of tokens long, you're calculating that ratio by multiplying the probability ratios for every single token together. Oh, wow. That multiplication must just create chaos. It does. If you multiply 100 numbers, even if they're all close to one, the final product could just zoom off to zero or explode, right?
2:56Exactly right. The numerical range of that product is just immense. And that injects this huge variance into the optimization. And with high variance, you can't get a stable gradient. And large-scale training becomes basically intractable. Okay, so that's why the standard approach, which this paper is validating, is to switch to this surrogate token-level objective instead. Right. Which you said looks a lot like the basic reinforce algorithm. But, and this is key, it has a token-level importance sampling weight. Yes. And the real genius of the first order approximation is how it gets you from that intractable product to a much more manageable sum.
3:35But it all hangs on one very specific assumption. Which is? That our target policy and our rollout policy are nearly identical. They have to be very, very close. Okay. So if the policies are close, then the ratio of their probabilities for any single token is just one plus some tiny little bit of change. A tiny change. Let's call it delta. Right. Delta. Precisely. So if that ratio is$1 plus delta and that delta is genuinely small, then that complex sequence level ratio, the product of all those$1 plus delta terms, can be approximated by something much simpler. Some of the changes. It becomes$1 plus some delta.
4:13You can just ignore all the second order terms like delta times delta because they're so vanishingly small. That feels like a huge mathematical leap. You're swapping this chaotic high-variance product for a clean, low-variance sum. It's a perfect swap. You're replacing that complexity with a stable circuit. But this brings us right to the core vulnerability. This entire thing is only mathematically valid if those two policies, the one we're optimizing and the one that sampled the data, stay close. And if they drift apart. The delta terms get bigger, those second order terms suddenly matter, and your token level objective just completely fails to maximize the sequence reward.
4:52The whole thing breaks down. Okay, this is where it gets really interesting for me. If policy closeness is the absolute linchpin, we have to know what causes them to drift apart in a real production system. And the research breaks this down into two completely independent factors that are always threatening that approximation. Right. The first one is a purely technical, almost numerical headache. It's called the training inference discrepancy. This is the difference between the policy as it's calculated during training versus how it's calculated during inference, even when you are using the exact same saved model parameters.
5:29Wait a minute. Are you saying that if I save a model and I run it on my training hardware and then I move that exact same file over to my fast serving hardware, the outputs might not be the same? That's exactly what I'm saying. That sounds like a catastrophic engineering problem. That's non-deterministic. It is non-deterministic and it's a fundamental issue in large scale systems. The engines you use for training like a Megatron or an FSDP and the engines you use for super fast inference like SGLang or VLLM, they use different computational kernels. They're optimized for different things. And those different kernels can lead to tiny numerical inconsistencies.
6:04Tiny but meaningful ones. The output probabilities can be different, even with the exact same input. And if the probabilities are different, well, the policies are already divergent before you've even done a single parameter update. Wow. And this gets way, way worse in mixture of experts' models, MOE models. Because of the router. Yes, the expert router. If that mechanism that decides which expert handles which token behaves differently between the two environments, the whole sequence can just diverge instantly. It completely shatters the assumption of policy closeness. Okay, so that's threat number one.
6:37Systems just not being numerically consistent. What's the second factor? You said it's more chronological. That's right. It's policy staleness. And this is a classic RL problem. The data you sampled with your old policy is now being used to train your current constantly changing policy. The data is stale. The data is stale. And this happens because we're always pushing for more computational efficiency. Generating responses takes a fixed amount of time. So to speed things up, we do off policy updates. Meaning you sample one huge batch of responses. And then you split it into smaller mini batches to do several gradient updates.
7:14Oh, okay. So the model updates its parameters after the first mini-batch. That moves the current policy away from the original one. Exactly. So when the model gets the second mini-batch from that same sample, the data is already out of date relative to the new parameters. Precisely. The responses you use later in the training step are optimizing a policy that has drifted further and further from the one that actually generated the data. This is the second huge factor we have to control. This is fascinating. So you have these two forces, numerical discrepancy and chronological staleness, both working against the core theory.
7:50So now let's talk solutions. The paper lays out these practical recipes to pull the policies back into alignment. We have to start with the important sampling correction, the IS correction. We already said this weight is what allows the token level objective to even exist as an approximation. What the research confirms is that this token level IS weight is the specific correction mechanism for that numerical training inference discrepancy. So it's not just part of the theory. It's the active fix. It is the fix. And the experimental evidence for this is just staggering. When they tried to train a model without the IS correction, even in a simple on-policy setting, the training just collapsed.
8:28Immediately. Rapidly and violently. It just confirms that the second you remove that mathematical guardrail, the numerical differences take over, the approximation shatters, and the model just drives itself straight into instability. Okay, so IS correction is mandatory. It handles the system-level inconsistency. But what about policy staleness, especially when we're intentionally making things worse with off-policy training? That's where the clipping mechanism comes in. This is used in baselines like Mini-RL, and it's inspired by PPO. Clipping directly attacks policy staleness. How? It stops those aggressive parameter updates from letting the target policy get too far away from the rollout policy.
9:08It puts a leash on it. A hard boundary, yeah. It's constantly watching the ratio, EALC, between the new policy and the old one for every single token. If that ratio gets too big, say above 1.2 or too small below 0.8. It cuts it off. The gradient for that token is immediately clipped, halted. This makes sure that even after a bunch of mini-batch updates, the model's behavior hasn't changed so much that the original data is useless. It preserves the first-order approximation. That's a really clean solution for general staleness. But you mentioned Moe models are a special, complicated case. They are.
9:43Because their expert routing gets tangled up with both the numerical discrepancy and the policy staleness. So how do you stabilize that extra layer of chaos? For that, you need routing replay, either R2 or R3. This solution basically forces the MoE model to use a fixed set of experts during the optimization step. It just bypasses all the non-determinism and the staleness from the routing layer. So you're basically recording which experts were used during the rollout. And forcing the optimizer to replay that exact decision. Exactly. And there are two versions, R2 and R3. Two key variations. Vanilla routing replay, R2, mostly deals with policy staleness.
10:22It fixes the experts based on what the rollout policy decided in the training engine. It keeps the pathway consistent as the parameters update. And R3. Rollout routing replay, R3, is the more robust one. R3 reduces both staleness and the training inference discrepancy. It does this by fixing the experts based on what the policy decided in the inference engine, the actual model that generated the data out in the wild. The ground truth, basically. It's the stronger mechanism for when things get really unstable. All right, let's get to the results. When are these techniques actually necessary? Let's start with the simplest case.
10:58On policy training, where staleness is low. In that low staleness world, the basic policy gradient algorithm, Mini-RL, did the best. But, and this is the key, only if it included the important sampling correction. The conclusion is? Pretty simple. When your data is fresh, just protecting the mathematical validity of the approximation is all you need to do. And what about other tricks? People use heuristics all the time. Right. And they tested that. They found that common heuristics like length normalization, where you divide the reward by how long the sequence is, that actually led to worse performance.
11:29Right. Because length normalization subtly breaks the math of the first order approximation. It's a great example of how being mathematically rigorous here is way more valuable than just applying a common sense heuristic. Don't mess with the approximation. Do not mess with it. Okay, so now what happens when we go off policy? When we crank up the staleness to get faster training by using mini-batches? As soon as you introduce off policy updates, both routing replay and clipping stop being optional. They become absolutely essential for stability. Not just helpful, essential. The experiments were crystal clear.
12:04If you omitted either of them, training collapsed. Prematurely and every time, you need clipping to limit the policy's movement, and you need routing replay to stabilize the MOE architecture. Both are required to keep the policies close. And they found a difference between R2 and R3 depending on how off-policy they went. Right. Tell us about that. Yeah. The degree of off-policiness, which they call nondollars, was the deciding factor. That's the number of gradient updates you do for each big batch of samples. Okay. If$1 was small, like just two updates, then R2, the less aggressive fix, actually did better than R3.
12:38It just alters the target a little less. But when you get really aggressive. When none dollars was large, say 1N4 or even N81, that's high policy staleness. And there, R3 completely surpassed R2. In fact, R2 just failed to keep training stable at all. The thinking is, it's like a car drifting. If you're only drifting a little bit, a small steering correction works best. That's R2. But if you're seriously veering off a cliff-high dollars, you need that aggressive full-throttle correction. That's R3. Even if it introduces a tiny bit of its own bias, it's better than crashing. So preserving that first-order approximation is the number one priority no matter what.
13:16No matter what. So what's the big takeaway here? It seems like this paper provides a recipe for success that's just rooted completely in theoretical validity. I think so. The success of RL for LLMs seems to hinge entirely on maintaining the delicate balance of that first-order approximation. All these techniques we talked about, IS correction, clipping, routing replay, they're just tools to keep the two policies close enough for the math to work. Absolutely. And that leads to what I think is the most compelling, almost paradigm-shifting conclusion from their research. Which is? Stability is decisive.
13:47That's it. They tested this on multiple big base models, WN3-Max, DSEC-R1, GPT-AUS-120B. And the results showed that once the RL training was stable and you let it run long enough, the models consistently converged to similar peak performance. It didn't matter where they started. That is profound. It's saying that the stability of the RL process itself dictates the final performance ceiling. It seems to, yes. So it shifts the whole focus. We spend so much time obsessing over, you know, cold start initialization, the exact data in the pre-training or the first alignment steps. But if stable, long-running RL training consistently gets you to the same place, regardless of the initial base model, should the focus of future LOM research maybe shift away from the quality of the base model and entirely towards just perfecting these stable RL recipes?
14:35Something to explore on your own. Thank you for diving deep with us. We'll see you next time.
From the publisher
The research paper proposes a novel formulation for applying reinforcement learning (RL) to large language models (LLMs), specifically focusing on how a **sequence-level reward** can be optimized using a **surrogate token-level objective** in policy gradient methods. The authors theoretically justify this approximation, showing its validity relies on minimizing the **training-inference discrepancy** and **policy staleness**. Extensive experiments, conducted with a 30B Mixture-of-Experts (MoE) model named Qwen, empirically validate that techniques such as **importance sampling correction**, **clipping**, and particularly **Routing Replay** are crucial for achieving **stable RL training**. The findings suggest that stable training is a more decisive factor than cold-start initialization for achieving comparable final performance across different training setups.




