KL-Regularized Reinforcement Learning is designed to Mode Collapse

27 Oct 2025 · 16 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

KL-regularized reinforcement learning (RL) fine-tuning for LLMs can mathematically force mode collapse, reducing output diversity. The episode argues that the KL-divergence direction (forward vs reverse) isn’t the key; instead, the regularization strength beta and reward/reference-model probabilities create two “traps” that collapse modes. It proposes Mode-anchored reward augmentation (MRA/MIRA) to flatten probability across multiple high-quality answers.

Guest backgrounds

No guests are mentioned; it’s a solo “Deep Dive” discussion.

Key claims

(1) Weak regularization (small beta) exponentially amplifies small reward differences, collapsing onto the single best mode. Example: reward difference 2.1 with beta=1e-3 yields ~2.6×10^43 likelihood ratio. (2) If multiple correct answers share equal reward, RL reverts to the base model’s original preference, never boosting initially rare correct solutions.

Notable examples

Uniformly generating 1 vs 2 (standard RL collapses; MRA yields ~50/50). Creative QA with learned reward model: MRA beats GRPO/RLOO on quality and diversity. Drug discovery: replacing RL with MRA in “reInvent” finds more unique high-reward molecules with fewer expensive reward calls.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Mode Collapse in AI

0:45 to 3:50

Exploring mode collapse in AI and its implications for creativity and diversity.

“If you want AI for creative writing or finding, say, new math solutions, or even speeding up scientific discovery...”

KL Divergence's Role in AI Training

3:50 to 6:10

Discussing KL divergence and its impact on diversity in output from AI models.

“And using that, they found these two mathematical traps.”

Mathematical Traps Leading to Mode Collapse

6:10 to 7:40

Examining the mathematical traps that lead to guaranteed mode collapse in AI outputs.

“So if the rewards are equal, the perfectly optimized RL policy defaults back to whatever the original model preferred.”

The Role of Regularization in Model Outputs

7:40 to 9:26

Investigating how regularization strength influences AI output diversity and performance.

“Beta controls the tradeoff between chasing the reward signal, which might favor novelty, and sticking close to the original model.”

Introducing Mode-Anchored Reward Augmentation

9:26 to 11:42

Explaining the proposed solution, Mode-Anchored Reward Augmentation, and its benefits.

“It's essentially reverse engineering that math we talked about.”

Testing MRA Across Different Domains

11:42 to 14:00

Reviewing MRA's performance in various tasks and its effectiveness over traditional methods.

“The MRA-trained models learned to generate 1s and 2s almost perfectly uniformly, 50-50, while still getting the high reward for correctness.”

Understanding KL-Regularized RL as Distribution Matching

14:00 to 14:48

Learn how KL-regularized RL can be reframed as a distribution matching problem.

“That's not the right frame for these flexible models.”

Potential of Targeted Reward Augmentation

14:48 to 15:30

Explore the implications of manipulating rewards to create diverse solution distributions.

“Which leads to a final thought, maybe something provocative for you, the listener, to consider.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Our mission today, we're cutting straight through a stack of research, really focusing on one of the most frustrating issues in modern AI. That's right. We're talking about diversity in large language models, LLMs, particularly when they get fine-tuned using reinforcement learning. Exactly. This RL post-training, it's amazing at making AI hit high-quality targets, right? Maximizing some reward function. Incredibly effective. But yeah, there's a big start effect. it often leads to what researchers call mode collapse. Mode collapse, meaning the AI gets really good, but kind of repetitive.

0:37Its range of answers just shrinks. Precisely. It gets smarter in one sense, but the variety, the diversity of its output just plummets. And why should you, our listener, care? Well, think about it. If you want AI for creative writing or finding, say, new math solutions, or even speeding up scientific discovery... You need exploration. You need a model that can find lots of different good answers, not just the single most obvious one it already knew. Right, not just hammering the most probable answer again and again. And this problem often starts right with the standard RL setup. You're trying to maximize a reward, call it$2, for quality.

1:11Okay. But you also add a penalty to keep the model from changing too much from its original base programming, its reference policy to perform... Let me say key tether. Sort of. And we measure that distance, that drift, using something called KL divergence, and there's this coefficient beta that controls how strong that tether is. And everyone thinks beta is the diversity dial, right? That's the common wisdom. But what this research suggests, and it's fascinating, is that the collapse isn't really a bug in the training. It's not like the AI can't find diverse answers. No. It's often because the mathematical goal we set, the objective function itself, is basically designed to end up with a single mode, especially in typical use cases.

1:52Whoa. Okay. So the math itself is pushing it towards non-diversity. Let's unpack. Yeah, it's a big claim. Okay, so first step. Let's challenge some basic ideas about that KL divergence thing. Yeah. Because, you know, in machine learning classes, there are these standard assumptions. Oh, yeah, the historical playbook. Everyone learns that reverse KL is mode-seeking, finds maybe one or two big peaks in the data, and just zooms in on them. Right. And forward KL. That's supposed to be mass covering. It tries to cover all the areas where the target has some probability. Sounds way more diverse, doesn't it?

2:26It sounds like it should be. So you'd think, OK, want diversity. Use forward KL. Makes sense. But, and this is the kicker, for these huge, super flexible models like modern LLMs, that intuition just falls apart. This research shows it's basically wrong in this context. Really? So the type of KL, forward or reverse, it doesn't actually dictate the diversity outcome for these big models. That's the crucial twist here. It's almost unbelievable. But yeah, when the model is flexible enough, both types can lead to diverse multimodal solutions if the conditions are right. The specific KL type? Turns out it's not the deciding factor.

3:02Okay, hold on. If the choice of KL divergence isn't the key, then what actually controls the diversity, the mode coverage? It comes down primarily to that regularization strength, beta-nodo we mentioned earlier, and how it balances against the rewards, Frenilize, and the probabilities in that original reference model, Pyrrhefic. It's about their relative magnitudes. We've kind of been looking in the wrong place, focusing on the type of KL instead of the strength of the regularization and the rewards landscape. We're blaming the tool, not the settings. Something like that. And this is where it gets really interesting mathematically.

3:35We move past intuition. Okay. The researchers actually analyzed the globally optimal solution. They derived a formula for the log probability ratio between any two possible outputs, let's call them Y dollar and Y dollar two. So a formula that tells you exactly how much more likely one answer is than another. After the RL has done its perfect work. Exactly. A closed form solution. And using that, they found these two mathematical traps. Guaranteed mode collapse, even with perfect optimization. Guaranteed. Okay, let's hear trap number one. Case A, you called it. The weak regularization trap. Right.

4:10This happens when belie is really small. Weak regularization, which, you know, people often do because they want to push hard for high rewards. Makes sense. Lower the tether strength to let the model reach further for better answers. Exactly. But look what happens mathematically. If you take two samples, Y$ and 202, that started off equally likely under the base model, Pyrex. Okay, same starting point. Their final probability ratio gets driven exponentially by how different their rewards are. And it's scaled up massively by$1 beta. Since we know it's tiny,$1 beta is huge. Extonentially. That sounds bad for diversity.

4:46What kind of numbers are we talking about? The paper gives a really stark example. Say the reward difference is tiny, just 2.1, and you use a common small d-ball value like$1 times 10.3. The math forces the sample with a slightly higher reward to become, get this,$2.6 times, 10 times more likely than the other one in the final optimal policy. Wait, 10 to the power of 43? That's astronomical. That's more than atoms in the galaxy territory, isn't it? It's an unbelievable number. It basically guarantees mathematically that the policy collapses onto that single highest reward answer it found. So trying to crank up quality by lowering beta actually ensures you obliterate diversity.

5:24It's not a trade-off. It's annihilation. That's what the math shows for this case. That exponential factor is just brutal. It forces the mode collapse. Wow. Okay. But you said there were two traps. What's case B? The verifiable task trap. Right. Right. This covers scenarios like maybe solving a math problem where there are multiple correct answers. Think different ways to prove a theorem, maybe. Okay. Multiple solutions, all equally valid. And crucially, they all get the exact same reward, say$1 for any correct answer. Makes sense. Correct is correct. So if one ball, one baller equals$1, what does that optimal probability ratio formula tell us now?

6:03Well, the reward terms cancel out completely. And the analysis shows the optimal probability ratio between two equally correct answers is just their original probability ratio back in the base reference model to Perry. Wait a minute. So if the rewards are equal, the perfectly optimized RL policy defaults back to whatever the original model preferred. Exactly. But then what was the point of the RL training if it just ends up where it started in terms of preference between correct answers? Well, if your goal included boosting diversity among correct answers, then yeah, it kind of defeats the purpose.

6:34This proves the RL step, as usually defined, never increases the relative probability of a correct answer that was initially rare compared to a correct answer that was initially common. So it won't learn to value that obscure but correct solution more. Never. Not relative to the common and correct one. And lowering Borelli does absolutely nothing here either. The objective function itself in this equal reward case inherently points back to the base model's biases. It locks in a unimodal preference, typically for the answer the base model already found easiest. Okay, this is really shifting my perspective.

7:06Yeah. So dig a belly, that regularization knob. If it's not really deciding between mode-seeking and mass-covering in these big models, what is its actual job then? It's still a dial and a predictable one, but it's balancing two specific forces. Think of it this way. On one side, you have those potentially high-reward answers that the base model thought were unlikely. High-reward, low-support. The novel ideas. Right. And on the other side, you have the answers the base model already produced easily, which might have slightly lower rewards. Lower reward, high support. The safe, common answers. Exactly.

7:42Beta controls the tradeoff between chasing the reward signal, which might favor novelty, and sticking close to the original model. It's not some magic diversity dial. It's a mathematical balancer between novelty and safety or adherence. That's a much clearer way to put it. Yeah. It's about balancing exploration for reward against staying grounded. And the research showed you can even calculate the exact balance point. Precisely. They showed, with an example figure 2 in their paper, that you can calculate a specific, unique value of beta dollars that will make two different high reward modes end up with the exact same probability in the final optimal solution, even if they started with different probabilities in paraffy.

8:22So you can pinpoint the beta needed to equalize two specific modes. Yeah, in their example it was around beta and called 0.1322. Setting beta to that value made two distinct high reward answers equally likely. It just confirms a beta is a predictable knob for balancing these known forces, not some general diversity adjuster. Okay, so if the standard objective function is mathematically flawed for diversity, either because small reward differences get blown up exponentially, or because equal rewards just revert to the base model's preference. Then we need a better target. We need to explicitly design a target distribution that values all the good answers more equally.

8:59Which brings us to their solution. Mode-anchored reward augmentation. M-I-R-A. M-R-A. Yeah, it's quite elegant, actually. Simple idea, theoretically sound. And the paper emphasizes it's like a two-line code change in practice. Two lines. Okay, that's appealing. How does it work? What does it change? It works by modifying the reward signal itself. Instead of just using the raw reward, it calculates an augmented reward. Let's call it high. Okay, so it's tweaking the reward the AI sees. It's essentially reverse engineering that math we talked about. It forces the final optimal distribution to treat multiple high-quality answers as if they should be equally likely.

9:35Clever. How does it do that calculation? It's a neat mechanism. First, you set a reward threshold. Tout anything scoring above tout counts as good enough. Okay, define good. Second, you pick one specific sample, call it Zell, to be your anchor. This anchor needs to be high-quality reward above tau and also something the original model already produces pretty reliably high probability. So Zell is like the gold standard answer, the best, most typical good answer the model already knows. Exactly. That's your anchor point. Then for every other sample dollar that also meets that high-quality threshold reward above tau, MAR calculates an augmented reward for it.

10:15And this calculation is specifically designed so that in the final optimal policy, this other good sample dollars is mathematically guaranteed to have the exact same probability as the anchor sets. Whoa. So it looks at all the good enough answers and says, OK, you're all getting rewarded in such a way that you end up just as likely as our best, most common good answer. That's the core idea. It forces uniformity amongst all the winners. It directly constructs a target distribution that has this nice, flat, high probability across all the high reward regions. Instead of hoping Teba Daval somehow stretches the probability mass thinly across them.

10:51Right. MRA just builds the desired flat distribution directly by manipulating the rewards going into that optimization formula. It sidesteps the mathematical traps. Okay, the theory sounds solid. Simple fix. Addresses the core mathematical issue. But does it actually work in practice? We need to see results. Agreed. And the paper tests it across three pre-different domains. First, a really simple, verifiable LLM task. What was it? Just training an LLM to generate the number one or the number two uniformly randomly. Both answers are correct. Reward equals 1.0. This directly tests that equal reward trap we discussed.

11:26Okay. And what happened with standard RL? Total collapse. Almost every run ended up generating only ones or only twos, usually one, because apparently the base model had a tiny, almost insignificant initial bias towards generating one. So the base model's tiny preference got locked in by the RL, even though both were equally rewarding. Exactly. But MRA... Did it work? Yeah. MRA successfully maintained diversity. The MRA-trained models learned to generate 1s and 2s almost perfectly uniformly, 50-50, while still getting the high reward for correctness. That's a clear win on the simple case. What about something harder?

12:02Okay, second test. Creative question answering. This is a more complex, non-verifiable task, using a learned reward model for alignment much fuzzier than just correct or incorrect. More realistic then. How did MRA do against standard alignment algorithms? It outperformed them, again using MRA, and they tested both forward and reverse KL versions, showing the KL type didn't matter. It scored higher on multiple metrics compared to baselines like GRPO and RLOO. Higher quality endeavor. Yes. Notably higher out-of-distribution reward, which measures quality on unseen data, and also significantly better on diversity metrics like Enneagram overlap and distinctness.

12:41It worked even with a fuzzy reward signal. Impressive. Okay, third domain, drug discovery. Using chemical language models, right? Why is diversity so crucial there? Because finding promising drug candidates involves evaluating molecules, and that evaluation, the reward function, is often incredibly expensive, like complex simulations or actual lab experiments. Ah, so you don't want the AI proposing slight variations of the same molecule over and over if only one evaluation is needed. Exactly. You want the maximum number of unique high-scoring molecules for the fewest possible expensive evaluations.

13:15Diversity directly translates to efficiency and discovery rate. Makes sense. So how did MARA fit in? They basically dropped MRA into an existing state-of-the-art algorithm called reInvent, replacing its standard RL component. And the result? Pretty compelling. MARA consistently found a higher number of unique, high-reward molecules, what they called yield. And critically, it did this using fewer expensive reward function calls shown by a lower OB100 metric. So better results, less cost. Yeah, essentially boosting the optimization efficiency while still keeping diversity metrics high. Yeah. It found more distinct promising candidates faster.

13:51Okay, three for three. Simple task, complex alignment, real-world scientific application. It seems to hold up. It really does. And I think the big lesson, the thread tying this all together, is that we really have to see this KL-regularized RL as a distribution matching problem fundamentally, meaning we need to let go of those old fuzzy intuitions about mode-seeking versus mass covering based just on the type of KL divergence. That's not the right frame for these flexible models. Instead, we need to be explicit about the target distribution we actually want the AI to learn. Exactly. We need to define it.

14:24Yeah. Because, as we saw, the lack of diversity often isn't some random training glitch. Right. That was the big aha moment for me. It's often a baked-in feature of the optimal solution, dictated by the math when regularization is weak or rewards are equal. An MARA gives us a simple, mathematically grounded way to fix that objective, to construct a target distribution that is explicitly diverse and high quality. Which leads to a final thought, maybe something provocative for you, the listener, to consider. Yeah, the researchers themselves hinted at this. They mentioned the potential for using this kind of targeted reward augmentation to construct an even wider class of desired solution distributions.

15:03So if we can manipulate the rewards so precisely to engineer properties like uniform quality and diversity. What else could we build in? Could we engineer safety constraints directly into the distribution or perhaps enforce certain kinds of predictability or maybe even target specific types of novelty beyond just general diversity? Engineering desirable properties directly into the model's target state through these kinds of mathematical fixes. That definitely gives us something to mull on.

From the publisher

The academic paper investigates the common belief that Kullback-Leibler (KL) regularized reinforcement learning (RL) objectives, particularly when used for post-training large language models (LLMs), inherently promote or inhibit output diversity based on the choice between reverse and forward KL divergence. The authors challenge this intuition, demonstrating both mathematically and empirically that mode coverage and diversity primarily depend on factors like regularization strength and the relative scales of rewards and reference probabilities, rather than the specific type of KL divergence. They prove that typical RL settings often construct an optimal solution that is unimodal by design, leading to an inevitable diversity collapse. To counter this, the paper proposes a new method called Mode Anchored Reward Augmentation (MARA), a theoretically justified algorithm that modifies the reward function to directly optimize for a target distribution that maintains high, uniform probability across all high-quality sampling modes, demonstrating success in LLM and chemical language model tasks.

More from Best AI papers explained

All 475 episodes
KL-Regularized Reinforcement Learning is designed to Mode CollapseBest AI papers explained · 16 min
Listen in VO