RL's Razor: Why Online RL Forgets Less

7 Sep 2025 · 25 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Catastrophic forgetting in continual learning, and a MIT paper’s claim that “RL’s razor” explains why online reinforcement learning (RL) forgets less than supervised fine-tuning (SFT).

Key claims

Forgetting is predicted by the KL divergence between the fine-tuned and base policy on the new task; RL achieves similar new-task performance while preserving prior skills better than SFT because RL is on-policy and tends to make smaller conceptual shifts.

Notable examples

LLM tests on math/scientific Q&A, tool use, and benchmarks like Hellaswag and MMLU; robotics tests on OpenVLA manipulation (e.g., pick-and-place, drawer opening/closing). Controlled environment ParityMist (multiple-correct answers) shows strong KL–forgetting correlation (R²≈0.96). Oracle SFT (data chosen to minimize KL) can outperform RL on the forgetting/learning tradeoff.

Guests

No specific guest names or backgrounds mentioned in the transcript (only paper authors: Aydin Shenfeld, Jyotish Pari, Pulkit Agrawal, MIT).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenge of Continuous Learning

1:54 to 2:52

Understand the limitations of current AI models in retaining prior knowledge while learning new skills.

“We all marvel at foundation models these days.”

Supervised Fine-Tuning vs. Reinforcement Learning

2:52 to 4:00

Learn the differences between supervised fine-tuning and reinforcement learning in AI training.

“And while making models larger or giving them more pre-training data helps somewhat, this research clearly shows it doesn't solve catastrophic forgetting entirely.”

Empirical Findings on Learning Methods

4:00 to 6:00

Discover how reinforcement learning preserves prior knowledge better than supervised fine-tuning.

“Just to give you a quick mental picture, SFT is a bit like teaching by example.”

Understanding Catastrophic Forgetting

6:00 to 8:00

Explore the concept of catastrophic forgetting and its implications for AI development.

“And also sophisticated robotic platforms like OpenVLA, performing precise manipulation skills, a pick-and-place task, for example.”

KL Divergence and Its Role

8:00 to 10:00

Learn about KL divergence and its predictive power regarding catastrophic forgetting in AI models.

“Well, it's a remarkably consistent finding.”

Oracle SFT Experiment Insights

10:00 to 12:00

Examine the findings from the Oracle SFT experiment and its implications for minimizing forgetting.

“Wow, that's almost a perfect prediction.”

Mechanisms Behind RL's Effectiveness

12:00 to 14:03

Delve into the mechanisms that make reinforcement learning effective in minimizing forgetting.

“Why does RL naturally maintain a smaller KL divergence compared to SFD?”

Exploring On-Policy Learning

14:03 to 14:36

Learn about the features of on-policy learning vs. SFT.

“Well, first, it's learning from its own experience, sampling from its current understanding, its own policy, rather than from prepackaged external examples like SFT does.”

Importance of On-Policy Sampling

14:36 to 16:46

Discuss the significance of on-policy data generation in reducing forgetting.

“They rigorously tested this by comparing different learning objectives, different algorithms like GRPO, 1.0 Reinforce, SFT, and Simpo, which vary in how they use rewards and sample data.”

Factors Affecting Forgetting

16:46 to 19:18

Examine other factors that influence forgetting in reinforcement learning.

“So it's always making the minimum necessary change in a sense.”
Show all 12 chapters

Scaling and Model Size

19:18 to 21:44

Discover the relationship between model size and forgetting in AI systems.

“While some of those showed some signal, none even came close to the predictive power and consistency of the standard forward KL divergence they identified.”

KL Divergence and AI Design

21:44 to 24:43

Understand how KL divergence can inform better AI learning strategies.

“I think this research doesn't say that as much as it highlights a fundamental principle.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine a future where AI isn't just intelligent, but truly a lifelong companion. Picture intelligent agents that don't just exist but continuously learn, adapt and grow with us day in and day out. That's the dream, right? Absolutely. But standing squarely in the way of that vision is a massive hurdle AI researchers call catastrophic forgetting. Like an AI learns a fantastic new skill only to immediately wipe out a crucial old one. You know, one step forward, maybe two steps back. Yeah, fundamentally limits how these systems can evolve over time. So today, our deep dive is all about tackling this fundamental problem.

0:38We're going to unpack some, well, genuinely groundbreaking research. It really is. It reveals why some AI models forget less than others when they're acquiring new tasks. We're looking at a fascinating paper from MIT that proposes this new principle called RL's razor. RL's razor, yeah. It's a clever name. It really seems like a game changer for how we think about AI that actually remembers. Exactly. And the central challenge here, it's fundamental to the very future of AI, isn't it? How do we enable today's powerful foundation models? You know, those incredibly sophisticated, large language models and cutting-edge robotic systems.

1:13The ones everyone's talking about. Right. How do we get them to acquire new skills without inadvertently eroding the vast amount of knowledge they've already gained? It's about designing AIs that truly evolve with us rather than having this constantly shifting, unreliable memory. And the key source for unlocking this crucial insight comes from a paper titled, RL's Razor, Why Online Reinforcement Learning Forgets Less. by Aydin Shenfeld, Jyotish Pari, and Pulkit Agrawal. A really solid piece of work from MIT. Definitely. And our goal for you today is to give you a direct shortcut, sort of cut through the academic specifics to grasp this vital understanding in AI's journey toward, well, continual learning.

1:53Let's get into it. Okay, let's unpack this. We all marvel at foundation models these days. I mean, whether they're crafting eloquent text, recognizing objects in images, or orchestrating complex robot movements. Their capabilities are pretty incredible, yeah. For sure. But as this paper highlights, for all their power, most of these models are largely static once they're put to use. They excel at what they were initially trained on, but they're not inherently designed to continuously learn and self-improve, adding new capabilities indefinitely without wiping out the old ones. And that brings us right to the heart of the problem.

2:29Catastrophic forgetting. Right. Simply put, it's the tendency for an AI model to lose its previously acquired capabilities when it's trained on new tasks. Think of it like this. An advanced language model that just learned a new programming language might suddenly struggle with basic math. Oh, wow. Or a robot learning a complex new manipulation skill, like assembling some specific component, might suddenly forget how to simply grasp a common object it previously handled with ease. That seems counterproductive. Totally. And while making models larger or giving them more pre-training data helps somewhat, this research clearly shows it doesn't solve catastrophic forgetting entirely.

3:08It remains a persistent obstacle to building what we envision as truly long-lived agents. You mean AIs that can adapt and grow with us for years? Exactly. Without needing constant retraining from scratch or losing vital skills along the way. So the urgency here is definitely palpable. I mean, for anyone involved in building or deploying these models, this is a major roadblock. It really is. To achieve that future of continuously adapting AI, we absolutely need methods that allow these sophisticated systems to acquire new skills without the constant risk of erasing all their old ones. It's a critical design challenge if we want AI assistants that, you know, truly endure.

3:47Couldn't agree more. Now, here's where it gets really interesting. The researchers dove deep into two of the primary learning approaches used to update these powerful models after their initial training. Right, the post-training phase. Exactly. Supervised fine-tuning, or SFT, and reinforcement learning, or RL. The two big ones. Just to give you a quick mental picture, SFT is a bit like teaching by example. You show the model specific labeled inputs and their desired outputs. Like flashcards, you said earlier. Yeah, exactly. It learns from that fixed external guidance. RL, on the other hand, is more about learning through experience.

4:23Right, learning by doing. The model interacts with an environment, tries things out, gets feedback rewards or penalties, and gradually optimizes its behavior based on that trial and error process. And the core empirical finding of this paper, well, it reveals a truly striking disparity between these two methods. It's quite surprising, actually. What did they find? What they observed was that even when SFT and RL achieved the exact same level of proficiency on a brand new task. RL consistently preserved the model's prior knowledge and capabilities significantly better. It just held onto its old skills with remarkable tenacity.

4:57Wow. So same new skill performance, but RL just remembers better. Way better. To help you visualize this, they have this great figure in the paper, figure one. Imagine a chart mapping out the trade-off. On one axis, how well the model performs on the new task it's learning. On the other axis, how well it still performs on prior tasks. its existing knowledge. Right, the forgetting axis, essentially. Exactly. And with SFT, as the model gets better at the new task, its performance on those prior tasks often drops sharply. It's like a zero-sum game, sacrificing old knowledge for new gains. You learn Python, you forget French.

5:34Kind of. Sort of, yeah. But with RL, that curve for prior tasks remains remarkably high, showing minimal forgetting, even as it reaches similar levels of performance on the new task. It's a stark, almost counterintuitive contrast when you see it plotted out. It really is. And they didn't just see this once. They validated it extensively across a really wide array of cutting-edge AI systems. Like what? Well, they tested advanced large language models on complex reasoning tasks like math and scientific Q &A, even tool use. And also sophisticated robotic platforms like OpenVLA, performing precise manipulation skills, a pick-and-place task, for example.

6:12So language and robotics covering the bases. Absolutely. And the measurements were incredibly thorough. They gauged new task success, of course, but also looked at a broad spectrum of general capabilities. For LLMs, things like Hellaswag, MMLU. Tendered benchmarks. Exactly. And for robotics, prior skills like opening and closing drawers. Crucially, they pushed the models hard, exploring lots of different settings, hyperparameters to find the best possible tradeoffs. The Pareto frontiers, as they call them. And RL always came out ahead on that tradeoff. Consistently. RL offered a far more favorable balance, what they call a better Pareto frontier, meaning it could master new skills without nearly as much forgetting compared to SFT.

6:53Okay, so the first big takeaway here, takeaway one from the paper, is crystal clear. RL is able to learn new tasks while incurring minimal forgetting, whereas SFT reaches similar new task performance, often only by sacrificing prior knowledge. That's the key distinction. This is a crucial distinction. It tells you that if you're building an AI that needs to keep its wits about it, the method of learning matters a whole lot, maybe more than just the raw performance on the new task itself. Right. It's not just if it learns, but how it learns. RL, in its very nature, seems to be a master of learning without losing its mind, unlike SFT, which often demands that costly trade-off.

7:32This isn't just theory. It's a fundamental design choice for building robust, continually learning AI. Absolutely. Which naturally leads to the next big question. Why? Right. This intriguing observation immediately sparks that more fundamental question. What's the root cause of this difference? What truly governs whether an AI holds on to its knowledge no matter how it's being trained? And this is where they found something really elegant. The researchers uncovered what they've termed an empirical forgetting law. A law? Sounds serious. Well, it's a remarkably consistent finding. They found that the degree of catastrophic forgetting is accurately predicted by the KL divergence between the fine-tuned and base policy evaluated on the new task.

8:14Okay, KL divergence. Break that down for us. Sounds a bit technical. It does, but the concept is actually quite intuitive. Think of KL divergence simply as a precise way to measure how much one probability distribution differs from another. In this context, how the model behaves. Exactly. It measures how much the model's new way of thinking or new behavior on the new task has shifted or diverged from its original way of thinking or behavior on that very same task before the fine-tuning started. So it's like a measure of how much the model changed its mind about the new task. That's a great way to put it.

8:50It's like a speedometer for conceptual change specifically related to the new thing it's learning. That's incredible. So this forgetting law isn't just some abstract theory. It's something we can actually measure and influence during training. Precisely. That's the practical utility here. That feels like a massive simplification for developers. You don't need to hold on to every piece of old data, maybe. You just need to monitor this one metric. Is that right? That's the idea. You can track this metric and even design your fine-tuning process to actively keep it low, all without needing to access or constantly replay vast amounts of potentially unavailable past-task data.

9:24Okay, how did they test that? Right, so to rigorously test this, they used a controlled, simplified environment called ParityMist. It's clever. ParityMist. Like MNist digits. Yeah, based on MNES digits, but designed to mimic real-world generative tasks where multiple correct solutions might exist. So, for example, instead of just classifying a 2 as the digit 2, the task might be to classify it as an even number by outputting any even number. Ah, I see. Multiple right answers. Exactly. And in this controlled setting, they observed a remarkably strong and consistent relationship between this KL divergence metric and the amount of forgetting they saw.

10:02The R-squared value was like 0.96. Wow, that's almost a perfect prediction. It's incredibly strong. It really solidified their hypothesis with clear statistical evidence, referencing Table 1 and Figure 3 in the paper. But what's genuinely fascinating here, I thought, was what happened with their Oracle SFT experiment. This blew my mind a bit. Oh yeah, the Oracle SFT, that was neat. The researchers literally constructed an SFT training data set that was provably designed to minimize this KL divergence. They basically cheated and gave SFT the perfect KL minimal answers. Right. They knew the answer that was correct for the new task, but also closest in KL space to the original model's prediction.

10:44And when SFT was trained with this perfectly optimized Oracle data set, what happened? It actually forgot even less than RL. It achieved the absolute best tradeoff between new learning and old knowledge retention they observed in any experiment. So SFT can avoid forgetting if you guide it perfectly to minimize KL. It absolutely proved the point. This showed that RL's advantage isn't some magical property exclusive to reinforcement learning itself. Instead, it's because RL implicitly, just by how it works, biases itself towards solutions that happen to minimize KL divergence. The core insight is this.

11:21If any training method is biased towards solutions that involve smaller conceptual shifts from the original model. Then forgetting is significantly reduced. Exactly. It decouples the method RL recess SFT from the principal KL minimization. Okay, this leads us to a crucial understanding, the second big takeaway. Catastrophic forgetting, whether you're using SFT or RL, seems to be accurately predicted by how much the fine-tuned model's thinking shifts from its original state on the new task, that KL divergence metric. It's like a universal barometer for knowledge retention during fine-tuning. A universal barometer.

11:57I like that. So if KL divergence is this universal predictor of forgetting, then the natural next question is why? Why does RL naturally maintain a smaller KL divergence compared to SFD? Right. What's the mechanism? Exactly. And this is where the paper's central concept, RL's razor, truly comes into play. Right. So RL's razor, as they define it, states that among the many possible ways to achieve high reward or good performance on a new task. The many possible solutions. Yes. On policy methods, like the kind of RL they studied, are inherently biased towards solutions that involve smaller conceptual shifts, smaller KL divergence from the original policy.

12:34Okay, so it prefers solutions that are closer to what it already knew. Think of it like a cautious explorer, maybe. If there are many paths leading to a treasure chest full of reward, the RL explorer naturally tends to choose the path that deviates the least from where they are right now, rather than, say, teleporting to a completely different, maybe equally good location that's much further away conceptually. That's a helpful analogy. So why does RL behave like that, cautious explorer? To understand why, let's look again at the fundamental differences in how SFT and RL learn. SFT primarily learns by minimizing the difference between its outputs and a fixed external supervision data set.

13:12The flashcards again. The flashcards. The correct answer is already provided externally, and the model simply tries to mimic it, pulling its distribution towards that external target. Okay. RL, particularly with the policy gradient methods they focused on, optimizes differently. It samples outputs from the model's own current distribution. This is the crucial on-policy aspect. Right, it learns from its own attempts. Exactly. It then refines those outputs based on the direct rewards or penalties it receives from the environment for those attempts. It's more like a child learning by trying things out, seeing what works, getting feedback, and gradually refining their own approach based on that real-time, self-generated experience.

13:55That distinction makes a lot of sense. So if I'm getting this right, it boils down to maybe two key characteristics of RL's training process. What are you thinking? Well, first, it's learning from its own experience, sampling from its current understanding, its own policy, rather than from prepackaged external examples like SFT does. Yes, the on-policy sampling is key. That's feature number one. And second, maybe it's the ability to learn from both positive and negative feedback. Yeah. It can actively learn what not to do, pushing probability away from poor outputs, a mechanism that's often absent in standard SFT, which just focuses on matching the correct answer.

14:30That's feature number two they highlight. Now, they actually tested the relative importance of these. Oh, interesting. What did they find? They rigorously tested this by comparing different learning objectives, different algorithms like GRPO, 1.0 Reinforce, SFT, and Simpo, which vary in how they use rewards and sample data. You can see this in figure four. And the critical factor that distinguished the low KL methods from the high KL methods wasn't just the presence of negative gradients or reward shaping. It was primarily the on-policy nature of the data generation itself. So sampling from the model's own current policy was the main thing keeping KL low.

15:08Exactly. Methods that did that, like GRPO and 1-0 Reinforce in their tests, consistently led to significantly smaller KL divergence and thus less forgetting. Okay, let's try that other analogy you mentioned for the paper, the warped bull. How does that fit in? Right, the warped bull image. Imagine all the optimal ways to perfectly solve a new past. All the best possible policies are represented by the lowest points scattered across the bottom of this uneven warped bull. Okay, multiple low spots. Yes. RL being on policy is like that cautious explorer again. It takes small iterative steps downhill from its current position on the bowl surface, always staying relatively close to its recent path as it seeks out a nearby low point.

15:51It finds a local optimum, maybe? Well, it finds an optimum, but it finds one that's reachable via small steps from where it started. SFT, however, because it's being pulled towards an external target distribution, the flashcards, might just jump across the bowl to a completely different low point. Even if that point is conceptually very far away from where it started. Exactly. It finds an equally low point, maybe even the globally lowest point, but it might get there by making a huge leap, potentially discarding all the terrain the prior knowledge it traversed before. It doesn't inherently prefer the nearest good solution.

16:24Got it. That makes the difference very clear. And from a more theoretical perspective, policy gradient methods can be understood as performing a kind of conservative projection. At each update step, they gently nudge the policy towards a nearby reward weighted version of itself, rather than yanking it directly toward a potentially distant external distribution as SFT can do. So it's always making the minimum necessary change in a sense. It's essentially performing something akin to a minimum KL projection onto the set of optimal policies at each step. It finds the best policy that's closest in KL terms to its current self.

17:01This mechanism, then, gives us our third big takeaway, maybe the core of RL's razor. Lay it out. On policy training fundamentally explains why RL maintains smaller KL divergence than standard SFT. Sampling from the model's own distribution inherently keeps it anchored, close to its base knowledge. while SFT, by relying on external, potentially arbitrary target distributions, can push it much further away conceptually, leading to more significant forgetting. That nails it. The on-policy nature is the key mechanistic explanation for RL's lower forgetting. Now, to be absolutely thorough, the researchers didn't just find this answer.

17:38They also meticulously ruled out a whole bunch of other potential explanations for why RL forgets less. What did they find? Wasn't the main cause. Right. They were very systematic about this. They looked at several common hypotheses people have had about forgetting. You can see the comparisons in Table 1 looking at the R-squared values. Like what kind of things? Well, for instance, changes at the level of the model's weights or parameters. How much did the numbers physically change? They measured things like L1 norm changes, Fisher-weighted changes, spectral norm. But these only showed weak correlation with forgetting.

18:14So it's not just about how much the weights wiggle. Doesn't seem like it, no. They also looked at changes in the model's internal representations. How did the patterns of activation inside the network shift? Again, using measures like L1 or L2 distance between activations, they found these didn't consistently align with forgetting either. Hmm. But surely the representations do change. Oh, they do change. Figure 7 shows there is representational drift. Especially with SFT, its internal similarity, measured by CKNNA, dropped quite a bit, down to like 0.56. But RL models actually retained high representational similarity around 0.94.

18:51So while drift happens, it wasn't the predictor of forgetting in the way KL divergence was. Okay, what else did they rule out? Things like the sparsity or mathematical rank of the updates. They found some observed sparsity was likely just an artifact of using B float 16 precision, not a fundamental property of RL causing less forgetting. Gotcha, technical detail. Right. And they also looked at other ways to measure the distance between the model's behavior distributions, like reverse KL or total variation distance. While some of those showed some signal, none even came close to the predictive power and consistency of the standard forward KL divergence they identified.

19:28Wow. So they really stress tested this KL divergence idea against all the usual suspects. They really did. It makes the finding much more robust. So what does this all mean for us, for people building and using these AI systems? It means this KL divergence on the new task is pretty unequivocally the most reliable compass we have right now for predicting and potentially controlling catastrophic forgetting. Absolutely. And this has profound implications for the future of AI development. For example, the paper looked at scaling. Does making models bigger solve this? Well, they tested models up to 14 billion parameters using the Quinn 2.5 models on Science Q &A, shown in Figure 8.

20:05And even these much larger models still exhibited the same fundamental SFT trade-off. Bigger models still forget significantly if trained with standard SFT. So size alone isn't the silver bullet for continual learning. Doesn't look like it. The method still matters critically. They also analyzed the optimization dynamics at a really granular level, step-by-step during training. Figure 10 shows this. And they found a strong correlation even there. Individual update steps that caused larger shifts in KL divergence were strongly aligned with gradients that actually increased forgetting of prior tasks.

20:39This shows the link holds true even for the tiniest learning adjustments. That's fascinating. It's happening at the micro level, too. It is. And ultimately, this whole principle, RL's razor and the KL connection, opens up a completely new design axis for future AI research, especially in post-training or continual learning. How so? Well, it suggests that algorithms should aim to not just optimize new tasks well, achieve high reward, but also to move conservatively in terms of KL divergence relative to the base model they started from. So efficiency and conservation. Exactly. It suggests a powerful new paradigm for how we should approach fine-tuning and adaptation in these powerful foundation models.

21:19That's a powerful principle. But just to play devil's advocate for a second, SFT is often simpler, right? Faster to implement, maybe uses less compute sometimes. That's an excellent point. And it's true. SFT definitely has its advantages in terms of data efficiency, labeling cost, and sometimes ease of implementation. So does this finding mean we should just ditch SFT and always lean towards RL for continual learning scenarios? Not necessarily ditch SFT, no. I think this research doesn't say that as much as it highlights a fundamental principle. The key is the KL divergence. Ah, okay. If we can adapt SFT methods, perhaps by carefully selecting the training data, like in the Oracle experiment, or develop new hybrid methods, methods that also implicitly or explicitly encourage minimizing KL divergence during learning, then we might actually get the best of both worlds.

22:11The speed and simplicity of SFT, maybe, with the memory retention of RL. That would be the ideal goal, right? Right. So the takeaway isn't just about picking RL over SFT blindly, but about understanding the underlying dynamics, the scale divergence mechanism that governs forgetting gives us a new lens, a new target for designing smarter learning algorithms, whatever form they take. Got it. So let's try and sum up. We've taken a really deep dive today into RL's razor. We have. And we've discovered that the core idea, the key to avoiding catastrophic forgetting, isn't just what an AI learns or how well it learns the new thing, but how conservatively its learning path deviates from its existing knowledge.

22:49Right. It's about keeping its conceptual shifts measured and sort of close to home, conceptually speaking. Exactly. And this concept of KL divergence as the universal predictor combined with RL's inherent bias towards these minimal changes, well, it's a truly transformative insight, I think. It feels like it. It strongly suggests that to build AI that truly learns for life, the kind we imagined at the start, we need to consciously design training methods that explicitly guide that learning along what we're now calling the KL minimal path. Keep the change small. And this isn't just some abstract academic finding floating out there.

23:26It could genuinely pave the way for foundation models that can truly adapt, evolve, and serve us over the long term without constantly having to relearn or worse, forget what they already know. It's really about designing AI with enduring memory, not just fleeting intelligence. Well said. It's about stability and growth together. Precisely. And maybe here's a final provocative thought for you, for everyone listening to consider. While we now have this really strong evidence that larger KL shifts predict forgetting, we still lack a full sort of deep mechanistic account of why exactly they so effectively disrupt prior knowledge within the complex dynamics of the neural network?

24:08What fundamental processes are truly at play at the neuron level when this divergence occurs? Hmm, the why behind the what. Exactly. And perhaps even more importantly, for future progress, how can we strategically combine the proven efficiency and data simplicity of SFT with the knowledge preservation properties linked to low-KL divergence, maybe inspired by RL? How do we engineer the ultimate continual learning algorithm, than one that is both fast and flexible to adapt, but also incredibly resilient against forgetting its past. That's the million-dollar question, isn't it? Finding that perfect balance.

24:41It really is. It's the next frontier. Yeah. Well, thank you for joining us on this deep dive into the fascinating world of AI's memory and learning and this concept of RL's razor. We hope this has given you a fresh perspective on how intelligent systems might, one day, truly learn for life.

From the publisher

This paper explores why **Reinforcement Learning (RL) fine-tuning leads to less catastrophic forgetting** in models compared to **Supervised Fine-Tuning (SFT)**, even when both achieve similar performance on new tasks. The authors introduce **"RL's Razor,"** a principle stating that **RL is implicitly biased towards solutions that cause minimal change (KL divergence) from the original model's policy** when learning new tasks. Empirical and theoretical evidence supports this, demonstrating that **KL divergence on the new task is a strong predictor of forgetting**, regardless of the training algorithm. The core reason for RL's advantage is its **on-policy training**, which samples from the model's current distribution and reweights those samples, leading to more conservative and KL-minimal updates compared to SFT's reliance on fixed external annotations.

More from Best AI papers explained

All 475 episodes
RL's Razor: Why Online RL Forgets LessBest AI papers explained · 25 min
Listen in VO