Self-distillation enables continual learning

7 Feb 2026 · 20 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Self-distillation fine-tuning (SDFT) for continual AI learning that avoids catastrophic forgetting by using the model itself as both “teacher” and “student,” turning in-context learning into on-policy updates.

Guest backgrounds

No guests are named; the episode is a two-host discussion.

Key claims

Standard supervised fine-tuning (SFT) is off-policy and causes a “seesaw” (skills learned later overwrite earlier ones). SDFT creates a split personality: a temporarily “cheat-sheet” teacher guides a blind student; the teacher output acts as a scoreboard via reverse KL (matching the teacher’s probability/confidence). Because the student generates answers itself, skills/facts integrate without deleting prior knowledge, and reasoning can deepen even when only final answers are provided.

Notable examples

Sequentially learning tool use, science Q&A, and medical diagnosis—SFT oscillates and drops earlier skills; SDFT keeps all. Injecting 2025 disaster facts (Myanmar earthquake, Typhoon Kalmagi, Hurricane Melissa): SFT ~80% strict accuracy vs SDFT ~89%, near oracle retrieval-augmented generation. Indirect questions about which countries needed aid: SDFT succeeds where SFT fails. Response length: SDFT ~4180 tokens vs SFT ~3200. Scaling: SDFT works poorly at 3B parameters but strongly at 14B (Quinn models).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Catastrophic Forgetting in AI

0:45 to 2:32

Discusses the concept of catastrophic forgetting in AI and its implications.

“You don't know a C major from a car door.”

The Limitations of Supervised Fine Tuning

2:32 to 3:42

Explains the shortcomings of current AI training methods, especially SFT.

“And honestly, looking at the research, it's one of the most elegant solutions I've seen because it sort of flips the script on who is doing the teaching.”

The Need for Active Learning

3:42 to 6:05

Outlines the importance of active learning vs passive learning in AI training.

“You take a base model, which is like a student who has read the whole internet but hasn't specialized in anything, and you just show it thousands of specific examples.”

Introduction to Self-Distillation Fine-Tuning

6:05 to 7:05

Introduces self-distillation fine-tuning as a solution to AI learning issues.

“The feedback is immediate and objective.”

How Self-Distillation Works

7:05 to 9:30

Explains the mechanics of self-distillation and its benefits for AI.

“I know it involves this in-context learning concept.”

Results from Experiments on Self-Distillation

9:30 to 12:07

Discusses the experimental results demonstrating the effectiveness of self-distillation.

“So if the teacher is 99 % sure the answer is blue, the student needs to learn to be 99 % sure it's blue, not just guess blue by accident.”

Comparison with Traditional Learning Methods

12:07 to 13:55

Compares self-distillation fine-tuning with traditional methods and highlights its advantages.

“You try to mix in some review sessions for the standard model.”

Understanding Oracle RAG Systems

14:01 to 15:12

Learn how Oracle RAG systems utilize external information for enhanced accuracy.

“ORAGAS stands for Retrieval Augmented Generation.”

The Importance of Understanding vs. Memorization

15:13 to 17:01

Discover the difference between memorizing facts and truly understanding concepts.

“The SDFT model achieved nearly perfect accuracy on these indirect questions.”

Teaching AI to Think Deeply

17:02 to 18:22

Explore how SDFT enables AI to develop reasoning and deep thinking skills.

“So you can teach a model to think deeply, even if you only feed it the final answers.”
Show all 11 chapters

The Shift Towards Continuous Learning in AI

18:23 to 20:12

Examine the transition from static models to AI that learns continuously and autonomously.

“The smarter the model, the better it can leverage that in-context cheat sheet to lift itself up.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, I want you to picture something for a second. Imagine you decide you're going to really commit to learning the piano. You buy a nice keyboard, maybe, I don't know, an upright if you're feeling fancy. You spend six months practicing scales, learning chords. Your fingers hurt. Oh, yeah. The pain is real. Right. But finally, you master a little bit of Mozart. You can play Rondo alla Turca without stumbling. You feel great. It's a fantastic feeling. Mastery is addictive. Exactly. But then the next year rolls around and you decide you want to get fit. You want to learn how to play tennis. Okay.

0:36You go out, you hire a coach, you practice your serve, you work on your backhand, and you get pretty decent. You're hitting winners. But then one rainy afternoon, you sit down at the piano and you stare at the keys blankly. A total blank slate. Complete amnesia. You don't know a C major from a car door. You don't know what your fingers are supposed to do. because somehow the act of learning how to swing a tennis racket completely overwrote the part of your brain that knew how to play music. That would be a nightmare. I mean, it's the definition of a one-track mind. Life would be impossible. It would.

1:09But thankfully for us humans, that is not how our biological brains work. We don't have to delete old files to make room for new ones. Our capacity is, well, it's expansive. Right. I can ride a bike and do long division. Usually not at the same time, thankfully. But I have both skills stored away. But for artificial intelligence, this is a massive, massive problem. It's what researchers call catastrophic forgetting. It is the Achilles heel of modern AI. We have these, you know, incredible models like the ones powering the chatbots everyone uses. But they are surprisingly brittle in this specific way.

1:44How so? They're static. They are frozen in time the moment their training finishes. If you try to teach them a new trick, say, specialized medical coding, after they've been trained on general literature, they tend to just overwrite the old data to make room for the new. So it's like they have a very, well, a very slippery memory. To learn the new thing, they have to push the old thing right off the ledge. Precisely. And that is the bottleneck we are talking about today. Because we don't want chatbots that just stay the same forever, you know, stuck with the knowledge they had in 2024, 2025. we want AI that grows with us.

2:20Lifelong learners. Exactly. AI that can learn to code today, learn medical diagnosis tomorrow, and not forget how to write a poem in the process. Which brings us to the breakthrough we are diving into today. It's a concept called self-distillation fine-tuning or SDFT. And honestly, looking at the research, it's one of the most elegant solutions I've seen because it sort of flips the script on who is doing the teaching. It really does. It moves away from this idea that humans have to constantly, you know, spoon feed the AI. It sounds super technical, but the core idea, struck me, is almost psychological.

2:54It's basically about the AI developing a split personality. A teacher and a student. Yeah, to coach itself. So that is a very apt way to put it. It's using its own latent intelligence to bootstrap its way to new skills. So today's mission, we are going to uncover how an AI can use a smarter version of itself to learn continuously without forgetting the past. And here is the kicker. Without needing a human to stand over its shoulder and grade every single quiz. Which is a game changer for scalability. Because if you need a human to grade every output, you eventually run out of humans. Or you run out of money paying them.

3:29But before we get to the solution, I think we need to slow down and really unpack why the current way we teach AI is so broken. What are we doing right now that causes this forgetting? So the standard method right now is something called supervised fine tuning or SFT. SFT. Think of SFT as the cramming method. You take a base model, which is like a student who has read the whole internet but hasn't specialized in anything, and you just show it thousands of specific examples. Right. If you want it to be a doctor, you show it thousands of medical records and diagnoses. So you're basically showing it the textbook and saying, memorize this, copy this.

4:04Yes. And while it works to get the model to perform that specific task, it comes with that heavy cost we mentioned. Catastrophic forgetting. The data shows a literal seesaw effect. A seesaw? Yeah. If you take a model that is good at general science and you use SFT to teach it tool use, like how to use a calculator or a calendar API, its science scores don't just dip, they plummet. So it is zero sum. To fit the tools in, the science has to go. Right. But there is a deeper reason for this that goes beyond just, you know, running out of space. It's about how the learning happens. SFT is what we call off-policy learning.

4:41Okay, off-policy. I saw this term in the notes. Break that down for me because that sounds like insurance jargon. It does, doesn't it? But it's actually about passivity versus activity. Off-policy means the AI is passive. It is just looking at static data demonstrations provided by others. So it's not actually trying to do the task. Exactly. It's just looking at what a human did and trying to minimize the error between its guess and the human's text. So to use a sports analogy, it's the difference between watching a YouTube video of someone shooting a basketball versus actually going out on the court and throwing the ball yourself.

5:15That is the perfect analogy. Watching the video is off policy. You can watch LeBron James shoot free throws all day. You might understand the mechanics intellectually, but you aren't building the muscle memory. Right. No, no. Shooting the ball is on policy. When you shoot the ball, you feel the weight, you see the arc, you adjust your muscles based on whether you missed short or long. That feedback loop creates a much deeper, more resilient memory. It changes your neural pathways in a way that just watching never could. Okay, so if we know that doing it yourself is on policy stuff is better, why don't we just always use the shooting the basketball method?

5:52That's reinforcement learning, right? That's how AlphaGo learned to play Go. It is. Reinforcement learning, or RL, is fantastic because it is on policy. The AI tries, fails, learns, but there is a catch. RL requires a reward function. You need a scoreboard. A scoreboard. In basketball, it's easy. Ball goes in hoop, points go up. The feedback is immediate and objective. In go or chess, you win or you lose. But in the real world? In the real world, life doesn't give you a score. If you write a polite email to your boss, a giant number doesn't appear above your head saying plus 10 points. Right. If a doctor diagnoses a patient, they don't get an immediate correct notification from the universe.

6:35It might take weeks to know if the diagnosis was right, or it might be subjective. Yeah, totally. Exactly. So we're stuck. We need the learning by doing of reinforcement learning to stop the forgetting. But we don't have the scoreboard to make it work. We have the desire to shoot the basketball, but we don't have a hoop. And that is exactly where self-distillation fine-tuning SDFT comes in. Yeah. It solves the scoreboard problem. It does. It essentially allows the AI to hallucinate its own hoop. OK. Walk me through this. How does it do that? I know it involves this in-context learning concept.

7:08Right. So to understand SDFT, you have to understand a weird quirk of large language models. If you show an AI a really good example in the prompt, just one good example, it suddenly gets way smarter. This is the few shot prompting thing, right? Like if I wanted to write a poem and I paste a Shakespeare sonnet in the prompt first, it writes a better poem. Yes. It's called in context learning. It's like giving a student a cheat sheet just for that one conversation. If you give the model a prompt that includes a demonstration of how to solve a physics problem, the model temporarily acts like a much more capable version of itself.

7:46It adopts the reasoning style of the example. Okay, so we have a way to make the model temporarily smarter. How does SDFT turn that into permanent learning? It uses that temporary boost to create the split personality we talked about. You have the teacher and you have the student. But they're the same model. They're the same underlying software weights, yes, but they are viewing the world differently. The teacher is the model with the cheat sheet. It sees the question, plus a perfect example. This conditions the model to act wisely, to use better reasoning. And the student. The student is the exact same model, but it's looking only at the question.

8:21No cheat sheet. It's flying blind. So the student is trying to solve the problem on its own, like a test, while the teacher is sitting there with the open textbook. Exactly. And here's the mechanism. The student tries to answer the question on its own. This is the on policy part. It is generating the answer. It is doing the work. It's shooting the basketball. It's shooting the basketball. Then the system looks at what the teacher would have done. Since the teacher is smarter because of the cheat sheet, the teacher's output serves as the target. So the teacher becomes the scoreboard. In a sense, yes.

8:55But it's not just grading right or wrong. The system calculates the difference between the student's attempt and the teacher's ideal distribution. Technically, this is minimizing the reverse KL divergence. Whoa. Okay. Pump the brakes. Reverse KL divergence. That sounds like something from Star Trek. What does this actually mean? Sorry, hazard of the trait. Let's strip the math away. Think of it as measuring the vibe or the confidence of the answer. The student isn't just trying to copy the teacher's words exactly. Okay. It's trying to match the teacher's probability distribution. It's trying to match the teacher's confidence.

9:30So if the teacher is 99 % sure the answer is blue, the student needs to learn to be 99 % sure it's blue, not just guess blue by accident. Exactly. It pushes the student to align its internal thought process, its neural pathways, to match the wiser version of itself. And then the student updates its brain, its parameters, to lock that in. This is the aha moment for me. Yeah. The AI is coaching itself. It doesn't need a human to score the test because the teacher version of itself creates the target. It's generating its own reward. You've got it. It is learning by doing rather than learning by reading.

10:07But the coach is just a temporarily boosted version of itself. And because it's on policy, because the student is actually doing the task, it retains the information much better. It integrates it. It integrates the skill into its muscle memory, so to speak. Well, talk is cheap. Let's look at the actual data, because the researchers didn't just propose this theory. They ran a learning gauntlet to see if it actually stops the forgetting. They did. They really put the model through the ringer. They set up a sequential learning experiment. They took a model and tried to teach it three distinct skills, one after another, specifically to see if its brain would get full and start deleting things.

10:46You know, three skills were pretty different, right? Very different. First, tool use-like, learning how to use software APIs, calling a calendar, using a calculator. Second, science Q &A, undergraduate-level reasoning. And third, medical diagnosis. That is a tough lineup. So first they tried the standard method, the SFT cramming method. What happened? Well, they taught at tool use, and it got good at tool use. Then they taught at science. And tools. Gone. The ability to use tools crashed. The researchers described it as oscillatory behavior. Oscillatory. So it just bounces back and forth. Yeah.

11:20It learns science, drops tools, learns medical, drops science. It couldn't hold all three. It's the piano and tennis player all over again. It is. It's catastrophic for getting an action. But when they used SDFT, the self-teaching method, it acted like a human learner. It learned the tools. Then it learned science. And it kept the tool skills. Kept them high. Then it learned medical and kept both the previous skills. It didn't have to delete the files. No. It achieved what they call Pareto efficiency. That's a fancy economic term. But here it just means it found the best possible balance where learning a new thing didn't cost it the old thing.

11:58It accumulated skills instead of trading them. I saw in the notes they even tried to help the standard model by reminding it of the old stuff. A method called re-invoke. Yes, that's a common trick. You try to mix in some review sessions for the standard model. Here's some science, but hey, remember how to use a calculator. Like a spaced repetition app for an AI. Right. And even that didn't work as well as SDFT. Really? Why not? Reviewing usually helps. Because the standard method is still passive. It's still off policy. The model is just looking at the review material, not doing it. The self-distillation process is active.

12:33It fundamentally changed how the information was integrated. It wasn't just reviewing. It was integrating the skills into its core policy. That covers skills, learning how to do things. But what about facts? learning what happened. We were recording this in February 2026. The models we use were trained on data that cuts off before 2025. The famous knowledge cutoff, the bane of every AI user who asks who won the Super Bowl last week and gets a total hallucination. Exactly. So if I ask a standard model about the massive earthquake in Myanmar that happened in 2025. It would have no idea. Or worse, it might just make up a fake earthquake from 2022 to try and make you happy.

13:11So they ran an experiment to see if they could inject this new history into the model. They fed it Wikipedia articles about the real 2025 disasters, the Myanmar earthquake, Typhoon Kalmagi, Hurricane Melissa. Right. And the challenge here is, can the model learn these facts permanently just by reading them using this SDFT method versus traditional fine-tuning? They checked something called strict accuracy. How did the standard cramming method do? Standard SFT scored roughly 80%. 80%. That's a B -. It's okay, but not great. It hallucinated details, it got dates wrong, or mixed up which disaster affected which region.

13:49You don't want a B - when you're dealing with disaster relief data. And SDFT. SDFT scored 89%. That is a massive jump in this field. But the comparison is what's crazy. It nearly matched an Oracle RAG system. Wait, Oracle RAG? Help us out here. Right. ORAGAS stands for Retrieval Augmented Generation. Basically, an Oracle ag system is allowed to cheat. It gets to look up the Wikipedia article every single time it answers a question. It has the book open in front of it. It has the open book and the SDFT model. From memory. It had to answer from memory. It had processed the articles during training, but during the test, the book was closed.

14:26And it still matched the accuracy of the system that had the open book. That's impressive memorization. But here is where it gets really interesting to me. It's not just about parroting back facts. They tested for understanding using indirect questions. This is the true test of intelligence. Instead of asking what happened in Myanmar in 2025, which you can answer by just regurgitating a sentence, they asked things like which countries required humanitarian aid in 2025. See, that requires you to know that an earthquake happened in Myanmar and that it was bad enough to need aid and then connect those dots to answer Myanmar.

15:03You can't just interrel f the tax. Exactly. It requires a worldview, a mental model of the event, not just a keyword search. The standard SFT model failed here. He didn't get it. It had memorized the sentences, but it hadn't connected the concepts. He didn't really understand the event. The SDFT model achieved nearly perfect accuracy on these indirect questions. It understood the events. That is the difference between memorizing a history textbook for a test and actually understanding the geopolitical landscape. And it goes back to the teacher-student dynamic. Because the student model had to generate the answers itself during training, guided by the teacher, it integrated that knowledge into its reasoning centers, not just its surface level memory.

15:43Speaking of reasoning, this brings up a bit of a paradox that I found fascinating. We usually assume that to teach an AI to think step by step, you know, that chain of thought process, you need to show it examples of thinking. Right. If you want a model to solve a math problem, you don't just show it the answer. You show it. Step one, do this. Step two, do that. Therefore, the answer is X. But often we only have the answer. In medical data sets, we might have patient symptoms and final diagnosis. We don't have the doctor's internal monologue explaining why. And this is where the SFT trap is dangerous.

16:17If you use standard fine-tuning on just the answers, the model gets lazy. It learns to just blurt out the diagnosis. Patient has a cough. Flu. Right. Because you trained it to just give the answer. You optimize for brevity. its responses get shorter, it stops thinking. But with SDF key, something magical happens. What's that? Remember, the teacher model has the cheat sheet. When a smart model sees a cheat sheet with a complex example, it naturally generates intermediate reasoning to bridge the gap between the question and the answer. So the teacher starts thinking out loud in a way. Implicitly, yes.

16:51The teacher creates the reasoning steps in its probability distribution. And because the student is trying to match the teacher's distribution, the student is forced to learn the reasoning process, even though the original training data didn't explicitly write it out. That is wild. So you can teach a model to think deeply, even if you only feed it the final answers. The data proves it. They measured the length of the responses. The SDFT model wrote longer, more detailed responses, averaging about 4180 tokens, compared to the standard model, which dropped to around 3200. Wow. It didn't get lazy.

17:25It retained its ability to think deeply. It effectively reverse engineered the logic needed to get the answer. But I'm guessing this relies on the teacher actually being smart. Right. Like if the teacher is dumb, the student is just learning from a dummy. Garbage in, garbage out. That is the crucial constraint. And the researchers actually tested this. They did a scaling analysis using the Quinn models. They tested a small 3 billion parameter model, a 7 billion, and a large 14 billion parameter model. And what did they find? On the small 3 billion model, SDFT didn't work that well. The teacher just wasn't smart enough to extract high-quality guidance from the examples.

18:03It was like a C student trying to tutor another C student. Not much progress there. You can't pull yourself up by your bootstraps if you don't have boots. Exactly. But on the 14 billion model, the gap between SDFT and standard methods was massive. So the smarter the model is to begin with, the better it is at teaching itself. It's a virtuous cycle. Intelligence begets more efficient learning. The smarter the model, the better it can leverage that in-context cheat sheet to lift itself up. So what does this all mean? We've covered a lot of ground skills, facts, reasoning. But zooming out, we're looking at a pretty fundamental shift here.

18:38We are witnessing the move from static models to continuous learners. For the last few years, AI has been trained in these massive, expensive runs that take months and then frozen. If you wanted to teach it something new, you had to be very careful or restart the whole process. Yeah. SDFT suggests that models can use their own latent intelligence to digest new experiences. They can effectively hallucinate their own rewards by comparing themselves to their smarter, example-aided selves. And it bypasses the need for those massive farms of humans clicking thumbs up or thumbs down on every response.

19:12It does. We used to think we needed human feedback, RLHF, to fix everything. This shows that the model, if properly set up with this split personality structure, can provide its own feedback signal. it becomes self-correcting. It's impressive, but I have to admit it's also a little eerie. In what sense? Well, think about the trajectory. We have an AI that can now effectively teach itself new skills. It can ingest new facts without forgetting the old ones. It doesn't need us to grade its homework. Right. If we're approaching the point where the training run never actually ends, what happens when a model runs this loop on the entire internet continuously forever?

19:50That is the provocative question. If the bottleneck of human supervision is removed and the bottleneck of catastrophic forgetting is solved, then the limit becomes simply compute and data. We might be moving towards systems that don't just update once a year, but evolve in real time every second of every day. A model that never sleeps, never forgets, and is constantly teaching itself to be better. That is a lot to process. That's all for today's Deep Dives into Self-Distillation Fine Tuning. We'll catch you on the next one. Keep learning!

From the publisher

This research introduces **Self-Distillation Fine-Tuning (SDFT)**, a novel on-policy learning method designed to help large language models acquire new skills without suffering from **catastrophic forgetting**. Unlike traditional supervised fine-tuning, which often causes models to lose prior knowledge, **SDFT** utilizes the model’s own **in-context learning** abilities by using a version of itself conditioned on demonstrations as a teacher. This approach generates **on-policy training signals** that allow the model to internalize new facts and reasoning patterns while remaining close to its original parameter distribution. Empirical results across **skill acquisition** and **knowledge injection** tasks show that **SDFT** consistently outperforms existing baselines in both task accuracy and the preservation of general capabilities. Ultimately, the research positions **self-distillation** as a practical and scalable path for enabling **continual learning** in foundation models.

More from Best AI papers explained

All 475 episodes
Self-distillation enables continual learningBest AI papers explained · 20 min
Listen in VO