In short
Sample-efficient parametric learning for LLMs that makes natural-language feedback permanent, unlike ephemeral in-context learning (ICL).
Guests
Not specified in the transcript (only two unnamed speakers).
Key claims
(1) ICL adapts instantly but disappears when the context window clears. (2) Traditional supervised fine-tuning (SFT) is permanent but data-hungry. (3) A three-step method—collect feedback, generate a corrected rollout, then SFT on the prompt/output while removing the natural-language instruction—forces true rule internalization; keeping the instruction causes “dependency” on the prompt.
Notable examples
Factual rule learning (no consecutive “0000”): 91.4% accuracy with 16 examples vs 87.1% ICL; dependency baseline 49.2%. Stylistic math task: 98.4% style adherence from 1 example vs 93.4% ICL; dependency baseline 2.7%. Iterative updates: aligned API v1→v2 works (100%), but adding v3 drops to 34%; cross-domain update preserves old API v2 (100%) but barely learns new style (7.8%). Proposed mitigations: KL-regularized SFT and data interleaving.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Challenge of Temporary Adaptation
0:45 to 2:12
Discussion on the limitations of in-context learning in large language models.
“Which brings us to our mission for this deep dive.”
Introducing Sample-Efficient Parametric Learning
2:12 to 4:00
Overview of a method to internalize user feedback effectively.
“You know, a lot of prior work kind of struggled with this.”
Step-by-Step Breakdown of the Method
4:00 to 5:44
Detailed explanation of the three steps to internalize feedback.
“Hold on, that feels completely backwards.”
Core Insight on Dependency Learning
5:44 to 6:28
Understanding the importance of removing instructions during fine-tuning.
“The proposed method SFT with the instruction removed hit 91.4 % accuracy with just 16 examples.”
Real-World Applications and Results
6:28 to 7:39
Examining how the method performs across different tasks.
“So training it with the instruction present literally made it incapable of learning the rule.”
The Downside of Permanent Updates
7:39 to 9:16
Exploring the trade-offs and potential issues with iterative learning.
“You have to remove the instruction to force internalization.”
Exploring Solutions to Maintain Flexibility
9:16 to 10:57
Discussing potential methods to mitigate the loss of learning capacity.
“That is a really fascinating and slightly terrifying problem.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. So if you work with large language models, you know this feeling. They're amazing at adapting on the fly. Oh, absolutely. You give them an instruction in context, that's ICL, in context learning, and, you know, tell it to use a certain tone or a specific structure, and it just does it instantly. It's like magic. But, and this is the big but, that adaptation is completely ephemeral. It's temporary. Right. The second you clear that context window, poof, the model just snaps back to its old self. that that new behavior is gone. And that's a huge problem for anyone trying to, say, build a system that learns from user feedback over time.
0:40I mean, if you want your LLM to permanently adopt a new style guide or a new API protocol, this transient memory is a serious limitation. It is. Which brings us to our mission for this deep dive. We need to find a way to bridge that gap. We want the permanence of, you know, traditional supervised fine-tuning, but we need it to be efficient enough to work with just simple, natural language feedback. The kind of feedback a normal user might give. Exactly. We're talking about customizability that actually sticks, even if you only have a tiny amount of data. And today, we're going to unpack a, well, a surprisingly simple three-step method that's designed to do just that.
1:15To bake that natural language feedback right into the model's weights. Make it permanent. So let's just quickly review the state of play. Like you said, LLMs are built around ICL. It feels amazing. You can modify behavior at inference time. But again, it's a witness when you want continuous improvement. Think about a customer service bot that needs to learn a new rule for escalating a call. You can't just keep reminding it in the prompt every single time. That's not scalable. It's not. And on the other side, you have supervised fine-tuning, SFT. Now, that's how models really learn. It adjusts the model's actual parameters, its weights.
1:51So that gives you the permanence we're after. It does. But traditional SFT is, let's just say it's data hungry. It needs huge amounts of structured data, and it's just not practical for this kind of on-the-fly user feedback. So we have two extremes, the permanence of SFT, which needs tons of data, and the sample efficiency of ICL, which is totally temporary. The real innovation is connecting those two worlds. Precisely. You know, a lot of prior work kind of struggled with this. It either focused on static instructions, which isn't really dynamic, or it assumed you had these huge data sets lying around.
2:25Which most people don't. Right. So what we're looking at here is a method for what the recruiters call sample-efficient parametric learning. It's all about making that learning step itself incredibly efficient. Okay, so that brings us to the core method itself. The researchers say it's simple. Let's see. What are these three steps to force a model to really internalize feedback? All right. Step one is the easiest. It's just obtain feedback. The model does something and you get a correction or an instruction in plain English. So it could come from a person, an automated tool. Yeah, another LLM, whatever.
2:58Let's say your model generates Python code with a bug. The feedback might be something like, make sure all API calls use version two syntax. Simple enough. What's step two? Step two is the sample generation, or what they call the rollout. You take that feedback, that V2 syntax rule, and you put it back into the context with the original prompt. Then you have the model generate a new answer. And because the instruction is right there, it uses its ICL ability to produce the correct code. We can call that new output, answer prime. Okay, so at this point, the model has done the right thing, but only because we were holding its hand with the in-context instruction.
3:37Now for step three, I hear this is the really clever part. This is the pivotal step. Fine-tune without feedback. You take that new, correct output answer prime, and you use it for supervised fine-tuning. But, and this is crucial, you remove the natural language feedback from the input. Wait, what? The training example becomes just the original prompt paired with the desired output. The instruction itself is gone. Hold on, that feels completely backwards. Right. I mean, if the instruction is what caused the right answer, Or wouldn't you want to keep it in during fine-tuning to, you know, strengthen that connection?
4:10You'd think so, wouldn't you? But that is the core insight of this whole thing. The Reedyard shows that if you keep the instruction in the prompt during fine-tuning, the model learns the wrong lesson. The wrong lesson. It learns dependency, not knowledge. It learns that it should only produce that correct answer when that exact instruction is present. It doesn't actually internalize the rule itself. So if I trained a model with write a report plus use bold headings, it doesn't actually learn to use bold headings in general. No. It just learns that when it sees the magic phrase use bold headings, it should perform that action.
4:48Take the phrase away and the behavior vanishes. It's still just relying on the context. I see. So by removing the instruction in step three, you're making the problem harder for the model. Exactly. You're forcing it to look at the original simple prompt and the desired output and figure out the missing piece itself. It has to bake the rule about V2 syntax or bold headings into its actual parameters because the external clue is gone. It forces true learning, not just prompt following. That's the idea. True parametric learning instead of shallow contextual dependency. OK, that distinction is huge.
5:22And they really put this to the test, right? They use two very different kinds of tasks. They did. First up was factual rule learning. This wasn't about style. It was a hard, logical rule. The task was to teach the model to generate binary strings that follow one specific constraint. No consecutive thousand sequences. That's a tricky one. It's a negative constraint it has to track across the whole output. Yeah. So how did the method do? How much data did it need? The efficiency was pretty remarkable. The proposed method SFT with the instruction removed hit 91.4 % accuracy with just 16 examples.
5:5816. Just 16. And get this, that's actually better than the ICL baseline, which only got 87.1 % accuracy, even with the rules sitting in the prompt every single time. So a tiny permanent update made the model better at the task than when it was constantly being reminded. Yes. But here's the real kicker. The proof. The baseline where they did SFT but kept the instruction in the prompt. The dependency model. It failed. Completely. It scored only 49.2 % accuracy. That's basically a cone flip. It's random. Wow. So training it with the instruction present literally made it incapable of learning the rule.
6:33It prevented any real persistent knowledge from forming. And they found the exact same thing when they moved to a more, let's say, real-world task. Stylistic adaptation. Right. A math style task. The instruction was, point out common wrong turns before solving the problem. This is a classic user request, like asking for a specific brand voice. And this is where sample efficiency really, really matters. I might only have one perfect example of my brand's voice, not 16. Well, you're in luck. The performance here was just wild. Yeah. The method got to 98.4 % style occurrence after training on just a single example.
7:09One example. One, a single training pair was enough to make that stylistic change permanent, and it did it without hurting the model's math accuracy. And that one-shot update was still better than the ICL baseline. Better again. The ICL baseline was 93.4%. And just to close the loop, the dependency baseline, the one that kept the instruction in during training. Let me guess. It failed. Miserably. It kept 2.7 % adherence. It just completely ignored the style. So it confirms whether the knowledge is hard and fast factual rules or soft stylistic preferences, You have to remove the instruction to force internalization.
7:42But now we have to talk about the big catch, the real-world problem of iterative learning. Right, because in reality, users are giving feedback all the time. You're not just making one update. You're trying to stack them. What happens then? The short answer is it gets messy fast. They found that small, closely related updates can work okay, like teaching the model about API version 1 and then updating it to version 2. That worked perfectly. 100 % accuracy on both. Okay, so sequential updates are fine if the topics are aligned. Where does it break? The very next step. When they introduced a third update to API version 3, the whole thing fell apart.
8:18Generalization just broke. The accuracy on the v3 task was only 34%. Yikes. So the repeated targeted updates made the model brittle, over-specialized. That seems to be it. It's almost like every time you make one of these permanent updates, you might be chipping away at the model's general ability to learn from context later on. Its ICL capability seems to degrade. So that's the trade-off. It's a huge trade-off. They ran one more test to prove it, a cross-domain experiment. They took the model that was an expert on API v2 and tried to teach it the math style rule. And the old knowledge. The old knowledge stuck.
8:56It still got 100 % on the API v2 task. That was the good news. But I'm guessing it didn't learn the new math style very well. It barely learned it at all. Only 7.8 % adherence to the style. out. The first permanent update for the API seems to have made the model deaf to new contextual instructions. It suppressed its ICL sensitivity. That is a really fascinating and slightly terrifying problem. We've figured out how to teach a model something permanently with almost no data, but the very act of teaching it seems to break its ability to learn other things in the moment. Exactly. The permanent updates work, but they might be calcifying the model, making it less adaptable to new, unrelated feedback down the line.
9:36So this is the new frontier then, figuring out how to mitigate this. How do we stop the model from breaking its own ICL? There are a couple of promising directions. One is about constraining the update itself, using things like KL regularized SFT to make sure the weight changes are really small and don't stray too far from the original model. Keep the update subtle. Right. The other approach is more about the data, using a technique called data interleaving, where you mix in a few examples from the model's old training data to sort of remind it of its general capabilities and slow down that catastrophic forgetting.
10:12What a journey. We have this clear, powerful and counterintuitive path to making knowledge permanent in LLMs. It all comes down to that one critical step. Yeah. Taking away the instruction during fine tuning. That's the key. It's the difference between true learning and, well, near total failure across every tasks they tested. But this path to creating truly customizable, ever-improving models is also a bit of a minefield. Specialization can break generalization. So here's the final thought I want to leave you with. Given that these permanent updates can damage a model's fundamental ability to follow instructions, how do we teach an LLM permanent new tricks without breaking its ability to learn in the moment?
10:52In other words, how do we get this powerful parametric learning without destroying in-context learning? That's the puzzle the next generation of LLMs will have to solve.
From the publisher
This research paper provides a novel approach for sample-efficient parametric learning in large language models (LLMs) using natural language feedback, addressing the transience of traditional in-context learning (ICL) and the data inefficiency of standard fine-tuning. The authors propose a simple three-step method: obtaining natural language feedback, sampling a generation conditioned on that feedback, and then performing supervised fine-tuning (SFT) on the new generation with the feedback removed from the prompt, which forces the model to internalize the instruction into its weights. This technique is evaluated against ICL and SFT baselines across both factual rule-learning (DFAs) and stylistic adaptation tasks, demonstrating superior performance with limited data budgets. However, preliminary results on iterative learning show that while small sequential updates are possible, the compounding of feedback quickly leads to catastrophic forgetting and interference.




