Fine-Tuning Strategies for Preserving In-Context Learning in Linear Attention

19 Mar 2026 · 19 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How fine-tuning can cause “localized amnesia,” degrading in-context (few-shot) learning even while improving zero-shot task performance, and what to do instead.

Guest backgrounds

No guests are named in the transcript; it’s a two-host discussion.

Key claims

Full fine-tuning (updating query/key/value) overwrites the query-key “reasoning blueprint” (modeled via inverse covariance), breaking few-shot pattern recognition. Value-matrix-only tuning (freeze Q and K, update V) preserves few-shot reasoning while learning new facts. Adding an auxiliary loss to optimize both zero-shot and few-shot for one domain improves in-domain few-shot but harms out-of-distribution tasks due to limited parameter space and “distance” between domains (cosine similarity).

Notable examples

HR handbook memorization vs forgetting vendor-contract parsing; Monopoly vs chess analogy; experiments on Qwen2.5 3B Instruct on MMLU humanities vs STEM (7-shot gains ~1.77 points in humanities, ~6.48-point drop in STEM).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Localized Amnesia in AI

0:28 to 1:40

Exploration of how fine-tuning affects an AI's memory and reasoning capabilities.

“We're talking about how we teach these systems to do what we want without, you know, breaking them in the process.”

Paradigms of Learning: In-Context vs. Fine-Tuning

1:40 to 3:30

Discussion on the differences between in-context learning and fine-tuning in AI.

“Let's start by looking at the paradigm split.”

The Tradeoffs of Fine-Tuning

3:30 to 6:20

Investigating the tradeoff between specialized knowledge and general reasoning skills in AI models.

“Fine-tuning, on the other hand, is like choosing to permanently memorize the rules for Monopoly.”

The Mechanics of AI Memory: Query, Key, and Value

6:20 to 10:30

A deep dive into the attention mechanism of AI and how it relates to memory and learning.

“The value matrix, on the other hand, acts as the pathway that encodes the specific target task knowledge.”

The Fine-Tuning Problem Explained

10:30 to 12:10

Detailed analysis of how full fine-tuning can degrade an AI's learning capabilities.

“The theoretical proofs show that you can achieve near optimal zero shot performance on a new task while keeping the few shot reasoning capabilities perfectly intact.”

Value Matrix Fine-Tuning: A Solution

12:10 to 13:30

Introduction to value matrix fine-tuning as a method to retain reasoning while learning new facts.

“But you know, in the world of machine learning, nobody is ever satisfied with a standard win.”

Empirical Testing of the Auxiliary Loss Method

13:30 to 14:02

Discussion of empirical tests conducted to validate the effectiveness of the auxiliary loss method in fine-tuning.

“They wanted to force the model to become a humanities savant.”

The Trade-offs of Fine-Tuning in AI

14:02 to 17:39

Explore the implications of hyper-specialization in AI fine-tuning and its impact on performance across different domains.

“It learned humanity's zero shot, and it stayed incredibly sharp at humanity's few shot.”

The Future of Modular AI Development

17:40 to 18:38

Discuss the potential for modular AI systems that allow for flexible reasoning and fact databases.

“It inherently shifts the balance of the entire system.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine you spend like millions of dollars fine tuning in artificial intelligence model to just perfectly memorize your company's proprietary HR handbook. Right, like every single detail. Exactly. It knows every policy, every vacation accrual rule inside and out. But then the exact moment it finishes learning that specific information, it suddenly completely forgets how to parse a simple vendor contract. Yeah, it develops this localized amnesia. Welcome to the deep dive. Today, we are looking at something that sits right at the center of this massive tug of war in modern AI. We're talking about how we teach these systems to do what we want without, you know, breaking them in the process.

0:40It's a huge problem right now. It really is. On one side of this tug of war, you have in-context learning, which is a few-shot prompting. And then on the other side, you have fine-tuning, where you're permanently updating the AI's internal parameters. Right. The mission of this deep dive is to explore a really fascinating puzzle. Like how can we teach an AI new highly specialized tricks without making it forget its general flexible reasoning skills? Because that localized amnesia is a very real, very frustrating phenomenon for developers. I mean, you want a specialist that knows your domain perfectly, but you absolutely do not want to lose the generalist reasoning capabilities that, well, made the model powerful in the first place.

1:21Exactly. And to figure out how to balance that, Today's stack of notes includes a pretty heavy theoretical mathematical analysis of something called linear attention models. And that is backed up by empirical tests on actual large language models. Yeah, the mask gets deep, but the implications are incredibly practical. Okay, let's unpack this. Let's start by looking at the paradigm split. For anyone listening who builds or deploys AI, you probably already know the pain points here. Paradigm one is in context learning. Right, the prompting approach. Yeah. You give the AI a prompt with a few examples, and it figures out the pattern.

1:55It's incredibly flexible. You can get it to do almost anything. But the drawback is the inference cost. I mean, you are passing massive amounts of context to the model every single time you query it. Which gets expensive fast. Very expensive. It's computationally demanding. It restricts your context window, and it introduces a lot of latency. The model basically has to process your entire list of examples from scratch for every single interaction. Right, which naturally drives everyone toward Paradigm 2, which is fine-tuning. Instead of fielding the AI examples in the prompt over and over, you feed it a bunch of training data once.

2:30You're permanently changing its internal parameters. Exactly. This adapts the model so it can perform zero-shot tasks. You don't give it any examples in the prompt, you just ask the question, and it answers based on its newly updated weights. At test time, that zero-shot performance is cheap and lightning fast. But the problem, and this is the core conflict the data explores, is that recent empirical studies reveal a severe tradeoff. The amnesia. Yes. Improving that zero-shot performance through fine-tuning actively degrades the model's in-context learning abilities. Oh, wow. So it gets worse at the very thing it was originally good at.

3:06Exactly. It literally makes the model worse at learning from examples in a prompt, especially when you try to get it to reason through unseen, out-of-distribution tasks. I picture it kind of like learning to play a board game. Prompting the AI is like reading the rulebook every single time you sit down to play. Right, which is tedious, but you can play anything. Exactly. You have total flexibility. You can play any game in the world as long as you have the rulebook in front of you. But like you said, super tedious. Fine-tuning, on the other hand, is like choosing to permanently memorize the rules for Monopoly.

3:40So you can just sit down and play Monopoly instantly. No rulebook needed. Right. But the sheer act of hardwiring those specific rules into your brain somehow, like, overwrites your ability to learn chess. You gain specialized efficiency, but you sacrifice generalized adaptability. That's a great way to think about it. And to understand the mechanics behind why that amnesia happens, we have to look into the hood at the AI's architecture. Into the math. Right, into the math. The analysis we are diving into uses linear attention models, solving linear regression tasks. Standard transformers use a complex function called softmax, which is, well, it's essentially a mathematical black box.

4:20Hard to see what's actually happening inside. Notoriously difficult to analyze cleanly. But by stripping that down to linear attention models, the math becomes tractable enough to solve perfectly on paper, while still accurately mimicking how the massive multibillion-parameter AI models actually behave in the wild. Okay. And the absolute heart of that behavior, the thing that determines whether the AI is actually reasoning or just, you know, reciting memorized facts is the attention mechanism. Yes, the core engine. And that relies on three distinct pillars, right? The query, key, and value matrices, the famous Q, K, and V.

4:53So I want to push back on this a little or at least clarify. If the AI is just crunching numbers, just multiplying massive grids of decimals together, what role do these specific matrices actually play in memory and learning? Like, if the query is basically the search bar, the mathematical representation of what the user is asking for, then the key matrix must be the metadata tag it's trying to match against, right? The organizational labels. That captures the dynamic, yeah. But it's a bit more structural than just a simple tag. The query and key matrices work together to form a relational map.

5:27A map. Yeah, a map. When the AI compares your query to the keys, it isn't just looking for a simple keyword match. It is calculating the relevance and the conceptual distance between what you want and what it already knows. Ah, okay. Which leaves the value matrix. So if Q and K are the search bar and the index, the value matrix is the actual content. Yes. It's the paper inside the folder that gets pulled out when a query successfully matches a key. Precisely. The value matrix holds the specific information that gets retrieved and synthesized. But the underlying question the research probes is how these matrices actually function during the learning process itself.

6:04Like, does the AI actually separate its reasoning engine from its factual database? Right. And the theoretical analysis reveals that it does. The querying key matrices act as the structural mechanism that enables few-shot learning. They're the pattern recognizers. They're the engine that allows the AI to recognize patterns, infer relationships, and make connections from the novel examples you provide in a prompt. The value matrix, on the other hand, acts as the pathway that encodes the specific target task knowledge. It holds the specialized facts. Okay, so Q and K are the overarching reasoning engine, the logic processor.

6:39And V is the factual database it draws from. If we connect this to the bigger picture, the mathematical proofs in the research show that a pre-trained model's ability to do few-shot learning, its ability to reason from examples, relies almost entirely on the state of the query key matrix. Just that specific part of the structure. Just that part. During pre-training, the model processes trillions of words, and as it does, the core mathematical structure of the query key matrix, specifically the middle block, converges to something called the inverse covariance matrix. Okay, let's translate that concept for a second.

7:12Covariance is basically the underlying structure and relationships in the data, right? Right, how things vary together. It's how different concepts, words, and ideas move together or apart. So if the query key matrix converges to the inverse covariance, it is fundamentally mapping out the deep structural logic of the universe the AI operates in. Yes. It's the ultimate pattern recognition blueprint. And that blueprint is what allows the model to look at three completely novel examples of a new task in a prompt and say, I recognize the structural pattern here. It leverages that deep map of relationships to infer the rules of the new game you are asking it to play.

7:49Which brings us to a massive glaring danger for anyone trying to train an AI. The fine-tuning problem. Exactly. The first major finding. When developers decide to fine-tune a model to get the absolute best zero-shot performance on a specific target task, they typically use full fine-tuning. They unfreeze all the attention parameters, the query, the key, and the value, and update everything to force the model to memorize the new data. Right. They just open the whole thing up. And according to the data we are looking at, doing that actively destroys the few-shot performance. It completely breaks the pattern recognition engine.

8:24The theoretical analysis proves this mathematically. They demonstrate that when you fully fine-tune, as the number of prompt examples increases during a test, the model's error rate doesn't go down. Wait, it doesn't go down? No, it actually converges to an equation representing the absolute baseline irreducible noise plus a massive structural penalty. Mathematically, it converges to sigma squared plus theta zero transposed times sigma times theta zero. Okay, so let's visualize that structural penalty. You are saying that the error rate becomes substantially larger than the baseline zero shot error.

9:00Like, the penalty represents the model actively fighting against its own overwritten reasoning pathways? Exactly that. Even on the exact same target task the model was fine-tuned for. Wait, really? Even on the task it learned? Yes. If you fully fine-tune the model to answer zero-shot questions perfectly about a specific topic, and then you try to assist it by giving it a few examples in the prompt, it gets confused. That's wild. The error spikes giving it examples actually hurts its performance because it can no longer process the pattern. Let's ground that for you, the listener. Why should you care about this structural penalty?

9:35because if you take a brilliant, off-the-shelf AI and fully fine-tune it on your company's proprietary data, you might get a model that knows your specific data perfectly. Sure, it knows the facts. But it has effectively been lobotomized when it comes to general problem solving. If you try to teach it a new trick tomorrow using a few examples, it will fail. You overwrote its underlying map of how things relate to each other in favor of forcing it to memorize a few specific facts. You completely overwrote that inverse covariance matrix. The deep logic is gone, replaced by surface-level memorization.

10:11So if full fine-tuning is a sledgehammer that breaks the model's flexibility, how do developers actually get a model to learn new specific facts without destroying the reasoning engine? Because, I mean, we still need models that know our specific domains without relying on massive expensive prompts every time. Right, and the solution found in the analysis is remarkably elegant. The theoretical proofs show that you can achieve near optimal zero shot performance on a new task while keeping the few shot reasoning capabilities perfectly intact. The surgical fix. The surgical fix. It's value matrix fine tuning.

10:44You freeze the query and key matrices entirely. You just lock them down. You do not let those weights change at all during the fine tuning process. You exclusively update the value matrix. I'm going back to my filing cabinet analogy. Full fine-tuning is like taking your beautiful, meticulously organized filing cabinet, burning the whole thing to the ground, and building a completely new cabinet that only holds one single file. Right. The new task. It's great for quickly grabbing that one file, but the organizational system is ruined for anything else. Value matrix tuning keeps the exact same indexing system, the exact same labels, the exact same folders.

11:23Those are the frozen query and key matrices. Yes. It just takes out the old papers inside the folders and swaps in new papers, the updated value matrix. You update the facts, but you don't touch the organizational logic. What's fascinating here is that the math confirms this dynamic perfectly. By preserving the pre-trained query key matrix, the model maintains its intrinsic few-shot performance across all sorts of different tasks. Its reasoning engine remains completely untouched. But it still learns the new stuff. Yes, because you aggressively updated the value matrix, it still achieves near optimal zero shot performance on the new target task.

12:00It learns the new facts without forgetting how to learn. It seems like the ultimate win-win. Freeze Q &K, update V, get a specialist that is still a generalist. But you know, in the world of machine learning, nobody is ever satisfied with a standard win. Developers always want to push the envelope and squeeze out every single drop of performance. what if they try to optimize that value matrix even further? Well, the researchers anticipated that exact instinct. They asked, what if we fine-tune the value matrix using both zero-shot goals and a little bit of few-shot training at the exact same time?

12:33Here's where it gets really interesting. They want to have their cake and eat it too. They introduce an auxiliary loss. They're essentially telling the model, learn these new facts for zero-shot retrieval, but also tweak the value matrix so you are really, really good at few-shot learning specifically for this one new topic. Right. They wanted to be a generalist, but a super genius few-shot learner in one specific domain. So how did that actually play out? Well, to test the viability of that auxiliary loss, the researchers stepped out of the purely theoretical math and ran an empirical test on a massive large language model.

13:08They used the Quen 2.5 3B Instruct model, and they tested it on the MMLU benchmark. And just for context, the MMLU is a massive, rigorous test of general knowledge across dozens of subjects, right? Everything from history to physics to law. Exactly. It's the standard for general intelligence. And for this empirical test, they decided to fine-tune the model specifically on the humanities category. They wanted to force the model to become a humanities savant. Yes. They applied both zero-shot and few-shot losses exclusively to the value matrix. And the numbers are very revealing. When they used this combined auxiliary loss method, the seven-shot performance on their target task humanities did indeed improve compared to other fine-tuning methods.

13:50Okay, so it worked. It did. It only dropped a minimal 1.77 percentage points from the original untouched base model, which is a highly successful retention of skill. I mean, a drop of 1.77 % is basically negligible. It learned humanity's zero shot, and it stayed incredibly sharp at humanity's few shot. It seems like the auxiliary loss worked perfectly. So where is the catch? The cost is paid somewhere else entirely. This hyper-specialization came at a severe cost to out-of-distribution tasks. Oh, no. Yeah. When they tested that exact same humanities-tuned model on STEM categories using 7-shot prompting, the accuracy plummeted by a massive 6.48 percentage point.

14:30Wow. Wait, so by forcing the value matrix to become overly optimized for humanities reasoning, they actively damaged its ability to reason about science, tech, engineering, and math. Because I sleep. Even though they didn't touch the query and key matrices. Just optimizing the facts inside the value matrix for one domain degraded the facts for another. This raises an important question about the fundamental physical limitations of AI generalization. The theoretical analysis proves this isn't just a quirk of the Quinn model. It is a mathematical certainty. Optimizing an AI for few-shot learning on one specific task mathematically forces it to degrade on mathematically distant tasks.

15:07Explain the mechanism behind that zero-sum game. Why does getting better at Shakespeare actively make the model worse at calculus? How do we even measure that mathematical distance? The distance is quantified using something called cosine similarity. Think of the model's parameters as coordinates in a high-dimensional physical space. Okay. The knowledge required for humanities points in one directional vector, and the knowledge required for stem points in a very different direction. Right. They're conceptually far apart. Exactly. And the physical space inside the neural network's weights is finite.

15:42If you use an auxiliary loss to actively drag the parameters in the value matrix toward the humanities vector to achieve few-shot perfection there, you are literally pulling them out of the alignment required for stem logic. You are stretching the blanket to cover your shoulders and your feet get cold. That's a perfect analogy. There is only so much parameter space to go around. If the concepts are mathematically distant, forcing perfection on one inherently drags down the other. The auxiliary loss makes you an absolute genius in your chosen neighborhood, but it aggressively shrinks the borders of your overall capability.

16:17The math is entirely unforgiving in that regard. There is always a trade-off when you demand hyper-specialization, even when you use surgical techniques like value matrix tuning. So let's step back and look at the whole journey we've just taken through this research. We started with the tug of war between prompting and fine-tuning. We saw that full fine-tuning is essentially a sledgehammer. It crushes an AI's underlying logic blueprint, that inverse covariance matrix, just to get you cheap zero-shot performance. Right, it destroys the general level. And value matrix fine-tuning, on the other hand, is a scalpel.

16:49It preserves the reasoning engine, the query and key matrices, while exclusively updating the facts in the value matrix. But even with a scalpel, if you get greedy and push too hard for few-shot perfection on one specific topic using an auxiliary loss, you will physically pull the parameters away from other domains, costing you heavy performance drops in unrelated fields. You preserve the engine, but you warp the database. So what does this all mean for you listening right now? It means that when you are deploying or utilizing specialized AI in your business, you must be acutely aware of what the model might have surrendered in order to learn its new skills.

17:28An AI fine-tuned perfectly to generate your company's marketing copy might have just lost its ability to reliably analyze a basic spreadsheet if the tuning process wasn't handled with surgical precision. You cannot treat fine-tuning as a magic bullet that only adds knowledge. It inherently shifts the balance of the entire system. And, you know, the strict distinction between the query key mechanism and the value pathway leaves us with a provocative possibility to consider. So, what's that? Well, if changing the value matrix alters the facts, but the query and key matrices dictate the reasoning, could we eventually see a future where we stop building massive monolithic AI models from scratch every few months?

18:09Wait, you mean like modular AI? Exactly. Could we simply buy off-the-shelf permanent reasoning modules, perfected query and key structures that map the fundamental logic of the universe, and then plug them into modular, swappable fact databases, the value matrices, depending on the task of the day? Oh, wow. Swapping out the monopoly rules for the chess rules without ever losing the fundamental knowledge of how to sit at the table, roll the dice, and play a game. Exactly. The filing cabinet stays firmly in place. We just get better and better at instantly swapping out the files. That is an incredible thought to end on.

18:43Thank you for joining us as we unpacked the math, the mechanics, and the undeniable trade-offs behind how AI actually remembers and forgets. We'll catch you next time on the Deep Dive.

From the publisher

This research examines the tension between in-context learning (ICL) and fine-tuning in Transformer-based models, specifically using linear attention to provide a theoretical foundation. While fine-tuning is often employed to enhance zero-shot performance on specific target tasks, the authors demonstrate that updating all attention parameters can inadvertently damage the model's ability to learn from demonstrations. They identify a superior strategy: restricting updates to the value matrix, which improves task-specific accuracy while maintaining the model’s original few-shot capabilities. The study further explores the use of an auxiliary few-shot loss, finding that it boosts performance on the target task but reduces the model's ability to generalize to out-of-distribution tasks. These theoretical insights are validated through both mathematical proofs and empirical experiments on the MMLU benchmark. Ultimately, the work provides a framework for optimizing language models without sacrificing their inherent flexibility as in-context learners.

More from Best AI papers explained

All 475 episodes
Fine-Tuning Strategies for Preserving In-Context Learning in Linear AttentionBest AI papers explained · 19 min
Listen in VO