Learning dynamics of LLM finetuning

9 Oct 2025 · 12 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Learning dynamics in LLM fine-tuning—how SFT and DPO change predictions across unrelated examples, causing hallucinations and repetition via an ENTK-driven “squeezing effect.”

Guests

No guests mentioned; it’s a single host (“Deep Dive”) summarizing a research paper.

Key claims

SFT spreads probability to “mathematically nearby” responses (ENTK similarity) but softmax creates global pushdown pressure that later suppresses similar alternatives. Hallucinations arise because some “preferred” answer styles/facts are reinforced strongly during their own training, surviving pushdown and leaking into other questions. In DPO, negative gradients applied to already-unlikely rejected responses squeeze probability mass into the argmax token (“rich get richer”), worsening degeneration/repeaters.

Notable examples

MNIST-style analogy; physics vs history questions; repetitive boilerplate/argmax token reinforcement; reported Xtend DPO win rate up to 69.28% using ChatGPT/Claude-3 judges.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Learning Dynamics

0:45 to 1:21

Explore the concept of learning dynamics in LLMs and their implications.

“It gives us this framework, like a lens, to interpret all those weird, sometimes counterintuitive things we see when we try to align these big models.”

Influence Spread in LLMs

1:21 to 2:22

How training on specific examples affects model predictions through similarity.

“Maybe start with the basics of how influence spreads.”

Global Pressure in Supervised Fine-Tuning

2:22 to 3:47

The impact of probability budget constraints in LLM training.

“Even though you didn't show it a nine in that specific update, it's like probability mass spreading through similarity.”

Hallucinations Explained

3:47 to 5:26

Insights into how unrelated training data can falsely boost confidence.

“Okay, that makes sense for why similar but wrong answers might go down.”

Dynamics of Preference Tuning

5:26 to 7:45

Examining how DPO training affects answer probabilities in LLMs.

“You ask about, say, politics, and it might spit out an answer structure or even a fact fragment.”

Addressing the Squeezing Effect

7:45 to 10:01

Proposing a method to mitigate the squeezing effect in LLMs during DPO.

“off-policy DPO, that rejected answer, U-minus, is often already quite unlikely based on the SFT model.”

Experimental Validation of Xtend Method

10:01 to 11:11

Results showing the effectiveness of the Xtend method in improving model performance.

“Yeah, the experimental results seemed to back it up really well.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're moving past just looking at the final scores of LLM's. You know, the benchmarks, we're actually digging right under the hood. We want to understand the mechanics of how these large language models learn when they get fine-tuned. If you've worked with LLMs, you know the terms, right? Supervised, fine-tuning, SFT. And preference tuning, like DPO, direct preference optimization. So our mission today for you, the listener, is to kind of shift focus. Stop looking only at the final model and start seeing the step-by-step process. You know, how knowledge actually shifts, where it builds up, and crucially, where it can go sideways.

0:35step by step. That's really the core of it. We're diving into what the researchers call learning dynamics. It sounds complex, but simply put, it just describes how learning from one specific training example ends up changing the model's prediction on other examples. All the others, really. It gives us this framework, like a lens, to interpret all those weird, sometimes counterintuitive things we see when we try to align these big models. And the research we're looking at today points to two really big things driven by these dynamics. First, we get a much clearer picture of why certain hallucinations pop up after SFT.

1:07And second, we uncover this mechanism, they're calling it the squeezing effect, which explains why trying to correct the model with DPO can sometimes make the answers you want actually become less likely, which sounds totally backwards, right? Okay, let's unpack this. Maybe start with the basics of how influence spreads. To really get how this works in something as complex as an LLM, maybe a simpler analogy helps. Let's think about a classic machine learning task, image classification, like recognizing handwritten digits, the MNIST data set. So if we train the model on an image of a handwritten 4, it's pretty obvious the model gets more confident that it's a 4, right?

1:43That's the direct effect. Right. That's the local update. But the interesting card is the global, maybe unintended effect. because of the, well, the deep mathematical structure of these networks, the researchers use something called the empirical neural tangent kernel, or ENTK, as a proxy for this. That update on the 4 doesn't just stay local. The ENTK acts kind of like a similarity measure. So training on 4 also indirectly pulls up the confidence for other inputs the model sees as similar. Ah, okay. So because maybe some handwritten 4s have strokes that look a bit like handwritten 9s. The model sees that similarity.

2:17and training on the four gives the nine a little confidence boost for free. Pretty much, yeah. Even though you didn't show it a nine in that specific update, it's like probability mass spreading through similarity. And that same idea, that spreading through similarity, is exactly what's happening when we move to LLM supervised fine-tuning, SFT. In SFT, we're essentially doing the same thing, applying a positive gradient to boost the probability of the answer we want. Let's call it Y +, the chosen response. So initially, that positive push does what you'd expect. And because of that ENTK similarity thing, it slightly increases the confidence of other responses that are mathematically nearby.

2:56This could be answers you might later reject in DPO, maybe the I - response, or even just good rephrases of the answer you want. They all get a tiny initial lift. Okay, but training doesn't happen in a vacuum. There's that soft max layer at the end, right? It forces all the probabilities for the next word, for all possible words, to add up to one. It's a fixed budget. So if the answer you want keeps grabbing more and more of that probability budget, something else has to lose out. It's exactly right. This creates what the paper calls a global push down pressure. It affects everything that isn't the main target of the update.

3:30And that's why those similar responses, the ones that got that little initial boost, eventually start to decrease in probability again later in SFT. Even though you never directly told the model they were bad, the global pressure just forces them down to make room for the main answer. Got to balance the books. Okay, that makes sense for why similar but wrong answers might go down. But here's where I thought the research got really interesting. This dynamic seems to explain hallucination. The researchers specifically track the confidence of a particular kind of answer. The preferred answer for a completely unrelated question also in the training data.

4:05So like question A is about physics, question B is about history. Both are in the SFT dataset. And what they saw was kind of surprising. When you train the model on the good answer for question A, the confidence in the good answer for question B also keeps increasing throughout the SFT process, even when you're not training on question B at that moment. Wait a second. If there's this big global push down pressure squeezing out similar wrong answers, why isn't it pushing down the answer from a totally unrelated question? Why is that one still climbing? Ah, well, it's like a tug of war between gradients.

4:37That unrelated answer, the question B answer, it gets reinforced so strongly and so often during its own training examples throughout the whole SFT process that this positive reinforcement is actually strong enough to consistently overcome the global pushdown pressure from other updates. The LLM essentially learns this very robust fingerprint, a preferred style phrasing structure. It's quite fascinating. The paper mentions that responses generated by models like chat GPT, even if they mean different things, often look very similar mathematically to the LLM itself. Okay, so the hallucination mechanism is the model has these highly reinforced facts or phrasings from its training data.

5:17Yeah, like islands of high probability. And they're so strong, so sticky, that they kind of leak into answers for other questions because the global pressure isn't enough to suppress them. That's the idea. You ask about, say, politics, and it might spit out an answer structure or even a fact fragment. It learned really well from a geography question in training just because both share that ingrained high probability fingerprint. Wow. Right. Now let's switch gears a bit to preference tuning. Think DPO. So here the game changes. We're not just pushing up on the good answer Y plus anymore. We're explicitly contrasting it with a rejected answer, Y minus.

5:53So DPO introduces both that positive gradient push and a significant negative gradient arrow aimed right at that rejected answer, trying to suppress it. Now here's where they found that really counterintuitive thing, especially with off-policy DPO, where you use a fixed data set of good-bad pairs collected beforehand. What they observed is that the confidence of both the chosen answer Y plus and the rejected answer Y - often starts to gradually decrease during DPO training. And this effect was apparently worse if the model had been through a really long SFT phase first. Okay, that just feels wrong.

6:24If we're actively optimizing for the chosen answer, why would its probability drop? And if we're pushing down the rejected one, where is all that probability going? Is the whole distribution just deflating? This is exactly where they introduced the squeezing effect. It seems to be a mechanical consequence of how softmax works. It happens specifically when you apply a large negative gradient that pushed down to a token prediction that is already very unlikely, something sitting in a low probability region, which you could call a valley in the probability landscape. So we're hitting something with a negative penalty that the model already thought was pretty unlikely to begin with.

7:02Precisely. And the math of the softmax layer handles this in a peculiar way. When you remove probability mass from that deep valley, it doesn't just spread out nicely among all the other options. Instead, it gets disproportionately squeezed into the single token that had the highest probability before the update. The argmax of tokens, think of it like the rich get richer effect, but for probability distributions. If you take away from the poorest option, the probability mass flows most easily to the richest, the highest peak. Whoa. So the negative feedback we're giving to punish the bad answer actually ends up accidentally reinforcing whatever the model's top default guess was anyway, even if that top guess is just some generic repetitive phrase.

7:42That seems to be the direct consequence, yeah. Especially off-policy DPO, that rejected answer, U-minus, is often already quite unlikely based on the SFT model. It's already in that valley. So when DPO hits it with a negative gradient, the distribution becomes extremely peaky right at the argmax token. And this peakiness, this concentration, strongly reinforces the model's existing biases or preferred structures. Gives us a really solid mechanical explanation for that awful text degeneration or repeater phenomenon you sometimes see. You know, where fine-tuned models start spewing out the same simple boilerplate phrases over and over.

8:15They get mathematically locked onto that single highest probability next token because of the squeezing. Oh, man. I remember seeing that in some early fine-tuned models. They'd write like five paragraphs, but it was basically the same sentence structure repeated. This finally gives a why. Okay, so the technical problem is DPO's negative gradient hits a low probability valley, the distribution squeezes, and it over-reinforces the default top prediction. So the big question is, if we know that's the dynamic, how do we prevent hitting that valley when DPO starts? Can we fix the valley beforehand?

8:48Exactly. And the analysis pointed towards a surprisingly simple fix. They call it SFT data augmentation, or the extend method. Sounds a bit crazy at first. The idea is to modify the SFT phase before you even start DPO. Instead of only doing SFT on the good answer Y +, the Xtend pipeline suggests you temporarily SFT on both the good answer Y +, and the bad answer Y -. Hang on. You're saying to make the bad answer less likely during DPO, the fix is to first train the model during SFT to make it more likely. That seems completely backwards, like you're training against your final goal. I know. It totally feels counterintuitive.

9:22But think about the dynamics we just discussed. By deliberately using SFT to pull that rejected answer, Y - out of that deep probability valley, you ensure that when DPO does start later and applies that negative gradient to Y-, it's hitting a response that now has higher initial confidence. It's hitting a peak, or maybe a slope, but crucially, not that deep valley floor. This gives the probability mass more room to redistribute when the negative gradient hits. It helps restrain that harmful squeezing effect because the mass doesn't just instantly funnel into the single argmax prediction. It can spread out a bit more normally.

10:00Okay, that makes a strange kind of sense dynamically. Did they test this Xtend idea? Did it work? Yeah, the experimental results seemed to back it up really well. First, they looked at the learning dynamics again. And just as the theory predicted, the confidence decay for other irrelevant responses during DPO was slower when they used the Xtend method first. And maybe even more tellingly, they tracked the squeezing effect directly. Using a metric focused on that top prediction, the argmax confidence. The model that went through the extend SFT first showed this immediate sharp drop in argmax confidence right when DPO started, which confirms the squeezing was being weakened.

10:36The negative gradient wasn't just piling onto the default top choice anymore. That's pretty compelling evidence for the mechanism. But what about the bottom line? Did the final model actually perform better? It seems so. When they put the final models head to head, the standard DPO versus the Xtend DPO and used external AI judges like ChatGPT and Cloud3, the Xtend model won by a pretty significant margin. They reported win rates up to nearly 70 percent, like 69.28 percent. So understanding and fixing that structural squeezing problem actually led to a measurably better, more robust and definitely less repetitive model.

11:11Right. So I guess the big takeaway for you, the listener, is that fine-tuning these LLMs isn't just like tweaking a few knobs. It's this really dynamic process. These hidden connections, the NDK similarities can unintentionally create hallucinations during SFT. And then the basic mechanics of the softmax layer can lead to this damaging squeezing effect during DPO, making models repetitive. Understanding these under-the-hood dynamics seems crucial to avoid these traps. Absolutely. And, you know, this whole analysis around the squeezing effect, it might have implications way beyond just LLM fine-tuning.

11:40Think about it. The core issue is applying a strong negative gradient to an already unlikely outcome when you have a soft max output layer. That mathematical reality isn't unique to LLMs. It makes you wonder, could this squeezing effect be unknowingly hurting performance in other areas of deep learning? Places that also use negative gradients, like maybe machine unlearning, where you're trying to force the model to forget something specific. Or adversarial training, pushing the model away from bad outputs. How many other hidden pitfalls might this mathematical quirk be creating across AI? And what other simple dynamics-aware fixes like the Xtend method are just waiting out there to be found?

12:17Something to think about.

From the publisher

This academic paper presents a novel framework for understanding the evolution of Large Language Models (LLMs) during finetuning by analyzing their learning dynamics from a dynamical perspective, contrasting with previous approaches focused on training targets or end-states. The authors formalize the change in model prediction using a decomposition into three key terms, which adapts to various finetuning algorithms like Supervised Finetuning (SFT) and Direct Preference Optimization (DPO). A significant finding is the "squeezing effect" caused by negative gradients during preference tuning, which reduces the confidence of most responses and is especially pronounced when the model is already confident or finetuning is off-policy. The framework is validated through experiments on both the MNIST dataset and LLM finetuning, demonstrating its ability to explain counter-intuitive phenomena like the confidence decay observed in DPO. Finally, the research inspires a simple yet effective method to improve alignment performance by mitigating the harmful aspects of the squeezing effect.

More from Best AI papers explained

All 475 episodes
Learning dynamics of LLM finetuningBest AI papers explained · 12 min
Listen in VO