Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models

16 Mar 2026 · 23 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Explains why long-form generation derails (teacher forcing + distribution shift) and why common fixes like RLVR can “hack” rewards, then introduces energy-based fine-tuning (EBFT) that matches sequence-level feature statistics in a frozen feature space.

Guest backgrounds

No guest identities or bios are provided in the transcript; two hosts discuss the research.

Key claims

Teacher forcing trains models on perfect prefixes, so small early errors snowball as conditional entropy/confusion grows with length. RLVR optimizes only final verifiable rewards, degrading validation cross-entropy and producing unhinged outputs. EBFT uses a frozen feature network at 25/50/75% layer depths, with alignment + diversity rewards and whitening to decorrelate features.

Notable examples

Coding: RLVR interleaves conversational prose inside Python. Translation: RLVR “multilingual runaway” repeats English, appends language tags (Spanish/Portuguese), inserts random words, and truncates mid-word. EBFT reports validation cross-entropy dropping 0.289 to 0.207 on a 1.5B model and maintains coherence at 64 tokens despite training on 8-token rollouts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI's Drift in Long Responses

0:45 to 2:12

Exploration of why AI models struggle with coherence in lengthy outputs.

“It's called energy based fine tuning or EBFT for short.”

Teacher Forcing and Its Implications

2:12 to 4:20

Discussion on teacher forcing and its effects on an AI's learning capability.

“It is constantly guided by flawless human data at every single microscopic step.”

Distribution Shift and Its Challenges

4:20 to 6:39

Examining the concept of distribution shift and its impact on AI performance.

“They look at the expected conditional entropy.”

Current Fixes: RLVR Band-Aid Approach

6:39 to 9:39

Overview of the RLVR method and its limitations in AI training.

“When you compare the failure modes of a standard supervised fine-tuned model and SFT against this RLVR band-aid, the difference in how they fail tells you exactly what is going wrong under the hood.”

Introducing Energy-Based Fine-Tuning (EBFT)

9:39 to 12:44

An introduction to EBFT and how it addresses AI training issues.

“And this is where we introduce the breakthrough, right?”

Mechanics of EBFT and Computational Efficiency

12:44 to 14:01

Explanation of how EBFT operates and its computational advantages.

“So the AI being trained isn't being told what specific words to use at all.”

Understanding Strided Block Parallel Sampling

14:01 to 15:00

Learn about how strided block parallel sampling enhances computational efficiency in AI text generation.

“But to make it computationally feasible, that use a specific technique called strided block parallel sampling.”

The Reward Signal in EBFT

15:01 to 15:46

Discover the nuanced reward structure in energy-based fine-tuning and its importance.

“of potential answers, sharing the computational load of the prefix across all of them.”

The Concept of Whitening in AI

15:47 to 17:41

Understand the mathematical process of whitening and its significance in feature matching.

“Without it, the model might find one safe, generic way of structuring a sentence and just use it for everything, collapsing into repetitive generation.”

Impact of Whitening on Model Performance

17:42 to 19:26

Learn how whitening and feature matching improve AI performance and coherence.

“By flattening the EQ, it ensures that all semantic features, the loud obvious bass line of the main topic and the quiet, subtle vocals of the underlying logic are heard and evaluated equally by the model.”
Show all 13 chapters

Practical Implications of EBFT

19:27 to 21:11

Explore how energy-based fine-tuning can be applied without unit tests and its scalability.

“And it also completely mitigates the snowball effect of distribution shift that we discussed.”

Future of AI Tools with EBFT

21:12 to 22:27

Discuss the future implications of EBFT on AI tools and their reliability.

“The research tested this architecture across multiple sizes of the Quinn 2.5 model from 1.5 billion up to 7 billion parameters and the Llama 3.21 billion model.”

A Philosophical Shift in AI Communication

22:28 to 23:24

Ponder the deeper implications of AI learning to communicate through concepts rather than words.

“Then we saw the destructive nature of rigid outcome-only rewards that turned models into test-hacking robots speaking gibberish.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Have you ever noticed how an AI model starts off answering a prompt brilliantly, but then slowly it loses the plot. Oh, yeah. All the time. Like you ask it to write a script or, you know, solve a really complex coding problem. And those first few lines are just absolute genius. Right. It looks perfect at first glance. Exactly. But then by paragraph four, it's wandering completely off topic or worse. It's just completely hallucinating facts that make zero sense. It is an incredibly common phenomena. And honestly, if you watch the output generate in real time, it's literally like watching a train slowly derail.

0:36Yes. Because the longer the output gets, the more confused the underlying system actually becomes. Well, today you and I are going on a deep dive into a cutting edge breakthrough in AI training that actually explains the mechanics of why this happens. It's called energy based fine tuning or EBFT for short. Yeah, EBFT is huge. It really is. So the mission for this deep dive is to explore exactly why current AI models struggle so much with long-form generation and why our current engineering fixes for this problem are, well, fundamentally breaking the AI's general intelligence. Which is a massive problem right now in the industry.

1:14Totally. And then we're going to get into how this new EBFT method solves the issue by completely changing what the AI actually pays attention to when it learns. Okay, so let's unpack this from the very beginning. Sure. To understand why EBFT is such a breakthrough, we first really need to understand why these incredibly advanced multi-billion parameter models drift off track in the first place. Because, I mean, when you look at the underlying architecture, it seems brilliant, but there's a hidden trap in how they are initially trained. Right. And the research actually refers to this trap as teacher forcing during cross-entropy training.

1:50Teacher forcing. Okay, break that down for us. So let's look at how an AI learns to predict the next word or the next token. During its standard training phase, the AI is always given the perfect, absolute ground truth prefix to predict that next token. Okay, so it has perfect context. Exactly. If the target sentence is, the cat sat on the mat, the AI is fed, the cat sat on the, and it just has to guess the word mat. Got it. It is constantly guided by flawless human data at every single microscopic step. So it's essentially like learning to drive a car with an incredibly nervous driving instructor.

2:30That is a great way to put it. Like you're behind the wheel, right? But the second you drift even a fraction of a millimeter out of your lane, the instructor just grabs the wheel and nanks you back to dead center. Right. They never let you figure it out. Exactly. You never actually learn how to correct a mistake yourself because the instructor, the teacher forcing, never lets you make one in the first place. Which creates a massive vulnerability when you finally take your driving test. Because when the AI is deployed in the real world to generate a long response for you, the instructor is gone.

2:58They've left the car. They're out of the car. And the AI now has to condition its next prediction not on perfect human data, but on its own generated text. But wait, let me push back on that for a second. Sure. If an AI is trained on terabytes of perfect data, shouldn't it just inherently mimic perfect data? Like if I'm trained perfectly by that driving instructor, I should theoretically drive perfectly on my own, right? You would definitely assume so. But this is where we run into a mathematical reality that the data calls distribution shift. Distribution shift. Yeah. The AI is only perfect locally, like one microscopic step at a time.

3:34Let's say it makes one tiny, almost imperceptible error early in a generated sequence. Okay, like a slight typo. Or just, you know, maybe it chooses a slightly awkward verb in paragraph one. Oh, I see. That awkward verb now becomes part of the context for the next word. Suddenly, the model finds itself in a state it was never explicitly trained on. The probability distribution of what comes next gets a little flatter, you know, a little less certain. So it's basically an unfamiliar territory. Exactly. And because it doesn't know how to self-correct, because of that teacher forcing, it makes another slightly weird choice.

4:09Right. And one weird word leads to a weirder sentence, which leads to a totally unhinged paragraph. The snowball effect. Precisely. And the research quantifies this beautifully, by the way. They look at the expected conditional entropy. Which is what exactly? It's essentially a mathematical measure of the AI's internal confusion. Oh, wow. OK. Yeah. And they found that even in supposedly perfect models with very low perplexity, this confusion naturally and inevitably grows as the completion length increases from one token to 64 tokens. So by the time it hits 64 tokens. By the time it hits 64 tokens, the model is exponentially more confused than it was at token number one.

4:49It's literally compounding its own errors. Exactly. Wow. Okay, so if standard next token training fails over these long sequences because of that teacher forcing, how are engineers currently trying to fix it? Because obviously we have AI right now that can write more than 64 words without entirely breaking down. Right. We use various optimization techniques to try and patch it up. Like a Band-Aid. Basically, yeah. The current industry band-aid for this is a method called RLVR, which stands for reinforcement learning with verifiable rewards. Verifiable rewards. Yeah. If token-by-token training is too narrow and microscopic, RLVR goes to the absolute opposite extreme.

5:28It ignores the token-by-token journey entirely and just grades the final output. Which logically sounds like standard reinforcement learning. I mean, if I ask the AI to write a Python script, I don't really care how it gets there as long as the code runs, right? Right. The end justifies the means. Exactly. So we just give it a pass or fail grade at the very end. If it passes a unit test, it gets a reward. If it's doing language translation, we judge it by a high BLEU score, which is a metric for translation quality. We just optimize for the destination. That is the theory, yes. But relying solely on verifiable rewards comes with a severe hidden cost.

6:05Always a catch. There's always a catch. The data reveals that while RLVR makes the AI exceptionally good at hacking those specific tests, it severely degrades the model's validation cross-entropy. Meaning what exactly? Like in plain English? Meaning it literally degrades the model's fundamental understanding of human language. Wait, really? Yeah. It forgets how to speak normally in its pursuit of the reward. So it basically becomes a test-taking robot that doesn't actually understand the subject it's being tested on. Exactly. It just learns the trick to pass. The findings on this specific degradation are honestly fascinating.

6:42When you compare the failure modes of a standard supervised fine-tuned model and SFT against this RLVR band-aid, the difference in how they fail tells you exactly what is going wrong under the hood. Oh, the difference is night and day. The SFT failures are usually quite predictable. Right. The SFT fails because it just misses subtle prompt logic. Like if a prompt asks the standard AI to count overlapping substrings in a piece of code, it will just jump forward and miss the overlaps completely. Yeah, it's a superficial error. Or if asked to find the greatest integer that meets a certain mathematical condition, it just returns the first integer it finds and then it just stops.

7:20Right. The syntax is perfectly fine. The grammar is fine, but it's logically lazy. Right. It is fundamentally sound in its structure, but it lacks the deeper reasoning constraints because it was only trained to guess the next likely token. Exactly. But when you look at how the RLVR models fail. Oh, man, it is spectacular. It really is. The RLVR failures aren't just lazy. They are completely unhinged. Yeah, they go entirely off the rails. The data shows that in coding tasks, the RLVR model will just randomly interleave regular conversational prose right in the middle of a Python script. Which is wild.

7:56It's insane. It'll just start chatting, which obviously breaks the code so it won't even compile. And my absolute favorite is the translation benchmark. Oh, the MTNT benchmark. Yes. They tested this on the MTNT benchmark, translating from English to French. And the RLVR model suffers from what researchers call multilingual runaway. Multilingual runaway. It is a profound glitch in the system. It really is. It's supposed to translate an English sentence to French, right? Instead, the AI repeats the English source sentence verbatim, then it randomly appends a tag that says Spanish colon, follows that with a Portuguese colon tag, spits out three completely random words, and then just truncates the output mid-word.

8:35Just completely crashes. It's like the AI's brain short-circuited. I mean, why does a machine suddenly start typing Spanish out of nowhere? Because it is blindly trying to maximize a rigid reward signal without any underlying semantic comprehension. Ah, so it's hacking the test. Exactly. The RLVR optimizer likely found some bizarre mathematical loophole. Perhaps the translation evaluation metric didn't heavily penalize extra text as long as certain target keywords were present. Oh, that makes sense. Or perhaps the model figured out that generating specific language tags artificially inflated a secondary reward metric during training.

9:11It found a shortcut that technically satisfied the math of the reward function, but it resulted in total linguistic gibberish in reality. So we are essentially caught between a rock and a hard place here. Very much so. Token-by-token training is too microscopic and leads to that snowball effect of errors, but then our LVR rewards are too broad, too rigid, and they turn the model into a hallucinating test hacker. Which brings us to the necessity of finding a middle ground. If looking at single words is too narrow and looking at the final test grade is too broad, how do we teach a model to look at the in-between?

9:46And this is where we introduce the breakthrough, right? Energy-based fine-tuning, EBFT. Yes, this is where it gets really exciting. Instead of matching individual tokens or chasing a pass or fail reward at the very end, EBFT matches sequence-level statistics in a hidden feature space. That is the core innovation right there. It completely shifts the target of what the AI is actually trying to learn. I want to build an analogy here just to make this concrete for you listening. If we think about judging a piece of music. Okay, I like this. Cross entropy training, that token by token method, is like judging a musician solely by whether every single individual note is perfectly identical to the sheet music.

10:24One long note and you fail, even if the song overall sounded okay. Right. And following that logic, RLVR is like ignoring the performance entirely and just waiting to see if the crowd clapped at the very end. Yes. The system doesn't care if you played the song on a beautifully tuned piano or if you just screamed into a kazoo. As long as the applause meter hits a certain decibel level, the system registers it as a perfect performance. Exactly. So where does EBFT fit into that? If it's matching sequence level statistics in a hidden feature space, it's not looking at the raw notes and it's not looking at the applause.

10:59No, it's looking deeper. It's like listening to the musician and judging if the vibe, the rhythm and the emotional arc match a master recording. It's judging the deep underlying structure of the song. That is highly accurate. And the mechanism they use to actually perform this vibe check mathematically is something called a frozen feature network. A frozen feature network. Okay, how does that operate under the hood? What is it actually looking at? So when the training process begins, they take a copy of the base AI model and they completely freeze it. Meaning it can't learn anymore. Right, right.

11:33Its internal parameters can no longer be updated or changed. This frozen model acts as the immovable judge. Okay. When the AI being trained generates a response, this frozen judge doesn't look at the final text output. Instead, it looks at the intermediate activations inside the neural network. Intermediate activation. Essentially, it is looking at the internal thoughts of the network as the text is being formed. Oh, wow. So it's reading the model's mind while it writes. Exactly. It's peeking under the hood. And it extracts these internal representations at three specific depths. 25%, 50%, and 75 % through the network's layers.

12:11Wait, why those specific depths, though? Why not just look at the final layer right before the text is generated? Because neural networks process information hierarchically. The early layers around that 25 % mark are primarily focused on low-level grammar, syntax, and basic sentence structure. Building blocks. Exactly. Then, by the time you get to the 50 % depth, the network is processing broad semantic meaning. It's understanding the actual concepts being discussed. Okay, and the 75 % mark. At 75%, it is handling complex structural logic, like how an argument is built or how a mathematical proof flows.

12:50So the AI being trained isn't being told what specific words to use at all. Not at all. It's being told to make sure its internal feature embeddings, its vibe across grammar, semantics, and logic perfectly match the feature embeddings of the ground truth human data. You nailed it. It's literally learning to think in the same structural patterns as the perfect data, rather than just blindly copying the surface-level output. And by doing this, it forces the model to align with the sequence-level statistics of human language. It learns the conceptual rhythm of how a good answer flows, which directly prevents that distribution shift we talked about earlier.

13:29Ah, because the foundation is stronger. Exactly. Even if it uses a slightly different word, its internal semantic compass remains aligned. Okay, conceptually, feature matching sounds brilliant. But from an engineering standpoint, how on earth is this actually calculated? It is a massive technical challenge, yeah. I mean, extracting the vibe of a generated sequence at three different neural depths, comparing them, and updating the model sounds incredibly computationally heavy. How do they do this without completely melting the GPUs? That is where the actual engine of EBFT comes into play. They utilize a reinforced-style gradient update.

14:03But to make it computationally feasible, that use a specific technique called strided block parallel sampling. Strided block parallel sampling. Okay, break that down for me. Because sampling multiple rollouts usually takes massive amounts of compute. How does this specific method save computational power? To understand the efficiency, you have to look at how an AI generates text. Calculating the context, what we call the prefix, takes a significant amount of memory and processing power. Right. If you want the AI to generate eight different possible endings to a paragraph to see which one has the best features, standard methods would require you to process the beginning of that paragraph eight separate times.

14:41Which is incredibly redundant. Highly redundant. Strided block parallel sampling eliminates that entirely. It calculates the prefix once, holds that massive context in memory, and then branches out. Oh, I see. It generates concurrent rollouts from nested prefixes simultaneously. So in a single forward pass, the system creates a rich, diverse dataset of potential answers, sharing the computational load of the prefix across all of them. So it's rapidly generating all these different branches of thought at once. How does it score them? I'm assuming the reward isn't just a simple thumbs up or thumbs down anymore.

15:15No. The reward signal in EBFT is beautifully nuanced. It actually consists of two distinct parts. The first is an alignment term. Which means what? This simply measures how close the generated features are to the ground truth. Does the semantic vibe match? But the second part is just as critical. A diversity term. Hmm. Why do we need a diversity term? It acts as a penalty, essentially. It actively penalizes the model if it just repeats the exact same sequence of features over and over again. Oh, to stop it from getting stuck in a loop. Precisely. Without it, the model might find one safe, generic way of structuring a sentence and just use it for everything, collapsing into repetitive generation.

15:57The diversity term forces the AI to be both conceptually accurate and structurally creative. Now, there is a highly technical concept mentioned alongside this process that I need you to translate for me. I think I know which one you're going to say. The data shows that for this feature matching to actually work, they have to use a mathematical process called whitening. Yes, whitening. They state that without whitening, the whole process gets dominated by correlated directions. What exactly is whitening? To understand whitening, we have to look at the second moment matrix of these features. Okay, losing me slightly.

16:30Let me simplify. In simple terms, this matrix represents how different semantic features cover or overlap with each other. In human language, some concepts naturally appear together very frequently. Like the words king and royal. Exactly. Because they are so highly correlated, their combined signal in the feature space becomes massive. Right. They create a very loud direction mathematically. Okay, loud directions. If you don't adjust for this, the loss function of the AI will obsess over minimizing that one massive vector. It will only focus on matching the loudest, most obvious features, like the main subject of the sentence.

17:06It will completely ignore the tiny orthogonal vectors that represent subtle structural logic or nuanced grammar. So if we bring back the music analogy, it's like listening to a poorly mixed track where the bass line is so incredibly loud that it literally rattles the speakers. Exactly. You literally cannot hear the acoustic guitar or the subtle backing vocals because the bass is just dominating the entire soundscape. That is exactly what happens in the neural network. Whitening is the mathematical process of decorrelating these features and scaling them to unit variants. So whitening is basically flattening the EQ on the stereo.

17:41Precisely. By flattening the EQ, it ensures that all semantic features, the loud obvious bass line of the main topic and the quiet, subtle vocals of the underlying logic are heard and evaluated equally by the model. That makes total sense. It forces the AI to align with the entire conceptual landscape, not just the most dominant signals. Okay. The math is elegant. The mechanism makes total sense, but let's get down to the actual result. The fun part. Right. Does flattening the EQ and matching features actually fix the AI for the end user? Because if it doesn't solve the long-form hallucination and the coding failures we talked about earlier, it's just a fun theoretical exercise.

18:20Well, the empirical results are what elevate this from theory to a true breakthrough. EBFT essentially shatters the historical trade-offs we've seen in AI training. Let's look at the specifics of that. When they tested downstream accuracy, which is, you know, how well the model actually performs on real-world tasks like the human evil coding benchmark or the MBPP EBFT flat-out outperforms standard supervised fine-tuning. Yes, it writes better code and translates languages more accurately. But the most crucial metric is what happens to the validation cross-entropy. Because remember how the RLVR band-aid caused the model to lose its fundamental grasp of language?

18:56Right, the RLVR models forgot how to speak normally and started hallucinating Spanish tags in the middle of French translations. Exactly. Well, EBFT does the exact opposite. It actually improves the validation cross-entropy. Wait, it improves it? Yes. When tested on a 1.5 billion parameter model, EBFT dropped the cross-entropy loss from 0.289 down to 0.207. Wow. That is a massive drop in AI terms. It's huge. So it's not a trade-off anymore. The model gets significantly smarter at specific complex tasks, and it becomes a fundamentally better, more coherent overall language model at the exact same time.

19:34And it also completely mitigates the snowball effect of distribution shift that we discussed. Oh, right. The confusion compounding over 64 tokens. Exactly. The data shows that EBFT achieves the lowest feature matching loss even on those long 64 token completions. And what is truly remarkable about that is the model was only trained using short 8 token rollouts. Wait, seriously. If it was only trained on bursts of 8 tokens, how is it remaining coherent at 64 tokens? Because it didn't just memorize a sequence of words. It internalized the underlying conceptual structure. Ah, the vibe. The vibe. Once an AI truly understands the structural logic of a good answer, it can naturally maintain that coherence over much longer distances without needing explicit guidance for every single step.

20:20That is incredible. There is also a huge practical implication here regarding how we gather data in the first place. Very true. The findings emphasize that EBFT works in what they call non-verifiable settings. With RLVR, you absolutely need a unit test, right? You need a rigidly defined pass or fail metric. Right. But what if you are training the AI on massive amounts of unstructured coding scraped directly from GitHub? There are no unit tests attached to that code. None. RLVR is completely useless in that scenario because it has nothing to verify against. But EBFT doesn't need a unit test. Not at all.

20:55It just needs to look at the code, extract the semantic features, and learn to match the conceptual vibe of good programming. Which means you can train on almost anything. Exactly. This unlocks the ability to train on vastly larger, unstructured datasets. And we know it scales. Right. The scaling results were impressive. The research tested this architecture across multiple sizes of the Quinn 2.5 model from 1.5 billion up to 7 billion parameters and the Llama 3.21 billion model. Across the board, they saw consistent improvements without hitting diminishing returns. So bringing this all together for you listening, what does this actually mean for the future?

21:35Why does a breakthrough in future spaces and whitening matter to anyone outside of an AI lab? It means the next generation of AI tools you interact with will be fundamentally more robust and faithful to your complex prompts. So no more derailed trains. No more derailed trains. When you ask an AI to write a comprehensive architectural document or complex piece of software, it won't hallucinate halfway through. It won't get lost in its own generated context. That sounds like a dream. And from a development standpoint, engineers won't need to manually write millions of external unit tests just to make the AI smarter.

22:09The AI is finally learning to align with human intent simply by observing the deep semantic structure of how we communicate. It really is a paradigm shift. I mean, we started this deep dive looking at models tripping over their own microscopic tokens, getting exponentially more confused the longer they talked. With that teacher forcing trap. Right. Then we saw the destructive nature of rigid outcome-only rewards that turned models into test-hacking robots speaking gibberish. The RLVR failures. And now we've arrived at a mechanism where models learn by matching the deep semantic essence of human knowledge.

22:43It's just brilliant. It is. Which leaves us with a fascinating philosophical question to ponder. Really. Ooh, lay it on me. If we are truly moving away from teaching AI to simply predict surface-level words, and instead teaching it to match abstract features hidden deep within its neural layers, are we finally crossing the threshold where AI communicates in a language of concepts rather than just a language of words? A language of concepts. Wow. That completely changes how we think about artificial intelligence. It really does. And hopefully with methods like EBFT, we finally have an AI that knows how to drive the car itself without the instructor constantly grabbing the wheel.

23:22Thank you so much for taking this deep dive with us. Catch you next time.

From the publisher

This research paper introduces Energy-Based Fine-Tuning (EBFT), a novel method for refining language models by matching feature statistics of generated text with ground-truth data. Traditional training relies on next-token prediction, which often causes models to drift or fail during long sequences because they lack global distributional calibration. By optimizing a feature-matching objective using a frozen feature network, EBFT provides dense semantic feedback at the sequence level without requiring a manual reward function. Experimental results across coding and translation show that EBFT outperforms standard supervised fine-tuning and equals reinforcement learning in accuracy. Furthermore, it achieves superior validation cross-entropy, proving that aligning feature moments effectively stabilizes model behavior and preserves high-quality language modeling.

More from Best AI papers explained

All 475 episodes
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language ModelsBest AI papers explained · 23 min
Listen in VO