In-Place Test-Time Training

9 Apr 2026 · 20 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

In-place test-time training (in-place TTT) to stop LLMs from “forgetting” early context by enabling continuous on-the-fly learning during inference, without retraining or infinite context windows.

Guest backgrounds

No guests are identified; the transcript shows two hosts discussing the research.

Key claims

Current LLMs follow a train-then-deploy paradigm with frozen weights; long prompts fail due to attention’s quadratic complexity and because in-context learning just stuffs text into prompts. Original test-time training (TTT) failed due to architectural incompatibility (replacing attention), slow per-token updates that bottleneck GPU parallelism, and misaligned objectives (reconstruction vs next-token prediction). In-place TTT unfreezes only the MLP block’s W-down matrix (W-up/W-gate remain frozen), uses chunk-wise context parallelism with associative updates, and trains fast weights with a next-token-aligned objective using 1D convolution.

Notable examples

“Goldfish” forgetting; RULER benchmark (needle-in-haystack, long instruction following) with up to 128k tokens and extrapolation to 256k; applied to QEN3-4B, LLaMA 3.18B, and Qwen3.14B; from-scratch training at 500M and 1.5B beating sliding window attention, gated linear attention, and DeltaNet on perplexity.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

In-Place Test-Time Training Explained

0:48 to 1:59

Explore the concept of in-place test-time training and its potential to evolve AI models.

“Which brings us to our mission for this deep dive, because today we are unpacking a groundbreaking new framework called in-place test time training or in-place TTT for short.”

Understanding AI Model Limitations

1:59 to 4:26

Examine why current AI models struggle with new inputs and the training-deployment cycle.

“Like, why do they suddenly max out when you give them massive amounts of new text?”

Limitations of Original Test-Time Training

4:26 to 6:36

Discuss the flaws of the initial test-time training (TTT) methodology and its impact.

“But scientists didn't just accept this limitation.”

The Flaws of Original TTT Mechanism

6:36 to 10:40

Delve into the three key flaws that hindered the adoption of the original TTT framework.

“The speed issue came down to computational inefficiency.”

In-Place TTT: A Paradigm Shift

10:40 to 14:00

Learn how the in-place TTT framework addresses previous issues and enhances AI models.

“And that is the brilliance of keeping W-up and W-gate frozen.”

Understanding Fast Weights and Context Prediction

14:00 to 14:48

Learn how fast weights help AI models align with autoregressive nature for predictions.

“Like a reader trying to pick up on context clues to guess the twist on the next page of a mystery novel.”

Empirical Data and Model Enhancements

14:49 to 16:15

Explore the practical test results of applying a new framework to existing models.

“The empirical data is actually quite striking.”

In-Place TTT Architecture and Performance

16:16 to 17:38

Discover how in-place TTT improves model performance and reduces perplexity.

“In both cases, the Ruler benchmark scores saw massive gains across the board, particularly at the 64 ,000 token links.”

Implications of AI's Evolution

17:39 to 19:20

Understand the broader implications of continuous learning in AI and its future.

“A model with lower perplexity feels sharply focused, coherent, and highly confident in its logic.”

The Future of Personalized AI

19:21 to 20:13

Ponder the unique implications of AI learning from individual user data.

“It seems the days of talking to a goldfish are finally ending.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, there is this very specific kind of modern frustration that I think we've all felt recently. You were working with an AI chatbot, right? Yeah. And you feeded this massive document. Or maybe you've just been talking to it for an hour, building up this really complex, nuanced context. Oh, I know exactly where this is going. Right. And then suddenly it just completely forgets something you told it at the very beginning of the conversation. It is, I mean, it's essentially like trying to have a deep philosophical debate with a goldfish. It really is. And, you know, the reason for that limitation is that we currently treat artificial intelligence like a static encyclopedia.

0:38You can write in the margins for a little while, but eventually you just run out of room. Yeah. And once that happens, the AI just stops retaining any of the new information you were giving it. It just hits a wall. Which brings us to our mission for this deep dive, because today we are unpacking a groundbreaking new framework called in-place test time training or in-place TTT for short. The research we're exploring today proposes a way to fundamentally unfreeze artificial intelligence. Unfreeze is a great word for it. The goal is to allow the model to continuously learn and adapt while it reads new information, like on the fly, rather than relying completely on the static data it memorized months ago during its initial creation.

1:22Yeah, and I mean, it represents a massive shift in how we approach machine learning as a whole. We are moving away from treating AI as this frozen snapshot of internet data, and we're moving toward an AI that functions more like an evolving adaptive thinker. Okay, I love that, an adaptive thinker. Exactly. This framework enables the model to dynamically rewire its own understanding as it processes the unbounded stream of information you feed it. That is so cool. But before we can really appreciate the mechanics of how this new framework fixes the goldfish problem, we should probably establish why current AI models hit this wall in the first place.

1:59Like, why do they suddenly max out when you give them massive amounts of new text? Well, it comes down to the standard life cycle of large language models. We essentially operate on a train-then-deploy paradigm. Train-then-deploy, right. Right. So during the training phase, the model absorbs massive internet scale corpora of text. This phase takes months, it costs millions of dollars, and it's where the model learns the fundamental patterns of language and logic. But the second that model is deployed to you, its internal weights, its digital brain, essentially are entirely frozen. So it can't learn anything fundamentally new, it's just applying what it already knows to whatever you type into the prompt box.

2:38Exactly. And to get around that limitation, we currently use something called in-context learning, which basically means we just stuff all the new information into the prompt itself. If I wanted to summarize a massive legal document, I'd just paste the whole thing in. Yeah, that's the current workaround. But if that's the workaround, why can't we just do that forever? I mean, why can't we just build computers with infinitely large context windows? If I have a massive 10 million line code base, why can't I just paste the entire thing in? Because of the underlying math of how these models actually read text, it relies on a concept known as the quadratic complexity of the attention mechanism.

3:14Okay, let's unpack that. Because quadratic complexity sounds intimidating, but it's really just a scaling problem, isn't it? It is a massive scaling problem, yeah. So for an AI to understand a sequence of words using its attention mechanism, every single word needs to mathematically look at every other word to establish context. Makes sense. Right. So if you input the sentence, I went to the bank of the river, the model needs to connect the word bank to the word river so it knows you aren't talking about like a financial institution. With 10 words, the AI calculates the relationships between all 10.

3:49That's roughly 100 connections. 10 times 10. Simple enough. Right. But now input 100 ,000 words. You aren't just calculating 100 connections anymore. You are calculating 10 billion connections. Oh, wow. 10 billion. Yeah. As the input grows, the computing power required doesn't just increase steadily or linearly, it explodes quadratically. The math itself turns into a black hole that consumes an insane amount of memory and completely bottlenecks the AI's ability to reason over long, complex tasks. So you really can't just buy a bigger memory box. The underlying architecture is fundamentally flawed for infinite context.

4:25Exactly. But scientists didn't just accept this limitation. The data shows that before this new in-place framework came along, there was an earlier attempt to solve this called simply test time training or just TTT. Yeah, the original TTT. And the core concept of that original framework was highly ambitious. It introduced the idea of separating an AI's memory into two distinct types, fast weights and slow weights. This is a really great concept to visualize. So the frozen weights from the initial training are the slow weights, like the AI's long-term memory, right? The massive textbooks it studied in college for four years.

5:01Right. Then the fast weights are like a dynamic stack of sticky notes. The AI uses those sticky notes to jot down new information during a live, real-time conversation. That's a perfect analogy. The fast weights act as an online evolving state. They compress and internalize the new context on the fly, storing it in those sticky notes so the model doesn't have to constantly look back and recalculate the entire massive prompt using that super heavy attention mechanism. Which sounds amazing. But while the theory was sound, the original implementation of TTT struggled to really gain traction in the real world, didn't it?

5:36It did. It had some pretty fatal flaws. Let's break down why it failed to catch on. Because the initial approach to getting those sticky notes into the AI required essentially ripping out the AI's existing architecture. Yeah. They built these original TTT mechanisms as standalone specialized layers designed to completely replace the standard attention mechanism. Which created a massive barrier. Yeah. Architectural incompatibility. Exactly. You couldn't just add this old version of TTT to a model that already existed. To use it, you had to throw away state-of-the-art models that cost tens of millions of dollars to train.

6:11You had to undergo the incredibly costly process of retraining an entirely new model from scratch just to incorporate these new specialized layers. And nobody's going to throw away a brilliant billion-parameter model just to get better memory. I mean, that's just bad business. Right. But even if a company did decide to bite the bullet and train a new model from scratch, they ran into a second fatal flaw. It was agonizingly slow. All right. The speed issue. Yeah. The speed issue came down to computational inefficiency. The original TTT updated its fast weights, its sticky notes on a strict per token basis.

6:47It read and updated sequentially word by word by word. And if you're listening to this and wondering why doing things sequentially is bad for a computer, it helps to understand how modern AI hardware actually works. The GPUs that run these models don't operate like a single person reading a book. No, not at all. They operate like an assembly line with thousands of workers. They get their speed by processing massive blocks of data all at the exact same time. Exactly. So when you force a GPU to process data sequentially, word by word, you are essentially leaving like 9 ,999 of those assembly line workers standing around doing nothing while one worker processes a single word.

7:26It completely bottlenecked the massive parallelism that modern accelerators rely on. You have to build a new model from scratch. And even if you do, it runs terribly on modern hardware. Yep. And yet there was a third, even deeper problem with the original DTT. The way it was trying to learn on the fly was fundamentally disconnected from what an AI is actually supposed to do. Right. It suffered from misaligned objectives. The original TTT used a generic reconstruction objective. What does that mean exactly? Essentially, as the model read a word, the fast weights just tried to memorize and reconstruct that current specific word.

8:00But language models are autoregressive. Their entire purpose is to predict what comes next. Right. Spending all your cognitive energy trying to perfectly memorize the word you are currently looking at doesn't actually help you guess the next word. It's focusing on the past rather than anticipating the future. Exactly. Which brings us out of the history lesson and into the core innovation of the research we are unpacking today. Because the real question is, how do you give an AI these fast-weight sticky notes without forcing it to read word by word, without focusing on the past, and most importantly, without ripping out its entire brain and starting over?

8:37And this is where the new in-place TTT framework totally changes the paradigm. Instead of adding brand new incompatible layers and throwing away the attention mechanism, this framework repurposes a ubiquitous part of the existing model. Specifically, it repurposes the multilayer perceptron blocks, commonly known as the MLP blocks. Okay, let's unpack the MLP block. Yeah. Because this drop-in design is really the defining feature of the framework. These MLP blocks are already sitting right there inside the transformer architecture of almost every major LLM today, right? Yep, they are everywhere.

9:09You can think of the MLP blocks as the feed-forward neural networks that act as key value memories for the vast general knowledge the AI acquired during pre-training. Okay. So inside a standard gated MLP block, there are a few distinct matrices of weights. You have W-up and W-gate, which initially process the input and decide what information passes through. Got it. In this new framework, those matrices stay completely frozen. They remain the slow weights. But the final projection matrix, called W-down, which actually produces the output of the block, that specific matrix is unfrozen and treated as the adaptable fast weights.

9:48It's essentially like the mechanics of a nightclub. Oh, a nightclub. Yeah, think about it. The W-up and W-gate matrices are the bouncers of the door. They are the slow weights operating on strict frozen rules about who gets in based on years of training. Okay, I see where you're going. But that final W-down matrix is the DJ inside the club. The DJ represents the fast weights. They can dynamically read the room, meaning the current prompt you're feeding the AI, and continuously change the playlist on the fly to adapt to the vibe. That is actually a highly accurate way to look at it. The input is still filtered through the foundational rules learned during pre-training, but the output is dynamically tuned by the DJ.

10:27Wait, but if we are modifying the original weights inside the model, even just the DJ in this scenario, doesn't that risk giving the AI amnesia for the things it already learned during its initial multi-million dollar training? And that is the brilliance of keeping W-up and W-gate frozen. Because they only adapt that final W-down projection matrix in place, the integrity of the original pre-trained weights is heavily preserved. The foundational logic of the model doesn't collapse. Oh, wow. Yeah. This means you don't need to retrain the model from scratch. You take a pre-existing model, unfreeze that specific matrix, and you have a true drop-in enhancement.

11:04Okay, so we've solved the architectural incompatibility. You get to keep your billion-parameter model, but now it has dynamic sticky notes inside its MLP blocks. But the framework still needed to solve the speed issue and the alignment issue to be functional in the real world. Let's cackle speed first. If per token, word-by-word updating bottlenecks the GPUs, how does in-place-TTT actually process the text? It completely abandons the slow per token update. Instead, it utilizes something called context parallelism through a chunk-wise update mechanism. Meaning it grabs a whole handful of text at once.

11:40Specifically, it processes large chunks of 512 to 1024 tokens simultaneously. And the reason it can do this without scrambling the chronological order of the words is because the new update rule relies on mathematically associative operations. Let's clarify what associative math means in this context, because it really is the secret to unlocking the speed. In simple math, associative means that adding A plus B and then adding C is the exact same as adding A to the combined sum of B and C. Right. The grouping doesn't change the final answer. And by designing the fast weight updates to be associative, the framework allows the system to use a parallel scan algorithm.

12:18It breaks a massive document into hundreds of chunks, sends those chunks to different GPU processors to be evaluated at the exact same time, and then recombines the results seamlessly. And if you're listening to this and wondering why context parallelism matters to you, just think about the last time you asked an AI to summarize a massive PDF. And you had to sit there for 30 seconds watching a blinking cursor while it generated an answer. Yeah, that wait is painful. This parallel scanning is what reduces that wait time to a fraction of a second. It perfectly saturates modern hardware. It achieves incredibly high throughput while maintaining the rigorous sequence of the data.

12:55It's a huge win. So it's fast, and it works on existing models. But we still have that third problem, the misaligned objective. How do we get the AI's fast weights to stop just trying to memorize the current word and start acting like a predictive language model? The researchers shifted the framework away from generic reconstruction and introduced a new objective explicitly aligned with the next Hoken prediction task. Okay, but how does the math actually force the model to look ahead? It does this by utilizing a 1D convolution operator when calculating the target value that the fast weights are trying to learn.

13:32A 1D convolution operator. Right. A convolution operator essentially slides a window across the sequence of data. Instead of looking at a single token in total isolation, the 1D convolution looks at a small neighborhood of overlapping tokens. Functionally, this means the target state the fast weights are optimizing for incorporates information from the future tokens in the sequence. So instead of staring blankly at the current word and trying to burn it into memory, the framework forces the fast weights to constantly scan for predictively useful information for the next words. Yes. Like a reader trying to pick up on context clues to guess the twist on the next page of a mystery novel.

14:10Precisely. By forcing the fast weights to compress information that is strictly useful for predicting the future, the model aligns perfectly with its core autoregressive nature. The theoretical analysis inside the data proves mathematically that this objective directly increases the logits, which are the raw probabilities the model calculates for predicting the correct next tokens. Okay, the math works beautifully on paper. The architecture drops in, the chunks process in parallel, the convolutions predict the future. But, you know, theoretical math often falls apart when you actually put it inside a real working AI model.

14:45So does this framework hold up in practice? this, what does the data say? The empirical data is actually quite striking. Let's look at the drop-in enhancements first. They applied this framework to the QEN3-4B base model, which is a highly competitive 4 billion pyramidal model. And they didn't retrain it from scratch, right? Nope, they didn't retrain it from scratch. They just applied a relatively cheap continual training curriculum to basically wake up the unfrozen in-place TTT modules. And to test how well this works, they evaluated it on the ruler benchmark. For context, the Ruler benchmark doesn't just test if a model can hold a lot of text in its memory.

15:22It tests if the model can actually use it. It involves tasks like finding a specific needle of information hidden in a massive haystack of text or following multi-step instructions across a huge document. How did the upgraded model perform? Its functional context window was blown wide open. The enhanced model achieves superior performance with contexts up to 128 ,000 tokens. That's huge. But even more remarkably, it successfully extrapolated out to 256 ,000 context lengths. It was maintaining accuracy on sequence lengths it had never even explicitly been trained on during its continual learning phase.

15:58Just to ground those numbers, 256 ,000 tokens is roughly the length of a 700-page book. You could feed this relatively small 4 billion parameter model an entire epic fantasy novel, and it wouldn't lose the plot halfway through. Exactly. And the data shows this isn't just a quirk of one specific model, right? Not at all. The exact same drop in enhancement was applied to different architectures, including LAMA 3.18B and the significantly larger QIN 3.14B base model. In both cases, the Ruler benchmark scores saw massive gains across the board, particularly at the 64 ,000 token links. The framework is highly model agnostic.

16:35But the applications of this framework don't stop at just upgrading old models. The data also explores what happens if you build a model from scratch with this in-place TTT architecture baked in from day one. Yeah. Testing from scratch is crucial to validate the fundamental merit of the architecture against the current cutting edge. They pre-trained models from scratch at the 500 million and 1.5 billion parameter scales. Then they pitted them against competitive, state-of-the-art architectures designed specifically for long contexts. Architectures like sliding window attention, gated linear attention, and delta net.

17:08Exactly. And those baselines are heavy hitters. Sliding window attention, for example, tries to save memory by only looking at a few surrounding words, but it often loses the big picture. So how did in-place TTT compare to those? Across the board, the models pre-trained with in-place TTT consistently achieve lower perplexity than all of those baselines. Which is incredible. And for anyone unfamiliar with the term, perplexity is essentially a measurement of how surprised or confused an AI model is by the text it's generating. If you've ever used a chatbot that suddenly starts hallucinating facts or completely loses the thread of a story, you are witnessing high perplexity.

17:46Right. A model with lower perplexity feels sharply focused, coherent, and highly confident in its logic. And by achieving lower perplexity than established, state-of-the-art efficient architectures, the data definitively proves that replacing a static projection matrix with dynamic fast weights yields a superior model. It keeps the high-resolution big picture without the quadratic explosion in computational cost. So what does this highly technical evolution of MLP blocks and junk-wise fast weights actually mean for your daily life? Well, it means the ceiling on what AI can process is being shattered.

18:20We're moving toward a reality where you can feed an AI entire libraries of legal documents, giant repositories of software code, or the transcripts of every single meeting your company has had for a year. Yes. And instead of grinding to a halt, eating up all your computer's memory, or just forgetting the first half of what you told it, the AI will adapt. It will internalize that information in real time, functioning much closer to a human taking detailed notes. If we look at the broader implications here, we are watching the static train-then-deploy paradigm begin to crack. In place, TTT represents a highly practical, scalable step toward true continual learning in large language models.

18:56Which is the holy grail, really. Yeah. It is. By repurposing existing MLT blocks, processing data in efficient parallel chunks, and aligning the learning objective with actual prediction, we are finally mimicking how humans learn from unbounded streams of experience. A person doesn't permanently freeze their brain's neural pathways the day they graduate college, right? They continue to adapt to new environments. Now we have a mathematical framework that allows AI to do the exact same thing. We are transitioning from a rigid frozen encyclopedia to a lightweight drop in architecture that can actually think on its feet, which leaves us with a fascinating and maybe slightly unsettling question for you to mull over.

19:36Oh, I like where this is heading. If this framework allows an AI to seamlessly and continuously update its own internal weights on the fly based entirely on the specific unbounded streams of information, you give it what happens to the concept of a universal base model. That is a wild thought. Right. If you and I download the exact same AI today, but we feed it entirely different documents, projects, and conversations for a month, will we soon live in a world where your AI's underlying architecture becomes fundamentally neurologically unique to you? It seems the days of talking to a goldfish are finally ending.

20:10The question now is, what kind of mind are you going to shape?

From the publisher

This paper introduces In-Place Test-Time Training (In-Place TTT), a novel framework designed to let Large Language Models (LLMs) dynamically update their knowledge during inference. Traditional models remain static after deployment, but this approach repurposes existing MLP blocks as "fast weights" that adapt to new information in real-time. By utilizing a chunk-wise update mechanism and a learning objective aligned with Next-Token Prediction, the system achieves high computational efficiency on modern hardware. Experiments demonstrate that this "drop-in" enhancement significantly improves performance on long-context tasks up to 128k tokens without requiring expensive retraining from scratch. Ultimately, the research offers a scalable path toward continual learning, allowing models to internalize evolving contextual data more effectively than standard attention mechanisms.

More from Best AI papers explained

All 475 episodes
In-Place Test-Time TrainingBest AI papers explained · 20 min
Listen in VO