In short
End-to-End Test-Time Training (TTT-E2E) for long-context LLMs, aiming for full-attention-like performance up to 128K tokens while keeping constant-cost inference.
Guests
No guests mentioned; this is a solo host “Deep Dive” episode.
Key claims
Full attention gives near lossless recall but costs O(T^2) compute; constant-cost models (e.g., Mamba 2, Gated DeltaNet) lose accuracy at long contexts. TTT-E2E treats the incoming document as a training signal at test time (continual learning) using an inner loop (optimize next-token loss on the context) and an outer meta-learning loop (gradients-of-gradients) to learn good initial weights.
Notable examples
Sliding window attention with 8,000-token SWA; updating only the last quarter of MLP blocks with a static “vault” MLP to prevent forgetting. Benchmarks: scales like full attention to 128K; 2.7x faster inference on H100. Trade-off: worse Needle-in-a-Haystack/passkey/number-in-haystack verbatim recall than full attention. Training latency: 3.4x slower; needs gradients-of-gradients support and better initialization.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenges in Large Language Models
0:45 to 1:48
Exploring the challenges of efficiently handling vast amounts of information in LLMs.
“So you double the length of the input, you quadruple the compute needed to process it.”
Redefining Long Context Problems
1:48 to 2:46
Discussion on the approach of using continual learning for long context issues in LLMs.
“How does this TTT E2E approach manage to get the performance scaling we see with full attention while also giving us the high-speed inference of the constant cost alternatives?”
The Mechanism of TTTE2E
2:46 to 4:21
Explanation of how test-time training (TTT) works and its benefits for LLMs.
“So how do you actually give a language model that kind of ability?”
Performance and Efficiency of TTTE2E
4:21 to 5:49
Analyzing the performance scaling and computational efficiency of TTTE2E.
“And what's really remarkable is the base architecture they started with.”
Trade-offs of TTTE2E Approach
5:49 to 9:59
Discussing the trade-offs between compression and lossless recall in the TTTE2E method.
“Yeah, that takes care of the performance requirement.”
Training Latency Challenges
9:59 to 12:26
Addressing the slower training latency issues of the TTTE2E approach compared to standard models.
“It gives you the gist fast and accurately.”
Future of Efficient LLM Memory
12:26 to 13:49
Exploration of potential advancements in LLM memory through efficiency and self-training.
“We started with this efficiency versus performance paradox for long context.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Our mission here is to take some of the most challenging new concepts in research and, well, distill them into knowledge you can use right away. And today we are wrestling with a perpetual challenge in large language models. I mean, it's one of the biggest efficiently remembering just vast amounts of information. Right. If you're building or using an LLM, you want it to handle these huge contexts. We're talking documents, 128 ,000 tokens long, maybe even more. The theoretical gold standard for this right now is the classic transformer with full attention. It gives you nearly perfect lossless recall, which sounds ideal.
0:38It does, but the cost is brutal. The computation grows quadratically with the context length. That's the whole O of T squared problem. So you double the length of the input, you quadruple the compute needed to process it. Exactly. And at lengths over, say, 100K tokens, that cost just becomes prohibitive for any practical, fast application. It creates this fundamental tension because on the other side, you have these really elegant architectural solutions, things like Mamba 2 or Gated DeltaNet. Yeah, and they achieve a constant cost per token O of 1, which means you get this blisteringly fast inference.
1:10But. There's always a but. There's always a but. When you push them into those really, really long contexts, they start to struggle. Their ability to actually use that massive history begins to degrade. So you solve the speed problem, but you lose that robust performance. Which is what makes the research we've been looking at so fascinating. The team behind it basically decided to stop fighting the battle on architecture alone. Right. They looked at this long context problem and they redefined it. They formulated it not as a structural bottleneck, but as an issue of, well, continual learning. They call it end-to-end test time training, or TTTE2E.
1:47And that's our mission for this deep dive. How does this TTT E2E approach manage to get the performance scaling we see with full attention while also giving us the high-speed inference of the constant cost alternatives? We really need to know where the magic is. And maybe more importantly, what are the trade-offs? What do you give up when you redefine a problem like this? The best way to get at the core idea is probably through the human analogy they used. Yeah, I like this. Think back to a foundational lecture you took years ago. Maybe a really tricky topic like machine learning. You probably don't recall the instructor's exact first word from the first day.
2:22Not a chance. Yeah. But the high-level intuition, the key patterns, the general concepts, that stuff is compressed into your long-term knowledge base. And that is the functional difference between compression and lossless recall. Full attention is going for lossless recall. It wants to remember every single word verbatim. While the TTT2E approach is aiming for efficient compression. So how do you actually give a language model that kind of ability? The insight is this. You treat the context itself as a training signal. You just continue training the model at test time, using the incoming document to update the model's weights.
3:00So you're compressing the document it's reading directly into its own structure. Now this idea, test time training or TTT, it's not entirely new, is it? It sort of echoes older ideas like dynamic evaluation, where a model updates itself mid-sequence. It does. But what makes this new method end-to-end, the E2E part, is what solves a really critical problem from those earlier attempts. Which was the objective mismatch. Right. If you trade a model for general knowledge, but then you ask it to learn dynamically for a specific context, those two goals can really clash. TTTE2E addresses this with two sophisticated interconnected loops.
3:38So the first one is the inner loop. This is happening live at test time. Right. This is standard continual learning. The model sees a sequence of tokens, and it optimizes its next token prediction loss on that specific context. It's learning the details of the document it's reading right now. And the outer loop, this is where the term meta-learning comes in. And for anyone listening, this feels like the most important innovation. Absolutely. The outer loop happens during the original pre-training phase. It uses a meta-learning technique specifically by calculating what they call gradients of gradients.
4:07To optimize the model's initial weights, what we'll call W0, you're not just training it to predict the next token. You're training it to be the optimal starting point for that inner loop. You're training the model to be an exceptional student. So when it sees a new, long document, it knows exactly how to absorb that context into its weights quickly and efficiently. That's a perfect way to put it. And what's really remarkable is the base architecture they started with. I mean, we've spent years debating Mamba versus DeltaNet versus the classic transformer, but they didn't invent a new mechanism.
4:41No, they used a standard transformer and just paired it with sliding window attention or SWA. And just for clarity, SWA is a way to drastically cut costs by only letting the attention mechanism look back at a recent fixed block of tokens. So maybe the last 8 ,000 tokens instead of the full 128 ,000. Correct. So the architecture is really just playing a supporting role here. It provides the short-term memory. The breakthrough is entirely in this continual learning process layered on top. Which transforms that short-term memory into long-term comprehension. So when we look at the core data, the strength of this TTT E2E approach is it's undeniable, especially in its staling properties.
5:20They benchmarked these three billion parameter models, and the results showed TTT E2E performance scales with context length all the way up to 128 ,000 tokens in the exact same way as a costly transformer with full attention. So they solved that performance degradation problem that hits the constant cost alternatives. I mean, while Mamba 2 and Gata DeltaNet saw their loss curves just climb as the context got longer, TTTE2E kept pace with the gold standard. Yeah, that takes care of the performance requirement. Okay, so now let's talk about the other side of that equation, efficiency, the speed.
5:55TTTE2E achieves constant inference latency, O of 1. It doesn't matter how long the context is, which puts it right in speed camp with those RNN-style baselines. Meaning the time it takes to process 128 ,000 tokens is basically the same as processing 8 ,000? Negligibly different, yeah. Yeah. And that translates to real world speed. On an H100 GPU, TTT E2E was measured to be 2.7 times faster than full attention for 128K contexts. That is not a marginal improvement. That's a massive efficiency boost. It makes long context processing actually viable for high throughput applications. But getting that combination of stability and speed, it required some very specific surgical decisions about the training itself.
6:36They didn't just update the whole model. No, they had to be really precise. Precisely. To save, compute, and keep it stable during that inner loop, they restricted the updates drastically. They used a standard sliding window of 8 ,000 tokens, with the TTT learning happening in mini-batches of 1 ,000 tokens, and they found that updating only the MLP layers was necessary. Leaving the attention, embedding, all that other stuff static. But wait. But wait, if you restrict the updates that aggressively, Only the MLPs, where most of the generalized knowledge is, how do you make sure the model doesn't just forget everything?
7:08The catastrophic forgetting problem. Right. Forgetting all its general pre-trained knowledge as it adapts to this one specific document. That's where the architectural nuance comes in. And it's a beautiful illustration of the whole compression strategy. They found they didn't even need to update all the MLP layers. Just updating the last quarter of the blocks gave the best result. Anything more was just extra compute for no real game. Exactly. And to address the forgetting issue you raised, they included a static second MLP layer in those updated blocks. And this static layer acts like a safe storage, a vault for the model's pre-trained general knowledge, ensuring that all this rapid adaptation...
7:47Right. But the full attention model has to optimize its single set of weights to be good at all future tokens in that massive window, which is an infinitely harder, more constrained task. So its constant adaptation gives it a performance edge. We've established the success, breakthrough speed, robust scaling, all through efficient compression. But here's the crucial pivot. Efficiency always comes with a cost. Always. If the whole mechanism is designed to compress massive experience into reusable intuition, what critical details does it intentionally leave behind? This is where we have to go back to that foundational difference between compression and lossless recall.
8:24If you're building long-term memory, you are by necessity leaving out details. And to test this limit, researchers use the gold standard for testing verbatim memory, the needle in a haystack evaluation, or NIAH. Right. These NIAH tasks specifically require nearly lossless recall. You hide a specific, often irrelevant, string-like, a unique idea, a random number, a passkey deep inside a very long, distracting passage. If the model is relying on compression, on summarizing, it's going to fail to retrieve that one tiny, specific detail. And the results here are extremely illustrative of the trade-off.
8:59They confirm the intuition of the method. The data from all the NIAH tasks, pass key retrieval, number in haystack, it clearly shows the transformer with full attention dramatically outperforms TTT E2E. And all the other constant cost methods. All of them, especially as the context length gets really long, beyond that SWA size. Full attention, just by its design, looks at the key and value of every single token. It's just the superior choice for verbatim memory. So this presents a critical choice for you, the listener, and it depends entirely on your application. If you need a model to act like a perfect legal archivist, flawlessly recalling every exact quote from a 100-page document.
9:39Full attention is still the engine you want. It excels at that granular, verbatim recall. But conversely, TTTE2E is the superior engine if you need a model that can quickly and efficiently extract generalized knowledge, intuition, trends, key summaries from a massive document load. Yeah, it prioritizes speed and high-level understanding over retrieving trivial details. It gives you the gist fast and accurately. Okay, so despite the constant O of one cost during inference, which is the dream, we have to address the practical limitation that the paper is very frank about, training latency. Right.
10:12While running the finished model is fast, training it is currently much slower than a standard transformer. Why? It's because that meta-learning process for the outer loop requires calculating gradients of gradients. And that procedure is computationally intense. It's far less optimized than the standard backpropagation we use for normal transformer training. Yeah, the data showed that training latency for TTT E2E is 3.4 times slower than full attention at the 8 ,000 token context length, which is where most of that pre-training compute is spent. And if we look at the computational realities, even though the method maintains a constant number of FLOPs floating point operations per token during training, the actual wall clock time still ticks up between 8K and 32K contexts.
10:57And that's because to manage memory, they had to increase the amount of gradient checkpointing through time. Exactly. Can we just break down gradient checkpointing for a second? Why does managing memory hurt speed? Sure. In short, training these huge models requires storing the activations of every layer for every token in memory. You need them to calculate the gradients accurately during backpropagation. But with long contexts, you just run out of GPU memory. You run out. So checkpointing is essentially recalculating those activations as needed rather than storing them all at once. When you do this through time, you're doing it repeatedly across the whole sequence.
11:34So it saves your memory, but the tradeoff is you have to pay the latency cost of recalculating everything over and over again. It just slows down training. That is the current technical hurdle for TTTE2E. But fortunately, the researchers did outline some clear paths forward. Right. They have a plan. First, they need a custom attention kernel that's specifically designed to support calculating gradients of gradients. The current optimized kernels, like flash attention, they just don't do that. And second. Second, they can initialize TTT E2E training from an already robust pre-trained transformer that was trained without TTT.
12:10Ah, so you minimize the time you spend in that slow meta-learning phase. You get the bulk of your general knowledge quickly with the standard method, and then you just spend a little bit of time adapting it to be the optimal student. Exactly. It drastically cuts the overall training bill. This has been a really insightful exploration. We started with this efficiency versus performance paradox for long context. And we learned that TTTE2E redefines the problem as continual learning, layering this learning process on top of a pretty minimal architecture. And this hybrid approach gives you breakthrough speed, 2.7 times faster inference at 128K, and robust performance scaling, all by prioritizing knowledge compression.
12:50The mechanism fundamentally creates this functional hierarchy that's, I mean, it's almost like biological memory. It is. The updated weights, which are absorbing the stream of tokens, they act as the model's long-term memory. The generalized knowledge, the intuition. And the core architectural part, the sliding window attention, that's the short-term memory. Just efficiently managing the immediate local context. Exactly. Which brings us to a final provocative thought for you to consider. If the efficiency of TTT E2E comes from focusing on the present context and, by design, leaving out seemingly irrelevant details for compression.
13:24And the researchers suggested it could be enhanced by using self-generated tokens during training, like a filtered or rephrased version of what it's currently reading. Then how could a model that is inherently designed to forget irrelevant details for efficiency be guided to generate its own training data? To selectively reinforce and remember the right things in the first place. The boundary between a model's self-training and its self-forgetting is thin. And it really defines the future of efficient LLM memory.
From the publisher
This research introduces TTT-E2E, a novel method for long-context language modeling that treats the task as a continual learning challenge rather than an architectural redesign. Unlike standard Transformers that struggle with the high computational cost of processing vast amounts of data, this model **compresses context into its weights** by learning at test time via next-token prediction. By integrating **meta-learning during training**, the system is optimized to initialize effectively for these **test-time updates**, ensuring the model improves as it reads more information. The authors demonstrate that while traditional RNNs and hybrid models lose effectiveness in very long contexts, **TTT-E2E scales performance** similarly to full-attention Transformers while maintaining the **constant inference speed** of an RNN. Ultimately, the method achieves significant efficiency gains, running **2.7 times faster** than standard models at a 128K context length while achieving superior language modeling accuracy.




