Beyond the Transformer: Titans, MIRAS, and the Future of Infinite Context

7 Dec 2025 · 39 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Long-context AI memory beyond transformers, focusing on Google Research’s Titans architecture and the Miras framework for “test-time memorization” (updating long-term knowledge during inference) to overcome quadratic attention scaling.

Guest backgrounds

No specific guest names or biographies are provided in the transcript; the episode is presented as a “Deep Dive” with two hosts.

Key claims

Standard transformers excel at short context but scale with O(N^2) attention, becoming infeasible for million-token inputs. Linear recurrent/SSM models scale O(N) but compress history into a fixed vector, limiting nonlinear abstraction. Titans adds deep long-term neural memory (MLP) plus selective update via a “surprise” (gradient/anomaly) metric, with momentum and adaptive forgetting (weight decay). Miras unifies sequence models as associative-memory optimization and argues for non-Euclidean, non-MSE objectives.

Notable examples

Quadratic FLOP estimates (10k→100M, 100k→10B, 1M→1T attention ops); legal/financial anomaly (“banana peel”); Babylon benchmark long-document reasoning; Titans MSC perplexity 19.93 vs Mamba-like 28.1 and transformer 21.18; scaling beyond 2M tokens; genomic (hyena DNA comparison) and time-series forecasting (Simba).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenge of Long-Term Memory in AI

0:46 to 2:50

Discussing the limitations of current AI models related to memory.

“All the mechanics that powered this AI revolution, they just they start to break down spectacularly.”

Understanding Transformers and Their Limits

2:51 to 5:06

Explaining the transformer architecture and its computational bottlenecks.

“So let's start with the status quo, the mighty transformer.”

Exploring Alternatives to Transformers

5:07 to 6:00

Overview of alternatives like linear recurrent models and their efficiency.

“And that just makes scaling for these massive data streams economically and technically.”

The Limitations of Linear Models

6:01 to 7:40

Discussing why linear models cannot reach the expressivity of transformers.

“The move toward linear recurrent models was very popular, wasn't it?”

Introduction to Titans and MIRA

7:41 to 9:21

Presenting Titans and MIRA as a solution to the long-term memory problem.

“It's like trying to summarize a library full of dense, multi-layered philosophical books onto a single, giant, but ultimately linear note card.”

Components of Titans Architecture

9:22 to 11:28

Deep dive into the components of the Titans architecture and their functions.

“It's often constrained to a local sliding window attention, SWA, or a very limited context.”

Understanding Dynamic Learning in Titans

11:29 to 12:35

Examining how the Titans model decides what information to store permanently.

“And they demonstrated much better scaling properties as the sequence link moved into that extreme long context regime.”

The Role of Surprise in Memory Formation

12:36 to 14:00

Illustrating how the model prioritizes memory based on surprise and anomalies.

“We've got short-term attention, the deep long-term MLP, and the fixed persistent knowledge.”

Understanding Memory Expectations in AI

14:00 to 15:05

Learn how AI models process legal documents and handle unexpected inputs.

“So if I'm feeding the model a thousand-page legal contract, and it expects the next input to be another legal clause, and it is.”

Momentum and Forgetting Mechanisms

15:05 to 17:48

Explore how momentum and forgetting enhance AI memory efficiency.

“And the selective update process has to be essential for efficiency.”
Show all 20 chapters

Maximizing Computational Efficiency in Memory Updates

17:48 to 19:33

Discover how advanced techniques optimize memory updates for speed.

“The memory has to have a way to actively prune or clear out old information that's no longer useful.”

MIRAS: The Unified Theory of Sequence Models

19:33 to 20:15

Understand the MIRAS framework that connects various AI architectures.

“Okay, so now we need to make a theoretical leap.”

Identifying Limitations in Current Memory Systems

20:15 to 21:49

Learn about the drawbacks of mean squared error in AI model performance.

“And within this new framework, Mirus defines any sequence model through four sort of concrete design choices.”

Transitioning Beyond Euclidean Memory Objectives

21:49 to 22:41

Explore how non-Euclidean objectives can improve AI memory systems.

“Well, MSE applies a really steep squared penalty for any error.”

Designing Robust AI Memory Models: YAD, Moneta, and MEMRA

22:41 to 25:06

Discover innovative AI memory models that enhance stability and robustness.

“And the researchers used the Marist framework to design and test new variants that explicitly move beyond that standard MSE paradigm.”

Integrating Memory with Attention in AI Architectures

25:06 to 28:04

Learn about the integration of long-term memory with attention mechanisms.

“Right, and this is where the Titans family of architectures comes in.”

Evaluating Memory Model Variants

28:04 to 29:58

Learn about the performance differences between memory models MA, MAG, and MAL.

“It performed very well on generalized language modeling tasks.”

Titans vs. State-of-the-Art Models

29:58 to 33:31

Discover how Titan models outperform existing models in language tasks.

“Do you have a specific data point for us?”

Broader Applications of Titan Models

33:31 to 35:57

Explore the versatility of Titan models across various non-linguistic tasks.

“For instance, in genomic modeling, DNA tasks.”

Theoretical Implications of Deep Memory

35:57 to 38:28

Understand the theoretical shifts Titans and Miras introduce to AI capabilities.

“Third, theoretical unification via Miras allows for robust non-Euclidean memory objectives.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. This is where we give you the essential shortcut to understanding the most fundamental shifts happening in AI research right now. And today we are tackling something that is, well, it's arguably the defining technical challenge for the future of AI. We're talking about memory or more precisely, the lack of truly effective, massive scale, long term memory in our models. Exactly. If you follow large language models at all, you know, they're they're phenomenal at short term context. No, absolutely. You know, what they just read in the last 10 ,000 or maybe even 100 ,000 tokens.

0:35But the second you need to go bigger. That's where it all breaks down. Right. When the context window needs to expand to handle, say, an entire library of documents or a massive genomic sequence or a data stream that spans years. All the mechanics that powered this AI revolution, they just they start to break down spectacularly. And that's what we're here to talk about. We are. We're diving into some really groundbreaking sources today. It's a collection of research papers and reports from Google Research, and they introduced two interconnected breakthroughs. First, there's the Titan's architecture.

1:07And second, the Mycorrhiz framework. And this isn't just, you know, a minor optimization. This is a foundational reimagining of how an AI remembers, how it learns, and how it processes these incredibly long sequences of data. That's absolutely right. Right. I think our mission in this deep dive is to really clearly lay out how this research gets over that seemingly insurmountable computational law. The scaling problem. The scaling problem. How do you process these enormous, messy, real-world data streams? The goal here isn't just about making it faster. It's not just efficiency. No. It's about combining that efficiency with the kind of deep recall and the expressive power that you expect from a world-class model, but doing it at an absolutely exponential scale.

1:52So these two things, titans and mirais, they're tightly coupled. Very tightly. You can think of it this way. The titans architecture is the specific functional tool. It's the design, the biologically inspired model built on deep neural networks. And miras. Miras is the overarching theoretical blueprint. It's a unifying theory for sequence modeling itself. And crucially, both of them are completely focused on advancing this one core concept. Test time memorization. Tip time memorization. That phrase is, it's really the fundamental concept here. These systems allow an AI model to fundamentally update and integrate new knowledge into its core structure while it's actively running.

2:34Right, while it's processing new data during inference. And this is the key. It does it without needing those traditional, incredibly expensive, dedicated offline retraining cycles that we all rely on right now. It's real time. It's dynamic adaptation. It's a game changer. Okay, let's unpack this. To really get the solution, I think we have to deeply understand the problem. So let's start with the status quo, the mighty transformer. The transformer. So the core of that architecture, the thing that made everything else possible, is the attention mechanism. Right. It completely revolutionized sequence modeling because unlike older models that had to process data one piece at a time sequentially, attention lets the model look back at all the earlier inputs in a given context.

3:15All at once. All at once. And it decides which of those pieces are the most important for making the current prediction. It calculates dependencies across, you know, vast distances within that context window. And that grasp of non-local relationships, that's why LLM sounds so intelligent, right? They're not just predicting the next word based on the last five words. No, not at all. They're synthesizing the meaning of the current sentence based on something that was said three paragraphs ago. It's the mechanism of superior in-context learning. Exactly. But, and this is the big but, that phenomenal power comes with an immediate and frankly catastrophic computational bottleneck as soon as you try to scale it.

3:54Okay, let's get into that. Because attention has to calculate the relationship between every single token and every single preceding token in the sequence, the resource consumption just explodes. We're talking about computation time, FLOPs, and memory. It increases drastically with the sequence length. Can you walk us through what that increase actually looks like? Because I think people hear the term ON2, complexity or quadratic scaling, but it's hard to really grasp the severity of that in practice. Sure. Think of it this way. N is just the number of tokens you're feeding into the model. The length of the input.

4:28The length of the input. Right. If you're using standard transformer attention, the computation grows with N squared. So if N is, say, 10 ,000 tokens, N squared is 100 million computations. Which is manageable. Manageable for modern hardware, sure. But what if N is 100 ,000 tokens, which is still a relatively short document? N squared jumps to 10 billion computations. And if you try to feed the model a million tokens, which you absolutely need for a full legal document or scientific paper, you're now talking about 1 trillion computations. A trillion, except for the attention layer. Just for the attention layer.

5:03And that right there is the ultimate blocker. Quadratic scaling means if you double the length of your input sequence, the computation doesn't just double. Not even close. It doesn't triple quadruples. And that just makes scaling for these massive data streams economically and technically. Yeah. Well, it's just unfeasible for a standard, dense attention transformer. And this is so frustrating because the need for long context is absolutely immense. We need full document understanding where your input sequences are in the millions of tokens. We need long-term time series forecasting for, say, financial modeling or climate modeling.

5:41And maybe the most demanding application of all is complex genomic analysis. Modeling DNA requires handling these unbelievably long, incredibly detailed sequences with perfect accuracy. And ON2 just makes all of those tasks a non-starter. A complete non-starter. So the research community, of course, they didn't just sit still. We've seen a lot of alternatives trying to solve this tradeoff between expressive power and efficiency. The move toward linear recurrent models was very popular, wasn't it? It was, and it produced some really exciting models. The key idea was to move away from that dense all-to-all attention and go toward more efficient linear recurrent neural networks, or RNNs, and also state space models, or SSMs.

6:23And the most notable recent breakthrough there would be Mamba 2. Mamba 2, exactly. And these models are incredibly fast. They achieve linear scaling, ON. So if you double the input length, the computation just doubles. Well, it just doubles, which is a massive, massive efficiency win. Okay, so here's a question then. If Mamba is linear, which is so much faster, why couldn't the researchers just use a much, much larger fixed vector in Mamba to capture more detail? What's the fundamental constraint that stops a huge fixed vector from having the same expressivity as a transformer's attention. That is the critical and maybe non-obvious point.

6:59It isn't just about the size of the memory. It's about how the information is encoded. Okay. The incredible efficiency of these linear models is achieved by compressing the entire history of the input sequence into a single fixed-size hidden state or vector. A single vector for everything that's come before. Everything. And even if you make that vector enormous, the model is still fundamentally linear in how it relates the past to the present. So the argument in the source material is that linear models, it doesn't matter what the vector size is, they fundamentally struggle to capture the rich, detailed, and highly non-linear abstracted relationships that define these complex, long sequences.

7:38So the fixed size compression is the fatal flaw. It's the fatal flaw. It's like trying to summarize a library full of dense, multi-layered philosophical books onto a single, giant, but ultimately linear note card. You get the speed, but you lose the expressive power you need for that deep, long context recall and synthesis. So for years, we've been stuck in this impossible tradeoff. You have transformers, which are expressive but scale quadratically and are just too slow for millions of tokens. Right. And then you have RNMs and SSMs, which are fast. They have that beautiful linear scaling. But they lack the memory richness and expressive power needed for really deep, really complex, long context.

8:18Precisely. And that, that is the genesis of the Titans of Marius approach. The researchers sought to combine the speed of recurrence, that linear ON scaling, with the deep abstraction and accuracy of attention-based models. And they did it by completely changing the architecture of memory itself. By moving beyond that single fixed size vector. Okay, let's dive into that. The core structure of Titans, the design philosophy here is just brilliant because it's rooted in biology and psychology. It is. It mirrors how the human brain works with its distinct interconnected memory modules. Right. You have short-term memory, working memory, long-term memory.

8:58The model doesn't just treat memory as one big monolithic system. And the research really meticulously breaks down the system into these functional components. This is especially clear in the MSC variant, which stands for memory as a context, and it proved to be highly effective. So what's the first component? First, you have the short-term memory. This is handled by the model's core processing stream. In the hybrid Titans variants, this is usually implemented with attention mechanisms. But not the full expense of attention. No, no. It's often constrained to a local sliding window attention, SWA, or a very limited context.

9:31This part of the system is really designed for that precise momentary recall. what just happened in the immediate vicinity of the current token. It's the model's working memory. And then we get to the really revolutionary piece, the long-term memory module. And this is the core innovation. This is what breaks that fixed-size state bottleneck we were just talking about. So instead of compressing history into a simple fixed vector or a matrix. The Titan's long-term memory is implemented as a deep neural network, specifically a multilayer perceptron, or MLP. So it's not just a storage unit? It is not just a storage unit.

10:05It is a processing unit. Why is that depth, that multilayer perceptron, so crucial compared to just having a giant vector? What does that get you? Well, the source material gives a really compelling theoretical argument for this. Deep memory modules are mathematically and theoretically just more expressive than linear or shallow models. A traditional fixed vector or a matrix-valued memory, it, it basically assumes that all the underlying dependencies in the historical data are linear. And that is a massive, massive simplification. Especially for something complex like DNA or a legal document. Exactly.

10:41Deep neural networks, on the other hand, are capable of universal approximation. They are strictly more expressive than linear models because they can encode the highly nonlinear abstraction and synthesis of past history. So it's about storing conceptual summaries, connecting distant ideas, synthesizing patterns. Not just storing raw keys and values. So a shallow memory might remember the dates of two separate events, but a deep memory could synthesize the relationship and the implications of those two events, even if they're a thousand tokens apart. Precisely. And the research really validated this through detailed ablation studies.

11:16They compared long-term memory modules that had the same capacity but different depths. What did they find? The modules with greater memory depth, specifically two or more layers, showed significantly lower perplexity in language modeling. And they demonstrated much better scaling properties as the sequence link moved into that extreme long context regime. So the depth isn't a luxury. It's not a luxury. It's essential for long-term competence. Okay, then there's the third component, persistent memory. This one feels a little different. It feels more like the model's skill set, separate from the actual running context of the sequence.

11:52That is a great way to put it. The persistent memory is a set of learnable parameters, but they are data independent. Meaning they're fixed during inference or test time? Exactly. This component is dedicated to encoding general task knowledge, the underlying rules, the grammar, the meta information about how a task should be performed. So if the long-term neural memory is fluid and contextual, it's always updating based on the input stream. The persistent memory provides the stable, input-independent foundation of the task itself. It allows the model to leverage these general learned skills without that knowledge being constantly overwritten or lost by the massive influx of new sequence data.

12:34Okay, so we have the three systems. We've got short-term attention, the deep long-term MLP, and the fixed persistent knowledge. Now we get to the dynamic part, the active learning. How does the model efficiently decide what is worth committing to that precious long-term storage, especially when it's processing millions of tokens a second? This part is arguably the most clever piece of the whole puzzle. It basically translates a human psychological phenomenon, how we prioritize what we remember, into rigorous mathematics. Right. The system is designed not to just passively log every single token.

13:08It has to actively decide what constitutes important information. And the research terms this the surprise metric. I love this concept. We remember the unexpected and we forget the routine. If my morning commute is exactly the same every single day, I don't retain those specific details. But if I see an ostrich crossing the road, that memory is locked in instantly. Exactly. The AI needs a trigger for that. And that's the perfect analogy. The AI's trigger for surprise is a mathematical measure of anomaly. The surprise metric is calculated by detecting a large difference between what the long-term memory currently expects based on its history.

13:43And what the new input is actually telling it. Precisely. And mathematically, this anomaly is measured by the gradient. It's the internal error signal of the neural network with respect to the input in what's called the associative memory loss function. So a large gradient means high surprise. A large gradient means the model's current expectation failed spectacularly, which signals a very high surprise. Okay, let's make that concrete. So if I'm feeding the model a thousand-page legal contract, and it expects the next input to be another legal clause, and it is. Then the memory's expectation aligns with reality.

14:17You get a low gradient, low surprise. So it skips permanent storage. Right. Why waste storage space on something that's routine continuation? But imagine that contract is being analyzed for sentiment and risk. Suddenly, in the middle of a really dense section about asset distribution, a new input token refers to an extreme unexpected change in market conditions. A sudden deep economic collapse, for example. That input dramatically violates the memory's current expectation for how the sequence should flow. The gradient, the surprise, is very, very high. And that anomaly gets prioritized. It's immediately prioritized for deep permanent storage because that high gradient is a signal that says this breaks the context, this is unexpected, and therefore this is vital information for future recall.

15:05And the selective update process has to be essential for efficiency. You can't just store millions of tokens in a dynamic high expressivity module without some kind of really rigorous gatekeeping. Absolutely not. And that selective process is refined even further by two crucial mechanisms that they borrowed from advanced optimization theory. They integrated them right into the recurrent update rule. And those are? Momentum and forgetting. Okay, let's start with momentum. So the update rule for the memory state, let's call it MT, is analogous to gradient descent with momentum. It doesn't just look at the momentary surprise, which is the current inputs gradient, GT.

15:41It also incorporates the past surprise or the momentum history, ST-1. Why is incorporating that past surprise so important? I mean, if the model just registered a high surprise event, shouldn't it ignore subsequent events that are related but maybe less surprising on their own? This is exactly where momentum shines. Let's go back to your banana peel anomaly from earlier. OK, in the middle of a financial report. Right. That's the high surprise event. The immediate next tokens might be something like warning, stop or hazard. Now, individually, those are relatively common words. So they'd have low individual surprise.

16:16Exactly. However, because the memory system just experienced that massive shock from the banana peel, the momentum turn carries the residual effect of that initial high gradient. This ensures that the relevant subsequent tokens that follow the big event also get captured and recorded into long-term memory. Ah, so it allows the model to memorize the entire related time frame of an event. Right. It prevents it from just immediately relaxing and forgetting the context that surrounded the anomaly. It's incredibly clever. The momentum term basically acts as a dynamic memory of surprise across time, and it's controlled by a data-dependent decay term.

16:53So if the context changes abruptly, say the document shifts from finance to poetry, that decay term can drop to near zero, and it effectively kills the prior momentum and prepares the memory for a whole new context flow. But if the tokens are still highly relevant, it stays high and maintains the focus. Exactly. That data dependency is what makes the whole mechanism adaptive. So the system isn't just reacting to a single moment. It's remembering the whole chain of events around the anomaly. Yeah. Okay, so what about the other mechanism, forgetting? I assume managing finite capacity means the model has to be pretty ruthless about discarding old, irrelevant information.

17:32It has to be. The model employs an adaptive weight decay mechanism, which acts as the forgetting gate. In any long recurrent memory system, especially one dealing with millions and millions of tokens, uncontrolled storage, just leads to saturation and degradation. The memory has to have a way to actively prune or clear out old information that's no longer useful. So this is the cleanup crew. The cleanup crew, exactly. This weight decay, or forgetting gate, manages the limited capacity of the neural memory. If a piece of information has proven to be irrelevant over a long sequence, the adaptive decay allows those memory parameters to shrink or clear out entirely.

18:10Making space for new, more surprising information. Right. And this concept is a generalization of the forget gates that you find in a lot of modern, efficient, recurrent models like Momba. The ablation studies confirm that the performance gains they got from this forgetting gate are just as important as the gains they got from the surprise-based remembering. Okay, but the technical efficiency point here is really important. Gradient calculations, especially for recurrence, sound computationally intensive. How do they make sure that integrating this deep, nonlinear, momentum-driven memory update stays fast and parallelizable?

18:44How do they keep that linear scaling advantage? They use some highly sophisticated techniques to tensorize the process. The source notes that by reformulating how the weights are calculated in the inner loop, they could use only massive matrix multiplication or mat-mull operations and sums. These are things GPUs are really good at. Things GPUs are incredibly good at, and crucially, they treated that momentum term, STT, as a linear recurrence that can be computed very quickly using parallel associative scan operations. That's the key. That is the key. By transforming what is an inherently sequential recurrent calculation into a parallelized matrix operation, they completely overcome the speed barrier.

19:24They maintain fast linear inference speeds, effectively getting the ON speed of an RNN while enabling the deep expressivity of that MLP memory. Okay, so now we need to make a theoretical leap. If Titans is the specific machine that's built to solve this memory problem. Then Miras is the theoretical foundation. The unified theory. The thing that proves that all major sequence models, past and present, are really just different ways of building a complex associative memory module. And this is a radical, really powerful paradigm shift. Mirus gives us this unified framework where we can stop seeing transformers and RNNs and SSMs as these totally disparate architectures.

20:04Right. Instead, we see them as just variations on the same fundamental optimization problem. Which is? How to efficiently integrate new information with old memories while preserving the essential concepts. And within this new framework, Mirus defines any sequence model through four sort of concrete design choices. Can you detail what those are? Sure. They're the fundamental knobs that you can turn on any memory-based system. First is the memory architecture. What is the actual structure that's storing the information? Is it a fixed vector, which is shallow? Is it a matrix? Or is it the deep nonlinear MLP that they use in Titans?

20:39Okay, that's number one. Number two is the attentional bias. This is the internal learning objective the model is optimizing for. It determines what the model prioritizes when it's integrating new data. Third. Third is the retention gate. This is the memory regularizer. Merace formally reinterprets mechanisms like weight decay or those forgetting gates we talked about as specific forms of regularization. They're just balancing the need for fast new learning against the stability of retaining past knowledge. And the last one. The last one is the memory algorithm. The specific optimization algorithm that's used to update the memory.

21:17Is it gradient based? Is it using specialized recurrent equations? That's the fourth choice. The power of this unified view, then, is that once you categorize all models this way, you immediately spot a massive shared limitation in, well, virtually all successful existing architectures. And that limitation is their over-reliance on mean squared error, MSE, or just simple dot product similarity for both the attentional bias and the retention objective. The Euclidean approach. A completely Euclidean approach to measuring memory loss and similarity. And why is relying on that Euclidean distance, or MSE, a restriction on the model's expressive power?

21:56Well, MSE applies a really steep squared penalty for any error. This means that if you have a massive data set and you get a single major outlier, a black swan data point or a big anomaly like our banana peel example, the MSE penalty is so high that the model is forced to overreact. It has to excessively skew its entire memory update just to accommodate that one single outlier. Which in massive, messy, real-world data sets must make the models really brittle. It makes them brittle and very sensitive to noise. It prevents the memory from forming stable, high-level abstractions. So Meris transcends this limitation.

22:32It provides a blueprint for designing memory systems that use non-Euclidean objectives and regularization, things that are informed by decades of advanced statistical and optimization research. Exactly. And the researchers used the Marist framework to design and test new variants that explicitly move beyond that standard MSE paradigm. They created three specific attention-free models just to demonstrate the flexibility of this theoretical leap. Okay, let's start with the first one, YAD. This one was designed to make the memory more robust to those messy data streams, right? That's right. YAD uses what's called Huber loss instead of the steep squared penalty of MSE.

23:08So mathematically, MSE penalizes errors quadratically. If the error is 10, the penalty is 100. Huber loss, on the other hand, is only quadratic for small errors. For large errors, for outliers, it switches to a much gentler linear penalty. So if the error is 10, the penalty might only be 10 or 20, depending on the threshold you set. And that gentler penalty is crucial when you're processing these colossal, messy documents where outliers are just. They're common. They're everywhere. A single typo, a minor anomaly, a surprising but irrelevant phrase. These things shouldn't destabilize the entire memory.

23:45And WIAD prevents the memory from catastrophically overreacting to that noise, which leads to much more robust memory updates. Okay, then we have Moneta. What was the design goal for this one? Moneta explored stability through mathematical strictness. It used really complex and strict mathematical penalties or generalized norms for both the attendance mechanism, what to pay attention to, and the forgetting mechanism, the retention gate. So the idea was to impose really disciplined, rigid rules on how the memory should be updated. Right, with the hypothesis that this would lead to a fundamentally more stable, powerful, and mathematically constrained long-term memory system.

24:22And finally, MEMRA, which focused directly on memory stability. MEMRA tackled stability by forcing the memory to operate like a strict probability map. And this is a very, very tight constraint. Every single time the memory state is updated, Memora ensures the changes are controlled and balanced. It guarantees a clean, stable, and theoretically controlled process for integrating new information. So it prevents memory drift. It prevents memory drift and ensures that the long-term context stays coherent, even after millions of updates. And the fact that these non-MSE variants actually delivered improved performance compared to the recurrent baselines, that proves that the Miras framework is more than just a theory.

Read the full transcript

25:02It's a powerful generative tool for building better AI memories. Okay, so if MERIS is the theory that proves we need deep adaptive memory, and the neural memory module is the long-term storage unit, the next critical engineering step is figuring out how to connect that deep recurrent memory module with the short-term attention mechanism of a hybrid transformer backbone. Right, and this is where the Titans family of architectures comes in. They propose three primary structural variants, and they're all categorized under this MERIA lens. And the underlying difficulty here is integrating a sequential recurrent component, the memory MLP, with a massively parallel component, the attention, without just destroying the linear ON speed advantage you work so hard to get.

25:46Precisely. So they explored three ways of integrating that long-term memory summary into the ongoing computation. And the most effective approach, especially for those extreme recall tasks, is MSC, or memory as a context. How does M &A ensure that the short-term attention leverages the deep memory without compromising speed? Well, in MSC, the memory module acts like a smart retriever. So when a new segment of data comes in, the model uses that segment to query the long-term memory module. And that module retrieves a synthesized, abstract, historical summary of the relevant past context. Okay. And crucially, the model then concatenates this retrieved historical summary along with the fixed persistent memory parameters with the current context segment.

26:30So it just sticks them together. Sticks them together. This concatenated extended sequence is then fed to the short-term attention module. And the advantage here is profound. Right. Because the attention layer is now operating on a sequence that explicitly includes the momentary context and the compressed relevant history that it just retrieved from the deep MLP. So attention can then perform its high-fidelity, high-expressivity calculations on this combined input. Allowing it to selectively attend to the retrieved historical context. And this mechanism proved to be the best performer in those long context recall tasks, like the NIH benchmark.

27:04That is a very elegant solution. The memory pre-processes the history into a digestible summary. And then the attention mechanism handles the complex synthesis of that summary with the immediate data. What about the second approach, MAG, or memory as a gate? MAG takes more of a parallel, multi-hat approach. In MAG, the core processing stream uses a local sliding window attention, or SWA, while the neural long-term memory module operates independently on the same input data stream. So you have two things running at once. Two things running in parallel. Both produce an output. One is focused on local precision from the SWA, and the others focused on long-term abstract synthesis from the MLP memory.

27:47And then they're combined. The outputs of these two independent processing streams are then combined using a nonlinear gating mechanism. This architecture sort of views the memory module as a fading memory, where the gate determines how much of that long-term synthesis should influence the final short-term prediction. And it performed well? It performed very well on generalized language modeling tasks. Okay, and the third option, MAL or memory as a layer. This one suggests a sequential stacking that sounds less effective. It is, generally. Memory as a layer is the approach that's most common in a lot of prior hybrid models, where you just stack recurrent layers before or after a transformer block.

28:27So in MAL, the memory layer compresses the past and current context before the attention module ever sees it. Exactly. And the problem with the sequential stacking, as the source material confirms, is that it generally results in poorer performance compared to MA and MAG. And why is that? Well, because it limits the complementary processing power that a parallel system like MAG offers. And it removes the crucial ability of the attention module to selectively attend to the retrieved historical context, which is the genius of MA. ML forces the attention module to process an already heavily compressed and potentially lossy representation of the past.

29:04So across all these variants, MA, MAG, and MAL, the core concept of test time memorization remains unique. We should probably reiterate this operational difference during inference. We absolutely should. During inference, or test time, you have to distinguish between the two types of knowledge. The persistent memory parameters, ones that hold the task knowledge, are fixed. They are not updated. Right. However, the long-term neural memory parameters are actively learning and memorizing during inference. As the data stream flows in, that deep MLP memory continues to update and absorb new information using the surprise metric.

29:39And that is the definition of a truly dynamic meta-in-context learner. That's it. So the theory is sound, the architecture is elegant, but the truth is always in the data. The experimental results comparing titans against the state-of-the-art linear recurrent models like Mamba 2 and standard transformer variants, they're extremely compelling. The performance victories were comprehensive. In general, language modeling tasks on foundational data sets like C4 and Wikitext, and then on reasoning tasks like PIQA and Hellaswag, the Titan's architectures consistently outperformed the state-of-the-art linear recurrent models of a comparable size.

30:15Do you have a specific data point for us? I do. Comparing 760 million parameter models, the Titan's MAC architecture achieved a perplexity score of 19.93. MAMA, which is a state-of-the-art linear model, lagged significantly behind at 28.1 HALF. Wow, that's a big difference. It's a huge difference, and even a highly optimized transformer baseline only achieved 21.18. And remember, lower perplexity means a better, more accurate understanding and prediction of the text. This really confirmed that the deep memory expressivity was delivering real, tangible gains in comprehension. But the most important finding, the one that validates this entire effort to overcome ON2, has to be the performance on extreme long context recall.

31:00This is the regime where standard models are just completely ineffective. This was tested most rigorously using the Babylon benchmark. It's specifically designed to assess a model's ability to perform complex reasoning across facts that are distributed throughout extremely long documents. So the model needs to connect the dots between pieces of information mentioned hundreds of thousands of tokens apart. Exactly. And the results here were, They were truly stunning. They were. Titans, specifically the MSC variant, was trained using the fine-tuning setup, and it achieved significantly higher accuracy across these long sequence lengths than even highly sophisticated cutting-edge baselines.

31:37Like what? Specifically, it outperformed models including GPT-4 and LAMA 3.1, even when these powerful models were augmented with retrieval augmented generation, or R. Okay, let's pause on that implication for a second. Titans, a model with vastly fewer parameters than something like GPT-4, outperformed it on a critical long-context reasoning cask that was designed to test memory synthesis. And it outperformed a R-augmented LAMA 3.1. So why did R fail where Titans succeeded? Well, this highlights the difference between passive memory retrieval and active memory synthesis. RA systems, they retrieve discrete chunks of information.

32:14They're excellent for factual lookup. Find me this paragraph. Exactly. But the Babylon tasks require abstract reasoning and the synthesis of complex relationships across those facts. Titan's deep MLP memory is designed not just to store chunks, but to synthesize abstract concepts and relationships over time. It's an integrated memory, not an external database. And that superiority in synthesis allows it to connect those distant facts better than a retrieval-based system can. And they pushed the scaling limit itself. They did. The research explicitly states that Titans demonstrated the capability to scale effectively to context window sizes larger than 2 million tokens.

32:532 million. 2 million. To put that in perspective for you, standard commercial transformers typically cap out around 128 ,000 tokens before the resource consumption just becomes prohibitive. 2 million tokens is a paradigm shift. It proves that the ON speed advantage is fully realized even with that deep memory integration. So the deep neural memory module, learning in context, is clearly superior to the shallow vector-based memory, like the models that compress history into a small fixed vector size. Absolutely. And the success wasn't just limited to standard language modeling. Titan showed a really impressive versatility.

33:27It outperformed specialized baselines in a bunch of different areas. For instance. For instance, in genomic modeling, DNA tasks. The LMM variant was competitive with state-of-the-art specialized architectures like hyena DNA. This demonstrates its ability to handle highly structured, non-linguistic, long sequences. And time series forecasting. Also, time series forecasting. The neural memory module, when they integrated it into the Simba framework, significantly outperformed transformer-based, linear-based, and even Mamba-based architectures across major benchmark data sets for financial and environmental forecasting.

34:04The fact that these fundamental principles, deep adaptive memory, and the momentum-based surprise metric apply across such diverse data types, it suggests this isn't some niche language breakthrough. No, it's a foundational shift in how we do sequence modeling. And we keep circling back to efficiency. But it's the key win, achieving transformer-like expressive power with RNN-like linear inference beads. The efficiency ablation studies confirm that while deeper memory gives you better scaling and lower perplexity, there is a small trade-off in training throughput. A small one, yes. But the overall architecture is a massive net positive because it unlocks tasks that were previously completely inaccessible because of quadratic scaling.

34:45And the finding that the largest performance gains came from the momentum and the weight decay mechanisms. I mean, the mechanisms of remembering the context of a surprise and actively forgetting irrelevant data, that is just a powerful statement on the nature of intelligent memory. It really is. So to summarize the critical contributions this research makes, I think we can define three major shifts that Titans and Miros represent for the future of sequence modeling. Okay, lay them out for us. First, deep memory is paramount. The move beyond shallow fixed vector or matrix memory to deep nonlinear neural networks MLPs is absolutely essential for effective abstraction, synthesis, and meaningful scaling in extremely long sequences.

35:27This is the difference between an AI that just stores data and one that truly understands and synthesizes complex relationships across time. Okay, that's number one. Second, dynamic adaptive memory management is the key to efficiency. The gradient-based surprise metric combined with momentum to capture context flow after an anomaly and adaptive weight decay for controlled forgetting. This provides memory management that is exponentially superior to existing models that are either static or rely on these really rudimentary momentary updates. Number three. Third, theoretical unification via Miras allows for robust non-Euclidean memory objectives.

36:03By viewing all of sequence modeling as associative memory optimization, Miras opens the door to using advanced mathematical concepts. It moves model design beyond the limitations of standard mean squared error, and that leads to inherently more stable and robust models capable of handling the noise and outliers of vast real-world data streams. So if we connect this all to the bigger picture for you, our listener, what does this breakthrough fundamentally change about future AI applications? This means that future AI systems are now structurally equipped to handle truly vast, specific, and often messy data streams instantly and adaptively.

36:40Imagine an AI legal assistant that can read five years of changing case law and complex financial filings. Right. I'm reading one new anomalous detail in the latest report. It immediately integrates that detail into its overall synthesized five-year knowledge base. It changes its entire risk assessment in real time. The era of context limits being measured in thousands of tokens is truly ending. We are now firmly in the realm of millions with the capacity for fluid, almost human-like long-term memory integration. So we've seen how the Titan's architecture and the Miras framework, through this biologically inspired design-dividing labor between short-term attention and a deep, actively learning long-term memory module, have successfully combined the speed of linear recurrent networks with the expressive synthesizing power of transformers.

37:27And here's a final provocative thought for you to consider based directly on the research. The paper mentions that Titans are capable of solving problems that are classified as being beyond TC0. Okay, you have to clarify that highly technical term. I do. So TC0 is a fundamental theoretical barrier in complexity theory. It represents the class of problems that are solvable by extremely fast, shallow circuits, so essentially simple, repetitive pattern recognition and low-level computation. Standard transformers and most diagonal linear recurrent models are fundamentally constrained to this class.

38:00They can't solve problems harder than TC0. But titans can. But titans can. And this raises an important question. If deep adaptive memory enables models to step beyond this fundamental theoretical barrier and perform mathematically richer, deeper, more interconnected computations, computations we previously thought were inaccessible to standard architectures, what new higher level cognitive tasks will this enhanced memory capacity unlock next in AI? What problems that we currently consider impossible for AI to solve will suddenly become routine. Exactly. That is a fascinating frontier. It's not just about scaling anymore.

38:36It's about fundamentally changing what the AI is capable of computing. Thank you for joining us on this deep dive into the future of AI memory. Keep exploring the intersection of deep learning theory and biological inspiration.

From the publisher

We explore Google's Titans and the MIRAS framework, a new paradigm in sequence modeling that replaces static context compression with active test-time learning. We discuss how Titans utilize deep neural memory modules to update parameters on the fly using a gradient-based "surprise metric," prioritizing unexpected information for long-term storage. We cover the theoretical MIRAS blueprint—which unifies sequence models through attentional bias and retention gates—and introduces robust new architectures like Moneta, Yaad, and Memora. We discuss how these models effectively scale to context windows exceeding 2 million tokens, outperforming GPT-4 and Mamba on complex long-context reasoning tasks.

More from Best AI papers explained

All 475 episodes
Beyond the Transformer: Titans, MIRAS, and the Future of Infinite ContextBest AI papers explained · 39 min
Listen in VO