MEMO: Memory as a Model

24 May 2026 · 18 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

MEMO (“Memory as a Model”) proposes separating an AI’s frozen reasoning (“executive model”) from a swappable, dedicated memory (“memory model”) so systems can learn new facts without retraining or losing prior skills.

Guest backgrounds

No guest identities or biographies are provided in the transcript; the episode is presented as a discussion between hosts.

Key claims

Current updates rely on RAG (context-window limits and retrieval noise), fine-tuning (catastrophic forgetting), or latent memory (representation coupling). MEMO uses a smaller memory model (e.g., 1.5B/14B QN) plus a five-step synthesis pipeline (fact extraction, consolidation, verification/rewrite into self-contained QA, entity surfacing to prevent reversal curse, cross-document synthesis). It then uses a three-stage multi-turn protocol (grounding, entity identification, answer synthesis). Benchmarks: 53.58% narrative QA accuracy vs RAG; with distractors, RAG drops up to 6.22% while MEMO stays stable. Continual updates via model merging reduce compute ~33% versus full retraining.

Notable examples

“Open book test” context overflow; catastrophic forgetting like forgetting phone number after cramming; reversal curse (knowing “Tom Cruise starred in Mission Impossible” but failing when reversed); pronoun rewriting (e.g., “Where did Earl go? Earl went to the store.”).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding MEMO and Its Importance

0:58 to 1:49

Learn about the MEMO architecture and its innovative approach to AI learning.

“And it's a new architecture designed entirely to teach an AI new facts without having to completely rewire its frozen brain.”

Current Workarounds for AI Knowledge Updates

1:49 to 3:00

Discover the three major workarounds for updating AI knowledge and their limitations.

“The first method is nonparametric updating.”

Challenges of Nonparametric Updating

3:00 to 3:35

Examine the flaws in nonparametric updating and the concept of retrieval noise.

“You can only feed it so many external documents before it gets completely overwhelmed and starts, you know, forgetting the beginning of the text.”

The Costs of Parametric Updating

3:35 to 4:42

Understand the financial and technical drawbacks of parametric updating in AI.

“It really struggles to connect the dots, especially if the real clues are scattered across 20 different documents.”

Limitations of Latent Memory in AI

4:42 to 6:32

Learn about latent memory and its coupling issues affecting AI memory transfer.

“Imagine you're studying for a ridiculously complex medical exam.”

Introducing the Memo Architecture

6:32 to 8:08

Explore the unique separation of thinking and memory in the MEMO architecture.

“You can't just take that compressed memory and plug it into any AI.”

The Five-Step Data Synthesis Pipeline

8:08 to 13:16

Understand how the data synthesis pipeline prepares AI memory for effective learning.

“Think about why this matters to you, the user.”

The Interaction Between Executive and Memory Models

13:16 to 14:00

Learn how the executive model interacts with the memory model to retrieve information.

“The executive receives your complex query, and instead of throwing the whole thing at the memory, it breaks it down into atomic, highly targeted sub-queries.”

Understanding Memo's Performance

14:00 to 17:00

Explore how Memo processes information and its superior performance in noisy environments.

“Knowing definitively who or what the subject is, the executive thinker gathers the final, highly specific supporting facts and synthesizes a fluid, natural language response for you.”

The Future of AI Memory Models

17:00 to 17:49

Discuss the implications of customizable memory modules for AI systems.

“Memo offers a scalable, noise-proof plug-and-play memory bank that lets any AI learn new tricks without forgetting the old ones or getting confused by irrelevant noise.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Have you ever noticed just how incredibly fast AI models become outdated? Oh, absolutely. It's almost immediate. Right. I mean, it is honestly a little wild to think about. The world is constantly changing around us. You know, new events are happening every single second. But an AI's brain is effectively frozen in time the exact moment its initial training run ends. Yeah. And that frozen state is, well, it's one of the most massive bottlenecks in the tech industry today. We've really come to expect these systems to act like omniscient assistants. Totally. You expect it to just know things. Exactly.

0:35But they are fundamentally static artifacts once the engineers hit stop on that initial multi-million dollar compute run. They don't learn from merely existing in the world the way you or I do. Which brings us to the topic of our deep dive today. We are exploring a fascinating new piece of research that introduces a radically different framework. It's called MEMO. Right, which stands for memory as a model. Yes, memory as a model. And it's a new architecture designed entirely to teach an AI new facts without having to completely rewire its frozen brain. It really represents a profound shift in how we build these systems.

1:10Historically, you know, the industry has tried to force a single model to act as both the reasoning engine, the part that thinks, and the storage hard drive. The part that remembers. Right. And Memo splits those functions entirely apart. Okay, let's unpack this, because our mission today is to understand exactly how memo works by separating the thinker from the memory. It's a huge distinction. It is, and we want to explore why this could completely change how you interact with AI in your daily life. But to really appreciate why memo is such a massive breakthrough, we first need to understand why updating an AI's knowledge right now is such a headache.

1:48Well, the research highlights that the industry has essentially been relying on three primary workarounds to deal with this frozen brain problem. And none of them are perfect, right? Far from it. Every single one hits a severe wall. The first method is nonparametric updating. In everyday terms, you might have heard this called RAG, or retrieval augmented generation. Right. So that is basically the open book test approach. That's a good way to put it. Yeah. Instead of expecting the AI to have the knowledge memorized, you just like handed a stack of external documents to read right before it answers your specific question.

2:20That is the goal. And it's popular because you do not have to alter the AI's underlying code at all. But it faces severe physical limitations, primarily what we call context window limits. Let's break that down for someone who isn't building AIs every day. The context window is essentially the AI's short term working memory, correct? Think of it exactly like human short-term memory. If I hand you a five-page report, you can hold most of those details in your active mind while we discuss it. Sure, that's easy enough. But if I hand you a thousand-page encyclopedia and ask you a nuanced question about page 3 and page 900...

2:58Oh, I'm definitely losing the thread. Right, you run out of working memory. And AI does the exact same thing. You can only feed it so many external documents before it gets completely overwhelmed and starts, you know, forgetting the beginning of the text. And on top of that, there is the issue the paper refers to as retrieval noise. Yes. Retrieval noise is the Achilles heel of the open book method. When you run a search to find documents to hand the AI, that search inevitably pulls up distracting irrelevant information. Because search engines aren't perfect. Exactly. So if the AI is trying to solve a complex puzzle and you hand it 10 relevant clues mixed with 5 pieces of total junk data, the reasoning engine gets confused.

3:39It struggles to filter out the noise. Immensely. It really struggles to connect the dots, especially if the real clues are scattered across 20 different documents. It's like asking a detective to solve a mystery, but handing them the actual evidence mixed in with like old takeout menus and random newspaper clippings. That is a perfect analogy. They just get lost in the clutter. Yeah. So if handing the AI external reading material is flawed, what is the second workaround? The second approach is parametric updating. which basically means fine-tuning the model. You take the new information, and you forcefully bake it directly into the AI's core mathematical parameters.

4:17By running a new, intense training phase. Yes, exactly. Which is incredibly expensive. I mean, prohibitively expensive. Oh, the compute costs for retraining massive models are astronomical. But beyond the financial cost, the research points out a severe technical side effect known as catastrophic forgetting. I love that term. Well, I mean, I don't love what it does, but catastrophic forgetting is such a visceral term. It really paints a picture. It does. Imagine you're studying for a ridiculously complex medical exam. You're cramming so much dense, highly specific new jargon into your brain for weeks on end.

4:55And then what happens? You finally sit down to take the test, and you realize you have suddenly forgotten your own phone number. Or, like, how to ride a bike. That captures the phenomenon perfectly. The model adapts to the new data, but to make room for it in its fixed neural network, it overwrites its prior capabilities. So it learns the new facts, but it loses its mind in other ways. Literally. It suddenly forgets its foundational safety alignments, or it loses its ability to perform basic logical reasoning. You gain new knowledge at the direct cost of destroying fundamental skills. Wow. So the open book test gets messy and overwhelmed, and retraining the brain gives the AI amnesia.

5:33What about the third workaround? So the third method discussed in the research is latent memory. This is a bit more abstract. Instead of storing knowledge as raw text, the system compresses the new information into what they call soft tokens. Let's clarify soft tokens. Instead of feeding the AI words that you and I can read, the system is translating that text into a dense mathematical barcode, basically. Exactly. A barcode that only the AI understands. It is a highly compressed, machine-readable format. And I imagine that compression makes it incredibly efficient to store. It does. But the research identifies a fatal flaw in this approach, too, which is representation coupling.

6:13That mathematical barcode is intricately tied to the specific architecture of the model that created it. Oh, I see. It's like buying a proprietary hard drive that only plugs into one specific brand of laptop from one specific year. Yes, exactly. If you buy a different computer or even upgrade to next year's model, that hard drive is completely useless. You can't just take that compressed memory and plug it into any AI. You really can't. And because of all these friction points, you know, context windows overflowing, catastrophic forgetting, destroying logic and latent memory being completely untransferable, the team behind Memo asked a brilliant question.

6:51Which was? What if the memory didn't have to be forced into the thinker? And what if it didn't have to be external text either? What if the memory was its own entirely distinct swappable entity? And this is where we get to the core architecture of Memo. It operates using two separate models working in a partnership. Correct. First, you have the executive model. Which is the thinker. Right. It is the frozen, highly capable reasoning engine. It could be one of the massive proprietary black box APIs that everyone knows, like Gemini 3 Flash. So its only job is to process logic. It doesn't store the new facts itself.

7:25Exactly. The executive model brings the deductive skills, the conversational ability, and the logic. And working alongside it is the second piece of the puzzle, which is the memory model. And the memory model is not some massive trillion-parameter behemoth, right? Not at all. The memory model is designed to be a smaller, dedicated model. The research uses models with a fraction of the parameters of the big systems, like a 1.5b or 14b parameter QN model. Something lightweight enough that it doesn't require a supercomputer to run. Precisely. Its entire structural purpose is to be specifically trained to encode and hold new knowledge.

8:03It acts as a highly intelligent, interactive database that the executive model consults. Think about why this matters to you, the user. The implications here are huge because it makes the system entirely plug and play. It really does. You don't need access to the underlying code or the weights of a massive corporate AI to give it a custom memory upgrade. You just build this small, dedicated memory model, and any advanced reasoning AI can be plugged into it. The separation of concerns is elegant, but actually building that memory is a pretty complex technical hurdle. Wait, if we are just taking a small model and training it on new text, isn't it just going to blindly memorize sentences?

8:40I mean, regurgitating facts isn't the same as understanding the context. How does the memory model actually absorb the data? You hit on the exact reason why you can't just feed raw text into the memory model. If you do that, it just acts like a parrot. Right. Totally incapable of answering complex queries later. Exactly. So to solve this, the research outlines a novel five-step data synthesis pipeline. Before the memory model ever sees the new information, a separate generator model digests the raw text and turns it into a data set of what they call reflections. Reflections. That implies the model is actually pondering the text, breaking it down into fundamental truths rather than just, like, scanning paragraphs.

9:20Yes, it is generating compositional representations of the facts. The pipeline takes the raw text through a pretty fascinating journey. Step one is fact extraction. So pulling out explicit facts. Explicit facts, of course, but it also pulls out indirect, implicit facts that require a degree of logical inference to identify. It's gathering all the raw materials, basically reading between the lines to find what is actually being said. What's step two? Once it has that massive pile of facts, it moves to consolidation. The pipeline analyzes the entire collection and merges redundant information. So if you're feeding it a novel and three different chapters mention a character named Earl lives in London, it doesn't need to memorize that three separate times.

10:02Right. It consolidates those instances into a single cohesive representation. That keeps things clean. But those facts need to be formatted in a specific way for the memory model to actually use them, right? That happens in step three. Verification and rewriting. This is critical. The goal here is to turn these facts into question and answer pairs that are entirely self-contained. How does a self-contained pair differ from a normal QA pair? Think about ambiguous pronouns. A normal system might extract a fact and generate a pair that says, Where did he go? He went to the store. Which is useless in a vacuum.

10:35Who is he? Exactly. The rewriting phase forces the system to replace the pronouns with the actual entities. It rewrites it to, Where did Earl go? Earl went to the store. Every single piece of knowledge has to make complete logical sense in total isolation. It is essentially a student creating their own highly advanced, perfectly phrased flashcards. If you just highlight a textbook, the information is still locked on the page. Right, but making the flashcard forces you to encode it. Exactly. Now what's step four? Step four is entity surfacing. The system actively makes key entities in their relationships explicit.

11:14This is a specific mechanism engineered to prevent a really common AI failure known as the reversal curse. Oh, the reversal curse. is so frustrating. That's when an AI learns a fact perfectly in one direction, but blanks if you reverse it. Yep. It knows that Tom Cruise starred in Mission Impossible, but if you ask the exact same model who starred in Mission Impossible, it hallucinates. Or just says it doesn't know. Right. Standard models struggle with bidirectional relationships. Entity surfacing trains the memory model to recognize entities from indirect descriptions from any angle. Curing that directional amnesia.

11:49Awesome. And step five. Step five is the pipeline's crowning achievement, cross-document synthesis. This is the most crucial step. This is where the magic really happens. How does a small model link facts that aren't located anywhere near each other? The pipeline mathematically forces the model to integrate evidence across entirely different documents. It scans for parallel properties between entities that might be located in completely different files. So it builds a web of connections before the user ever asks a question. Exactly. And the data from the research underscores how vital this is.

12:22If you remove step five, basically turning off that cross-document synthesis, the model's accuracy absolutely collapses on complex tasks. It's a huge drop. Yeah, on the narrative QA data set, it dropped from 24 % to a dismal 6.37%. Yeah. Without that web of connections, the memory is just disconnected trivia. The memory model has to do the heavy lifting of connecting the dots beforehand. hand. And once this pipeline finishes, it has internalized its flashcards perfectly. Okay, so the flashcards are built. But when you sit down and ask a really complex question, how does the executive thinker actually get the right flashcard out of the memory model?

13:01You can't mind read. No, it relies on an interactive multi-turn protocol. It's a three-stage inference pipeline. The executive model and the memory model actually have a back and forth conversation under the hood. A conversation between models. Walk us through that. Stage one. Stage one is grounding. The executive receives your complex query, and instead of throwing the whole thing at the memory, it breaks it down into atomic, highly targeted sub-queries. It probes for initial clues. This is exactly like playing a game of guess who or 20 questions. That's a great way to think about it. Right. The thinker doesn't just ask for the whole answer at once.

13:37If I'm playing guess who, I don't yell out, isn't the guy with glasses and a red hat who works in finance? I start small. Do they wear glasses? Exactly. It asks targeted questions to narrow down the suspect, which leads to stage two, entity identification. Narrowing it down based on the clues. Right. The executive model iteratively narrows down the candidates. It asks a question, gets a clue, formulates a new specific question, and loops until it converges on the exact entity your prompt is asking about. So it confirms the target. Yeah. And then stage three. Answer seeking and synthesis. Knowing definitively who or what the subject is, the executive thinker gathers the final, highly specific supporting facts and synthesizes a fluid, natural language response for you.

14:23It's an incredibly elegant solution. But here is the ultimate test. A game of 20 questions between AIs sounds clever in theory, but does it actually perform better than just handing the AI a stack of documents? Well, the empirical validation in the research is exceptionally strong. They tested Memo on brutal benchmarks like narrative QA, which involves understanding whole books or movie scripts, and music, which demands multi-hop reasoning. Multi-hop reasoning meaning connecting disparate facts to find an answer. Right. And paired with Gemini 3 Flash, Memo achieved 53.58 % accuracy on narrative QA.

14:57Which is huge. That completely destroyed steed-of-the-art rag methods like HipRag 2. It really did. But the most revealing metric isn't just about reading books, it's about how the system handles garbage data. The noise factor. Oh, right. How does memo handle being handed irrelevant distracting information? The researchers deliberately injected retrieval noise. They flooded the system with distractor documents that looked relevant but were actually junk. And what happened to the baselines? The standard RAG baseline suffered massive performance drops, up to 6.22 % drops on the BrowseTomp Plus benchmark the moment distractors were added.

15:32They took the bait and got confused. But Memo? Memo remained completely stable. Its accuracy barely twitched. Think about your daily life. If you hate when an AI hallucinates or gives you a wrong answer because it got distracted by a random, irrelevant search result, Memo is the cure. Because it deeply internalized the connections during training, it inherently filters out the noise. I am completely sold on this. But wait, what happens when more new information comes out tomorrow? Do we have to rebuild this entire memory model from scratch every time the world changes? That would be a logistical nightmare.

16:05But the architecture bypasses this through a concept called continual knowledge integration via model merging. Model merging. Instead of retraining on the old and new data combined, you do what? You just train a brand new memory model exclusively on the new data and then merge their task vectors directly in the parameter space. A task vector being a mathematical representation of what it learned. Exactly. Techniques like ties or dare linear allow engineers to mathematically blend those vectors. This cuts compute costs by 33 % when merging to corpora. And that savings scales up dramatically the more you add.

16:42It is exactly like downloading a lightweight software patch for a video game rather than having to buy and reinstall the entire 100 gigabyte game every time there's a small update. Perfect analogy. And while there is a slight accuracy drop compared to full retraining, the merged model still outperforms all the standard retrieval baselines. So let's wrap this up. Memo offers a scalable, noise-proof plug-and-play memory bank that lets any AI learn new tricks without forgetting the old ones or getting confused by irrelevant noise. It's a massive step forward. And it leaves us with a really fascinating thought to mull over.

17:16What's that? Well, if AI can now utilize distinct, swockable memory modules tailored to specific subjects like plugging in a medical memory drive or a legal memory drive, what happens to the concept of a single omniscient AI? Oh, wow. Will we soon be curating and trading personalized memory models for our AIs, essentially choosing what reality our personal assistants remember? That is wild. We are essentially giving it the ability to learn and adapt exactly what we need it to right as the world happens. Talk about a completely different way to think about the future of intelligence.

From the publisher

MEMO (Memory as a Model), a modular framework designed to integrate new, domain-specific knowledge into Large Language Models (LLMs) without the need for expensive retraining. By encoding information into a dedicated, smaller MEMORY model while keeping the primary EXECUTIVE model frozen, the system avoids catastrophic forgetting and remains compatible with proprietary, closed-source models. The process involves a five-step data synthesis pipeline that converts raw documents into a structured question-answer dataset of "reflections" that capture complex, cross-document relationships. At inference, the EXECUTIVE model retrieves information through a structured multi-turn protocol, decomposing difficult queries into targeted sub-questions. Empirical results across multiple benchmarks demonstrate that MEMO is more robust to retrieval noise than standard methods and achieves superior performance by leveraging internalized parametric knowledge. Furthermore, the framework supports continual knowledge integration through model merging, allowing new data to be added efficiently while maintaining a retrieval cost that is independent of the overall corpus size.

More from Best AI papers explained

All 475 episodes
MEMO: Memory as a ModelBest AI papers explained · 18 min
Listen in VO