Memory Management for AI Agents (The Agents Season, Episode 4)

10 May 2026 · 25 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Memory management for AI agents constrained by finite LLM context windows, covering MEMGPT (virtual-memory-style paging), RAG vs agent-controlled retrieval, compaction (e.g., Claude Code-style summarization), context hierarchy (system prompt, pinned project file like claude.md, then conversation), and failure modes like lost-in-the-middle and context rot.

Guest backgrounds

No guests mentioned; it’s a solo host episode.

Key claims

Agents should manage their own “RAM vs disk” boundary (MEMGPT) using tool calls (read/write external memory, retrieve, summarize/compress). Compaction enables long runs via rolling summarized history, but summaries are lossy and can cause context rot. Context is hierarchical: system prompt and durable project instructions are re-injected and treated as higher priority than chat history.

Notable examples

Analyzing long documents and multi-session conversations; Claude Code compaction with a claude.md file; “compaction creep” where earlier decisions get contradicted.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Context Windows

0:45 to 1:30

Exploration of context windows and their role in AI agents' functionality.

“very large context window, it's really the beginning and end that are going to be the most useful.”

The Challenge of Memory Management

1:30 to 2:32

Discussion on the limitations of context windows and the need for effective memory management.

“clawed code, which you might know from your day-to-day work.”

Introducing MEMGPT

2:32 to 3:41

An introduction to the MEMGPT concept and its relevance to AI memory management.

“So the question that they're asking in this paper is, what's the right way to think about this problem?”

Virtual Memory in Computing

3:41 to 5:05

Explanation of how virtual memory works in operating systems and its application to language models.

“It's immediately accessible, it's limited.”

Applying Paging Concepts to AI

5:05 to 6:28

How the principles of paging in operating systems inform memory management for AI agents.

“that you started yesterday or the day before.”

The Distinction of MEMGPT

6:28 to 8:07

Comparison of MEMGPT with traditional retrieval methods in AI tasks.

“That's what it's going to need, of course.”

Active Memory Management

8:07 to 9:40

Detailed look at how MEMGPT approaches active memory management in AI systems.

“I'm not making the decision about what to retrieve in the first place.”

Compaction in Memory Management

9:40 to 11:09

Exploration of the compaction process and its implications for AI memory use.

“All right, so let's go back to memory management in the context of the agent doing its work.”

Challenges of Summarization

11:09 to 12:51

Discussion on the potential downsides of summarization in AI memory, known as context rot.

“So if you're following along closely and you're wondering at this point, oh, so it's like the agent wrote itself a note.”

Importance of Preserving Key Information

12:51 to 14:00

How AI agents manage to retain crucial information during compaction.

“And to dive into the mechanics just a little bit more, because I think they are worth understanding, you can illuminate what's actually being preserved and what isn't.”
Show all 14 chapters

Understanding Context Reconstruction

14:00 to 15:43

Learn about the systematic reconstruction of context in AI agents.

“After that summarization has been generated, there's a systematic reconstruction of the context that happens.”

Hierarchy of Context in AI

15:43 to 19:38

Explore the hierarchy of context in AI systems and its implications.

“context more carefully than older context.”

Memory Management Systems Explained

19:38 to 22:48

Discover different memory management systems and their roles.

“The context window isn't this flat space where everything is on equal footing.”

Teaser for Next Week's Topic

22:48 to 23:03

Get a preview of next week's discussion on planning in AI.

“for any of you who are out there who are working with these systems day to day.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi, welcome to Linear Digressions. This is our fourth episode in a season that's all about AI agents. So if you haven't heard the first three, don't worry, you're going to get plenty out of this one. But if you have heard the first three, or especially the last one, this is going to be a really interesting episode. So what we know at this point is that AI agents are AI systems that go through this reasoning, acting, and observing loop. They can take actions and use tools. And everything that they have to know to do their job has to fit inside of their context window, which is a limited amount of real estate.

0:36Sometimes it can be a lot of real estate, but it's not an unlimited amount. And in particular, what we looked at last time was this idea that even if you have a very large context window, it's really the beginning and end that are going to be the most useful. Stuff that happens in the middle of that context window tends to get lost for architectural reasons. And so between both that lost in the middle and the fact that context is finite, it introduces this overall challenge of how to manage the memory and that context of an agent so that it has the information that it needs to effectively do its job, especially over really long time horizons or long-running complex tasks where there's a lot of things going on at the same time.

1:19That is what we're going to talk about today, the memory management of AI agents, starting with some of the foundational concepts from computing, but then pretty quickly getting into things like clawed code, which you might know from your day-to-day work. You are listening to Linear Digressions. So if you listened last week, you might have been wondering at the end of the episode, introduced this challenge of the context window as this technical constraint, this coming out of the architecture of the LLM. It has this structural bias towards the beginning and end of whatever the model can see, which effectively starves the middle of attention.

1:56So that's where we're starting out today. We have this constraint. What are we going to do about it? The first place we will visit is a paper out of UC Berkeley in late 2023 called MEMGPT. The idea is the following. Start with this simple observation that language models have this fixed context window. There's many useful tasks that require more information than a fixed context window can hold. So examples here might include analyzing a really long document, maintaining coherence across a conversation that stretches out over multiple sessions, keeping track of everything that's happening during a long-running agentic task.

2:34So the question that they're asking in this paper is, what's the right way to think about this problem? And they landed on a pretty clever answer, which derives from how operating systems solve the same problems for a computer. So computers themselves also have limited working memory. It's called RAM, and it's finite. but programs routinely need to work with data that's much larger than what you can hold in your RAM. So operating systems developed this concept decades ago at this point of virtual memory and that effectively makes it look like there's this large seamless memory space and the way that you achieve that illusion is between moving data between the fast RAM that's right at the front of what your computer has access to and then this slower disk storage that sits a little bit farther they're back behind the scenes.

3:22So programs don't have to do this. The operating system is handling this. It's called paging in and out. Operating system is identifying the data that needs to actively be loaded into RAM, pulls it in, pushes data that isn't going to be used, pushes that back to disk. And what memgpt does is it takes this architecture and applies it to language model agents. So think about it this way. The context window is RAM. It's fast. It's immediately accessible, it's limited. And you have this external storage, this disk. It's slower to retrieve from, but it's effectively unlimited. And the agent itself is managing the paging back and forth.

4:01So the agent is deciding what to load into its context, what to push out, and when it goes to retrieve something from storage. And then what's fun about this is we're starting to use some of the fundamentals that we've talked about in a couple of the previous episodes. In particular, how does the agent do this? It does it through tool use. So the agent has access to functions that read and write to its external memory, searches for relevant past information, and summarizing and compressing information from its context window before it moves it out to make space for the next piece of information that it has to pull into that context window.

4:39So it's not a separate memory system that's bolted on. It's the agent actually managing its own cognitive resources using the same tool calling architecture that it uses to do everything else. So in this UC Berkeley paper, the researcher tested this concept in two domains where the context limitations bite the hardest, namely analyzing long documents that exceed any reasonable context window and multi-session conversations. So those are the kinds of conversations where you might be coming back to the same conversation that you started yesterday or the day before. you have the expectation in that experience that the agent is basically remembering what you talked about yesterday, how are you going to achieve that?

5:17So in both these cases, MEMGPT, which is their operating system-inspired system, significantly outperformed standard approaches because it was actually maintaining continuity across the kind of length that will break when you have a fixed context model. I'm also going to introduce an idea at this point that's going to pay off a lot in a few minutes later in the episode, which is that in this operating system analogy, this isn't just a cute metaphor. This is doing some important conceptual work, namely that when you have RAM and disk, you don't just dump everything into your RAM and hope. You're making deliberate decisions, or the operating system is, about what's the hot data that you need.

5:59So this is the data that's frequently needed, keep it close. And then data that's colder, which is to say it's rarely needed, you can push it out, that's being managed by that operating system. And memgpt introduces the same distinction for agents. So privilege, important or immediately active data for the in-context memory, and the external memory is for everything else. And it's up to the agent to develop some intelligence about the difference. And it's actually a harder problem than it sounds because the agent doesn't always know in advance That's what it's going to need, of course. So at this point, a little bit of a digression, especially if you've been a longer time listener and you tuned into our episode a few months ago about RAG, retrieval augmented generation, which also introduces this idea of there being a knowledge store that the LLM, or if you want to think of it this way, the system the agent has access to.

6:52So it can go into that knowledge store, look for information that it needs to accomplish its task or answer its question, retrieves it, and then injects it into the prompt. for the LLM to have in context when it generates a response. So what makes this memOS fanciness different from just RAG? And it's a fair question to be asking, and the distinction is a real one. Namely, that RAG is retrieval that's done to the model. So usually there's this external system that sits around the model. It can be a system that's got some AI components to it itself, don't get me wrong. but the idea is that there's this external system that goes out identifies and retrieves the relevant information puts it into the context window and then that entire package is sent to the the core llm the core model in this system so the model is kind of sitting there on the retrieving end of all of that upstream action of assembly and retrieval the model is kind of this passive thing that just takes that all in and then spits out the best answer that it can.

7:56It's not doing things like making decisions on its own, like, hey, I don't have the information that I need, better go back to the knowledge system and try again. I can't, in this case, as the model, write back to the retrieval store. I'm not making the decision about what to retrieve in the first place. So those are things that really make MEMOS really different. So those are things that make MEMGPT really different. This is retrieval that's being done by the model. The agent is deciding when something needs to be stored externally. It's deciding what to retrieve and when that needs to happen, how to compress information before moving it out of context, and so on and so forth.

8:35And it's doing this because in this formulation, the agent is aware of its own memory architecture. So it knows it has a limited context window and an external store, and it's managing the boundary between them as part of doing its work. If you'd like an analogy, I always find these a little bit more helpful if you have like a mental model or a picture of something. RAG is kind of like a researcher that at the beginning of their research task, they're handed a stack of relevant papers that has everything that they need in there, at least should, for them to do their job. MemGPT is a little bit more like a researcher who is deciding throughout their research project which notebooks to pull off the shelf, which to put away, and they can write to the notebooks in between.

9:15So in both cases, you have this underlying library, but there's this very different relationship. And then, of course, in practice, modern agent systems often have both of these components operating within them. So you might have something like a RAG-style retrieval system for bringing in background knowledge, and then MEMGPT-style active memory management for maintaining the task state across a long-running process. But they're solving different problems. All right, so let's go back to memory management in the context of the agent doing its work. I've mentioned a couple times, but it's worth unpacking now, the whole idea of compaction, of moving things out of context in an intelligent way.

9:58And this may be familiar to you in particular if you've used something like Claude Code or another coding agent for a really long session. At a certain point, you're going to get a message that pops up about compaction being on the horizon and then actually happen. So here's what's going on with that. The point at which that triggers is the point at which you've had a bunch of stuff that's gone into the context window. So it can be a long session with lots of back and forth. It could be that there's a particularly large read of the code base that's happened, maybe all of the above. This starts to eat up the context window and at some point you're approaching the capacity.

10:37Now at this point what we want is not for the agent to crash out and say like, I'm all done, you have to start over. Compaction is what comes in to give you the illusion of the infinite continuation of the conversation. So what it's doing in the background is the agent will summarize everything that came before. They compressed the history of the session into a structured summary. And then in the context, it replaced the actual history with that summary and resumed from there. So that all that old context is gone and the summary is what's remaining. That's what the agent has to work with. So if you're following along closely and you're wondering at this point, oh, so it's like the agent wrote itself a note.

11:20Exactly. In this note, it wrote what it was doing, what decisions it made, what the current state of the code is, what still needs to happen. And then it threw away the original and kept only the note. This is genuinely useful. It's what allows the agent to work on a complex task across hundreds of turns without hitting a hard wall. But I will say it's a bit of a sleight of hand to call it an infinite context window. Sometimes tools do this, I think, kind of as a marketing gimmick. What it really is, is this rolling window of summarized history. And as summaries are prone to do, the summaries are lossy.

11:58So if you've encountered compaction with Claude Code, you may have encountered this as well, a failure mode. It's sometimes called context rot. What it feels like is you start your coding session with this sharp, clear context, tight reasoning from the LLM, bang, bang, bang, you're getting stuff done. And then as the conversation progresses, things start to get fuzzy or muddled where decisions that happened earlier in the conversation might have gotten forgotten. The LLM is asking questions that you already answered. It might make suggestions that contradict something that it said an hour ago.

12:36And so what could be happening there, in addition to potentially lost in the middle, is compaction creep, where you have these summarizations that are the breadcrumb trail of what's happened in that conversation, but they're losing detail. And especially after you've done this multiple times where you have summaries of summaries of summaries, those losses can accumulate. And to dive into the mechanics just a little bit more, because I think they are worth understanding, you can illuminate what's actually being preserved and what isn't. And when you start to exceed that threshold, there's a summary prompt that gets injected as a user turns.

13:11So in the same place where you as a user might ask your next question. And what this prompt is doing is it's asking the model to generate the structured summary and then the conversation history is cleared. And then it resumes with only the summary. If you go to the Cloud Code API documentation, you can understand a little bit more what's going on in their compaction prompt. Some of the stuff that it asks the model to preserve include things like the state, next step, learnings. So essentially what it's making is this little project status update that's written by the agent to its future self.

13:44The stuff that tends to survive are the high-level decisions, the current task state, what's been completed, and what's next. And the things that get lost are specific reasoning behind a decision that might have been made a while back, exact exceptions you might agree on, some subtle constraint you mentioned in passing, those details. After that summarization has been generated, there's a systematic reconstruction of the context that happens. So this will include the summary, but it also includes files that you've recently read with some limits on that. It'll reload any skills and tool definitions and instructions from the claude.md file.

14:24Claude.md is the name of a file. It's a bit of a technical detail. And it's pretty important if you want to be a power user of Claude code. It's a file in your project where you store things that must survive compaction. So this includes things like project conventions, architectural decisions, things the agent should never forget. And this is going to get re-injected at every compaction. So it's effectively pinned context that the summary process can't wash away. And so connecting this back to MEMGPT, where we started that operating system of the agent, the insight from MEMGPT was that the agent is who should be managing its own memory.

15:03And so compaction is this more mechanical version of this same idea, where a threshold triggers, the model summarizes, the history is replaced. The agent isn't necessarily making decisions that are quite as nuanced as they were thinking in MEMGPT regarding what it should be keeping. It's a little bit blunter than that. So MEMGPT is this very elegant architecture, but what we're actually using in practice is something much more direct, which is compaction. There's one other thing that connects nicely back to the last episode, if you caught it, where we were talking about lost in the middle. Compaction systems typically preserve recent context more carefully than older context.

15:45So the last few turns in the conversation stay more intact. They're represented with more detail than the conversation you were having two hours ago. And the thing that's fun is that mirrors the model's own recency bias that we talked about with Lost in the Middle. The architecture is already weighting recent tokens more heavily, and the compaction systems have learned to align with that rather than fight it. So the engineered solution and the architectural bias are pointing in the same direction. Now we've built up to another important turn in the description of memory management, which is that there's a hierarchy of context.

16:25Let's go back to what we mentioned a second ago about this claw.md file and how that's special. It gets re-injected every time. That's a hint of a broader pattern, which is that there's a hierarchy of context and it's not accidental. It's constructed very carefully. So when you're interacting with a language model agent, not all of the context in the window has equal standing. This is a little bit different from lost in the middle. At the top of the hierarchy is something that's called the system prompt. And these are instructions that are set by whoever built or configured the agent. This specifies its role, its constraints, what it should and shouldn't do.

17:01In a consumer product like Claude, that's Anthropic that writes the system prompt. In an enterprise deployment, it's the company building on top of the API in Cloud code. It's the core instructions that define how the coding agent behaves. The system prompt is always first in the context window, and it's always present. It's the thing that the agent is supposed to take most seriously. And then below that is where we're going to get something like CloudMD. That's where you'll store a layer of project-specific instructions that an operator or a user has has designated to be durable. This is where something like your team's coding conventions would live, or the architectural decisions that the agent should never lose track of, or constraints that should survive compaction.

17:45So this is privileged over the flowing back and forth conversation history because it's being re-injected deliberately. This isn't accumulated organically. And then below that is the conversation itself. So that's the back and forth the user turns between user prompts something, the agent responds, tool calls, results, things like this. So this is obviously the layer that's changing the most, it's the most dynamic. This is the one that grows with the task. And this is the layer that's the most vulnerable to the problems we've been discussing around lost in the middle, compaction loss, and context draw.

18:20So what we have here in this hierarchy is partly just management of the position bias from the last episode. Again, remember that the LLM will implicitly privilege the stuff that's early and late in the context window, and the stuff in the middle gets lost. So what we're seeing here is practitioners and system designers who have been aware of this for years are actually working around this architectural finding even before it was fully proven. You put important stuff like system prompts at the beginning of the context. That's where the model is reliably going to be paying attention. Your Claude.md gets re-injected every compaction, so it's always near the start of the fresh context.

19:04You're putting the things that matter in positions where the architecture is most likely to look. And this isn't true just for Claude code. It's how the instruction hierarchy works across LLM systems generally. And the technical terminology can vary from system to system, but the general structure is the same where you have the most privileged layer is a configuration that sits above the user interaction and it's designed to be more durable and more attended to by the LLM. And then the conversation flows beneath that. So if you're building an agent, understanding the hierarchy is one of the more practically useful things that you can know.

19:42The context window isn't this flat space where everything is on equal footing. It has this architecture to it. And if you're working with that architecture, you're more likely to build a system that stays coherent. So we've covered a lot here. Let's zoom out and give a quick compaction before we wrap up for this week. We've talked about RAG. We've talked about the different levels of memory management that are happening in a system like MemGPT. We've talked about compaction. These are all responses to the same underlying problem, which is that the context window can't hold everything. But I think it's worth a moment to think about what each of these systems are doing that's different in terms of the kinds of information, because that is a distinction that matters.

20:29RAG is mostly about what you might call semantic memory, which is factual knowledge about the world. RAG is going to help you retrieve relevant background information before the model reasons about a problem. MEMGPT is about episodic memory. So what's happening during a task, what the agent has observed, what is decided. And compaction is a hybrid. It's trying to compress this episodic memory into a semantic summary that the agent can work from. There's one more kind that's worth briefly naming because you may have already encountered it, which is procedural memory or how to do things. When you have an agent that has a skill, like OpenClaw installing a Gmail skill or a GitHub skill, what that's doing is externalizing procedural knowledge, knowledge about how to do things.

21:15The agent doesn't have to hold in context all the know-how about navigating Gmail's API. It loads that skill when it needs it, brings that information into context. And then when it's done, it can discard it. So skills are just kind of a packaging and like a conceptual layer around getting the right information into that constrained context rather than trying to fight it. It's just focused on a different category of information, namely how to do stuff. So where does this leave us now after four episodes? So the context window, now we've talked a lot about that. That is a constraint that's not going away.

21:52Context windows are getting larger. We're at millions of tokens for some models now, but there's just structural dynamics that make the middle hard to use, and those haven't fundamentally changed. What has changed is that the field has gotten much more deliberate about context as a resource to manage. People started to talk about context engineering as a discipline, and I think that this is why. You have this insight from memgpt, which is that you have this active agent-controlled memory management rather than just a bigger context window, and that's increasingly influential and important in how production systems are built.

22:27And then you have pieces like compaction, as unglamorous as it is, that's doing real work in every long coding session that doesn't crash. And so now that we've taken some time to go through what it's doing and where it fails, namely context rot and lossiness in the summary using claw.md as a workaround, that's all gonna be really useful knowledge for any of you who are out there who are working with these systems day to day. So with that, a teaser for next week, we're gonna talk about planning. If context and memory are about what an agent can hold in its head, planning is about how far ahead it can think.

23:03And the gap between deliberate search-based planning and what most production agents actually do in practice is larger and more interesting than you might expect. So we'll catch you next week for that. Thank you for tuning in this week. As a quick reminder, if you haven't yet, head over to Substack, subscribe to the newsletter. We have summaries and links for the content that we cover here every week, as well as little bits and bytes of content that I think is fun because there's a lot of good stuff out there. but that hasn't made it into any of the main content. So you'll get a few more things than just from the podcast alone.

23:40So head on over to Substack, look for Linear Digressions. If you haven't already, please subscribe on iTunes or Spotify or wherever you're getting this. Thanks, we'll talk to you next week. And I plan to talk to you about planning.

23:59This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

24:35Thank you.

From the publisher

Context windows are powerful — but finite, and surprisingly easy to overwhelm. When an AI agent is tackling a long, complex task, the information it needs has to fit inside that limited real estate, and research shows that anything buried in the middle tends to quietly disappear. So how do you design a system that actually *remembers* what matters? This episode digs into memory management for AI agents, from foundational computing concepts to practical lessons from tools like Claude Code.

---
Website: https://lineardigressions.com
Apple Podcasts: https://podcasts.apple.com/us/podcast/linear-digressions/id941219323
Spotify: https://open.spotify.com/show/1JdkD0ZoZ52KjwdR0b1WoT
Substack: https://substack.com/@lineardigressions

More from Linear Digressions

All 35 episodes
Memory Management for AI Agents (The Agents Season, Episode 4)Linear Digressions · 25 min
Listen in VO