In short
Why LLMs “forget” and how context engineering builds persistent, personalized agents using sessions (short-term conversation state) and memory (long-term user-specific facts).
Guest backgrounds
No guest names or bios appear in the transcript; it’s a two-speaker discussion (host/interviewer plus expert).
Key claims
LLMs are inherently stateless and only see a single API call’s context window; sessions are a chronological event log plus a scratch-pad state; memory differs from RAG (RAG retrieves static documents; memory extracts dynamic user preferences from dialogue).
Notable examples
peanut allergy changed over time triggers conflict resolution (update/create/delete); memory retrieval scores blend relevance, recency, and importance; memory placement risks “over-influence” (system instructions) or “dialogue injection” confusion (history injection); memory generation runs asynchronously via an LLM-driven ETL pipeline with provenance, PII redaction, and memory-poisoning defenses.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Sessions in AI
0:45 to 3:30
Discussing the role of sessions as a temporary workspace for continuous conversation.
“Into a truly persistent, personalized agent.”
Dynamic Strategies to Manage Context
3:31 to 7:28
Examining methods for managing session history and addressing context rot.
“We call those the events, so the turn-by-turn dialogue.”
The Difference Between Memory and RAG
7:29 to 11:21
Clarifying how memory management differs from retrieval augmented generation.
“So once you've extracted the signal, the next step you mention is consolidation.”
Memory Generation and Provenance
11:22 to 13:32
Exploring how memories are generated, managed, and validated for reliability.
“This is what defines the trust hierarchy.”
Key Takeaway on Context Engineering
13:33 to 13:52
Highlighting the importance of context engineering for building robust AI systems.
“So it's an alternative to slow offline fine-tuning.”
Transcript
Automatic transcript. May contain errors.0:00Antonio Gulli:So why is it that you can spend, I don't know, 20 minutes carefully explaining your entire life story, your goals, the context of a really difficult task to a sophisticated AI, only for that same assistant to completely reset and forget everything like three questions later? It's so frustrating.
0:17Kimberly Milam:It is. And that frustrating amnesia, it really stems from a fundamental reality. Large language models are inherently stateless.
0:26Antonio Gulli:Stateless, meaning they just, they don't have a built-in memory.
0:29Kimberly Milam:Exactly. They operate only within the context window of a single API call. So to bridge that gap, to build an AI that actually remembers who you are and evolves with you, that requires this advanced discipline we call context engineering.
0:44Antonio Gulli:Context engineering. Okay, that sounds pretty technical, but it feels like it's the core practice that shifts an AI from being, I guess, a powerful but temporary calculator.
0:52Kimberly Milam:Yeah, that's a good way to put it.
0:54Antonio Gulli:Into a truly persistent, personalized agent. Absolutely.
0:57Kimberly Milam:Our mission in this deep dive is for you to really get how context engineering works. It's the dynamic assembly and management of all the necessary information, the instructions, the history, the retrieved data, all within that very limited context window.
1:14Antonio Gulli:And we're going to unpack the two essential components that make this whole thing possible, right? Sessions and memory.
1:19Kimberly Milam:The two symbiotic components, yeah.
1:20Antonio Gulli:Okay, so to frame this for everyone, let's use a simple analogy. Think of the session as your messy, temporary workbench. It's got all the tools, scribbled notes, and just the immediate mess of a single conversation.
1:31Kimberly Milam:And then the memory is the meticulously organized, climate-controlled filing cabinet.
1:37Antonio Gulli:Ah!
1:38Kimberly Milam:It holds only the most critical, processed, and finalized knowledge. It's data that's stored for the long term, accessible across all future conversations, not just the one you're having right now.
1:48Antonio Gulli:Got it. Okay, so let's untack context engineering first. Most of us are familiar with prompt engineering, which is, you know, about crafting the perfect set of instructions. How is this different?
1:57Kimberly Milam:Well, prompt engineering is static. It's all about crafting the perfect recipe for one specific instruction. Context engineering is dynamic. It addresses the entire payload.
2:09Antonio Gulli:The entire payload.
2:11Kimberly Milam:Yeah, it's like the entire mise-en-place in a kitchen. You are orchestrating the delivery of not just the instruction, but all the ingredients, all the historical context, the procedural knowledge, everything the model needs.
2:22Antonio Gulli:So you're making sure it has what? No more and no less than the most relevant information.
2:26Kimberly Milam:That's the goal.
2:27Antonio Gulli:So this is really about coordination then. You're not just writing a better prompt. You're coordinating all these external systems, RAG databases, session stores, memory managers, to construct a fully state-aware payload for every single turn.
2:40Kimberly Milam:Precisely. And this orchestration is absolutely crucial because it directly addresses the biggest operational challenge we see, which is context rot.
2:48Antonio Gulli:Context rot. I like that term.
2:49Kimberly Milam:As that context window grows, your cost and latency just skyrockets. Yeah. And what's worse, the LLM's ability to pay attention to the critical information in that sea of text, it just diminishes.
3:02Antonio Gulli:It gets distracted.
3:03Kimberly Milam:It gets distracted. Context engineering combats this by implementing dynamic strategies like compaction and summarization to prune and curate that history on the fly.
3:13Antonio Gulli:Okay, so that problem of context rot leads us straight to the first component, sessions, the workbench. If this history is on the hot path, meaning it has to be sent with every single turn, how do we define what a session actually contains?
3:28Kimberly Milam:A session is really just a container for that single continuous conversation. It holds the chronological history. We call those the events, so the turn-by-turn dialogue.
3:37Antonio Gulli:User says this, agent says that.
3:39Kimberly Milam:Exactly. And it also holds the state, or what some people call a scratch pad. That's the temporary structured working memory. Think of it as the items currently sitting on the counter right next to your cutting board.
3:49Antonio Gulli:Like the items you've added to a shopping cart during a conversation.
3:53Kimberly Milam:Perfect example. This temporary state for the task at hand.
3:56Antonio Gulli:But because of that performance constraint, that session can't just grow indefinitely. So this compaction imperative is huge. Are there advanced ways to do this? To use the AI's own power to summarize and manage the context instead of just, you know, crudely cutting it off?
4:12Kimberly Milam:Yes, and the strategies really range from simple to complex.
4:16Antonio Gulli:Yeah.
4:16Kimberly Milam:The simplest is just a sliding window. Keep the last end turns.
4:20Antonio Gulli:Okay, easier.
4:21Kimberly Milam:A bit more token efficient is token-based truncation. It just counts backwards from the most recent query and cuts off older messages once you hit your token limit.
4:29Antonio Gulli:And the really sophisticated approach, the one that uses the LLM itself.
4:33Kimberly Milam:That would be recursive summarization. This uses the LLM to condense the older parts of the conversation into a nice tight summary. That summary is then prefixed to the current verbatim messages.
Read the full transcript
4:44Antonio Gulli:That sounds expensive, computationally.
4:46Kimberly Milam:It is. And crucially, you never want the user waiting around while the AI summarizes the last hour of conversation. So it must be performed asynchronously. It happens in the background, gets persisted, and then it's ready for the next turn. It keeps the history lean without, you know, losing the thread of the dialogue.
5:04Antonio Gulli:But the workbench isn't enough, is it? The session only captures the immediate chaos. us, we need a place to put the lasting lessons.
5:11Kimberly Milam:Right. And that's where memory comes in.
5:12Antonio Gulli:The filing cabinet. Memory is that snapshot of extracted, meaningful information that's persisted across multiple sessions to build that continuous, personalized experience.
5:22Kimberly Milam:And this is where we have to clarify, I think, the biggest point of confusion in AI architecture today. Memory versus RAG, retrieval augmented generation.
5:31Antonio Gulli:Oh, yeah. I hear RAG used for everything. How do these two concepts actually differ?
5:36Kimberly Milam:They're fundamentally different in purpose. RAG, as it's usually defined, is about accessing external static data like documents, FAQs, knowledge bases.
5:45Antonio Gulli:So RAG makes the agent an expert on facts. It retrieves fixed data and uses it as evidence.
5:51Kimberly Milam:We can think of RAG as the research librarian.
5:54Antonio Gulli:Okay, the research librarian. I like that.
5:56Kimberly Milam:Exactly. The memory manager, on the other hand, is the personal assistant.
6:00Antonio Gulli:Ah.
6:00Kimberly Milam:Its data is dynamic, it's user-specific, and it's derived from the conversation dialogue itself. It's not about accessing a PDF. It's about extracting and consolidating facts, like the user lives in Boston, or the user prefers to fly American Airlines.
6:16Antonio Gulli:Or the user's current project involves optimizing server latency, things you'd never find in a static document.
6:22Kimberly Milam:Precisely. Memory makes the agent an expert on the user.
6:26Antonio Gulli:And these little atomic pieces of memory, they unlock personalization, and critically, they help with that efficient context management we were just talking about. But building that filing cabinet, it's not a passive thing. You describe memory generation as an LLM-driven ETL pipeline. What does that mean?
6:42Kimberly Milam:ETL stands for Extract, Transform, and Load. It just means we're actively automating the process of turning that messy conversational noise into structured, usable knowledge.
6:52Antonio Gulli:Okay, walk us through those stages. Let's start with extraction.
6:55Kimberly Milam:Extraction is basically targeted filtering. It's separating the signal, the facts, the preferences from all the conversational noise.
7:02Antonio Gulli:How does it know what the signal is?
7:04Kimberly Milam:Well, the LLM is guided by programmatic guardrails, and often it's using what we call topic definitions. These can be formal schemas or just natural language descriptions of what is meaningful for that particular agent's purpose.
7:19Antonio Gulli:And developers can kind of show it what to do.
7:21Kimberly Milam:Yeah, that's right. They often use C-shot prompting, where you show the LLM a few examples of input text and the ideal structured memory it should output.
7:29Antonio Gulli:So once you've extracted the signal, the next step you mention is consolidation. Why is self-editing so vital for a memory system?
7:37Kimberly Milam:Because if you just stuffed every single piece of paper you extracted from every session into the filing cabinet,
7:42Antonio Gulli:you'd have a useless mess.
7:43Kimberly Milam:You'd have a useless contradictory mess. Consolidation is what keeps that knowledge base clean and reliable. It handles duplication merging, the same fact mentioned multiple times. And it also handles forgetting.
7:55Antonio Gulli:Forgetting. You want it to forget things.
7:57Kimberly Milam:Proactively. You prune old or low-confidence memories through something called memory relevance decay. You don't want outdated information cluttering things up.
8:07Antonio Gulli:Okay, so what happens when the user's state changes? Let's say I told the agent last year I was allergic to peanuts, but this year I mentioned I had outgrown the allergy.
8:15Kimberly Milam:That's the most critical part. Conflict resolution. The system has to evaluate which piece of information is more current or more reliable. This then results in one of three operations. It'll update an existing memory, create a new one, or just delete the old one.
8:30Antonio Gulli:And this whole process is why the memory generation has to be so robust. It's literally determining what the agent believes is true about you.
8:38Kimberly Milam:It's the agent's source of truth about the user.
8:41Antonio Gulli:So once we have these curated memories in our filing cabinet, we need to find the right ones for the job and do it fast. We have to balance usefulness with a really strict latency budget. So what makes a memory the belt fit?
8:54Kimberly Milam:You can't just rely on, you know, basic semantic similarity. It's not enough. A good agent has to blend multiple scoring dimensions.
9:01Antonio Gulli:Okay, like what?
9:02Kimberly Milam:Well, you need to score on relevance, which is that semantic similarity to the query. But also recency. How recently was the memory created or updated? And finally, importance. Importance.
9:15Antonio Gulli:A score you assign when the memory is created.
9:18Kimberly Milam:Exactly. A significant score. Blending those three prevents the agent from, say, surfacing a highly relevant but totally trivial memory from two years ago over something important that happened last week.
9:29Antonio Gulli:Right. And once you've retrieved the best memories, where do you put them in the context window? Does the placement actually matter?
9:35Kimberly Milam:It matters immensely. It changes the way the LLM reasons. The two main options are highly strategic. You can append memories to the system instructions. This gives them very high authority. It's ideal for stable global information. like a user profile.
9:50Antonio Gulli:But there's a risk.
9:51Kimberly Milam:The risk is over influence. The agent tries to relate everything back to its core instructions, even when it's totally irrelevant.
9:57Antonio Gulli:And the other option, if we just inject them into the conversation history.
10:01Kimberly Milam:That's what we call dialogue injection. Usually you place it just before the latest user query. It's more subtle. But the big risk there is also called dialogue injection.
10:10Antonio Gulli:Wait, same name.
10:12Kimberly Milam:Yeah, it's a bit confusing. It means the model mistakes the memory, which, remember, it didn't actually say, for something that was actively said in the conversation.
10:20Antonio Gulli:Which could lead to some very confusing replies.
10:23Kimberly Milam:Very confusing. That's why a lot of advanced agents are moving towards a memory-as-a-tool approach. The agent itself decides when to call a specific tool to retrieve context. It makes the whole process more reactive and a lot less risky.
10:37Antonio Gulli:Moving this into production reality, we touched on this, but memory generation is computationally expensive. It involves LLM calls, database rights.
10:45Kimberly Milam:Which confirms why memory generation has to be an asynchronous background process. Right. If you try to generate and consolidate synchronously, you are blocking the user experience. Yeah. You're making them wait. Decoupling that whole pipeline from the user-facing runtime is absolutely essential for a fast, resilient system.
11:04Antonio Gulli:So for the agent to make reliable decisions, especially with that conflict resolution piece, it has to be able to evaluate the quality of its own memories. And this is where the concept of provenance comes in.
11:16Kimberly Milam:Provenance is essentially the memory's detailed receipt. It's a record of its origin and its history. This is what defines the trust hierarchy.
11:24Antonio Gulli:So not all memories are created equal.
11:26Kimberly Milam:Not at all. For instance, bootstrap data that's been preloaded from a trusted CRM is very high trust. But implicit user input, something the model just inferred from dialogue, that's generally much less trustworthy.
11:39Antonio Gulli:And the system uses this to assign a dynamic confidence score.
11:43Kimberly Milam:Yes, and that confidence isn't static. It increases with corroboration. So, when multiple trusted sources confirm the same piece of information.
11:51Antonio Gulli:And it decays over time.
11:52Kimberly Milam:It decays over time, or when contradictory evidence is introduced. This score isn't for you, the user, to see. It's injected into the context window so the LLM can properly weigh the evidence when it's reasoning.
12:03Antonio Gulli:Finally, we have to touch on security and privacy. This is a huge deal, especially since memory contains so much user-specific data.
12:11Kimberly Milam:It's paramount. The cardinal rule is strict data isolation at the user level, and that's enforced via access control lists or ACLs.
12:20Antonio Gulli:Just dictating who can read or write what data.
12:22Kimberly Milam:Simple as that.
12:23Antonio Gulli:Yeah.
12:23Kimberly Milam:But two other critical safeguards are needed before any session data is persisted. First, PII redaction, ensuring sensitive data is never stored long-term in the memory manager. And second, a robust memory poisoning defense.
12:36Antonio Gulli:Defense against what?
12:38Kimberly Milam:Against malicious users trying to corrupt the agent's persistent knowledge through sneaky, prompt injection attacks, disguised as just benign dialogue. The system has to validate and sanitize that incoming information.
12:50Antonio Gulli:Wow. Okay, that's a really comprehensive look at how statefulness is actually built. The key takeaway for me is this. Context engineering, which is really the symbiotic relationship between sessions managing the now and memory managing the long term. That is the difference between a brittle chatbot and a truly stateful personal AI.
13:08Kimberly Milam:Absolutely. And, you know, we focused mostly on what's called declarative memory. That's the knowing what facts user preferences. But the most complex and I think exciting frontier is procedural memory.
13:17Antonio Gulli:The knowing how.
13:18Kimberly Milam:The knowing how, yes. Procedural memory is the system that would allow an agent to distill a reusable strategy or a playbook from a successful interaction. This knowing how guides complex task execution, and it allows for fast online self-evolution of the agent's logic.
13:34Antonio Gulli:So it's an alternative to slow offline fine-tuning.
13:38Kimberly Milam:A much faster one. Which raises, I think, a really important question for you, the listener, to think about. If an agent is constantly curating and editing its own operational playbooks, how will we evaluate the trustworthiness of an agent when it designs its own steps for success?
From the publisher
This whitepaper by Google titled **"Context Engineering: Sessions & Memory,"** authored by Kimberly Milam and Antonio Gulli in November 2025, which provides a detailed guide to building stateful, intelligent Large Language Model (LLM) agents. The document defines **Context Engineering** as the process of dynamically managing information within an LLM's context window, emphasizing two core, interconnected components: **Sessions** and **Memory**. **Sessions** manage the immediate, chronological dialogue and working state of a single conversation, while **Memory** is a decoupled system for long-term persistence, capturing and consolidating key information across multiple sessions to enable personalization. The paper extensively covers architectural considerations for both sessions (e.g., compaction strategies for managing long context) and memory (e.g., types of memory, storage architectures, and the LLM-driven process of extraction and consolidation), contrasting the dynamic, user-specific role of memory managers with the static, factual role of Retrieval-Augmented Generation (RAG) engines. Finally, it outlines critical production requirements, including **privacy**, **security**, and **asynchronous processing**, to ensure robust and efficient deployment of these state-aware agents.




