In short
Memento proposes “fine-tuning LLM agents without fine-tuning LLMs” by enabling continuous, real-time learning for autonomous LLM agents using an external memory (case bank) rather than retraining the base model.
Guest backgrounds
No guests are named; the episode is a host-led discussion of the Memento paper.
Key claims
Fine-tuning is too costly and risks catastrophic forgetting; Memento achieves low-cost continual adaptation by storing trajectories (state/action/reward outcomes) and using case-based reasoning to plan and act. It adds an online-learned Q function to select which retrieved cases are most likely to succeed, without updating the main LLM.
Notable examples
Planner uses GPT-4.1 as a CBR agent to retrieve K similar past cases (often K=4) and generate plans; executor uses MCP tools (web search via Cerexing, crawling via Crawl4AI, multimodal processing, and a sandboxed code tool). Benchmarks: GAIA (87.88% validation success), Deep Researcher (~67 F1; ~2x over static chain-of-thought), HLE (2nd overall; 95% accuracy). Ablations show tools, explicit planning, and CBR each improve results.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding LLM Agents and Their Limitations
2:02 to 4:19
Discover the distinction between LLMs and agents, and the flaws in current approaches.
“Let's get clear on LLM agents and what's holding them back currently.”
How Memento Mimics Human Memory
4:19 to 5:53
Explore how Memento uses memory-based learning to improve agent adaptability.
“This is where Memento gets really interesting, taking inspiration from us, from humans.”
Memento Architecture: Planner-Executor Setup
5:53 to 7:50
Learn about the Memento architecture and its planner-executor cycle.
“And they formalize this with something called an MMDP, Memory Augmented Markov Decision Process.”
Memory Retrieval Methods in Memento
7:50 to 10:15
Understand how Memento retrieves memory through non-parametric and parametric methods.
“The executor, maybe a different general purpose LLM, takes each subtask from the plan.”
The Role of Tools in Memento's Functionality
10:15 to 12:50
Discover the various tools Memento utilizes for research and data analysis.
“It's trained separately, online, typically using a simple classification method to predict success failure.”
Evaluating Memento's Performance on Benchmarks
12:50 to 14:00
Examine Memento's performance results and its effectiveness across various tests.
“Segment five, let's look at Memento's results on those benchmarks.”
Exploring the Memento Paper's Findings
14:00 to 18:02
Learn about the insights and implications of the Memento paper on LLM agents.
“It shows static knowledge just can't keep up for real-time research.”
Paradigm Shift in AI Learning
18:02 to 18:15
Discover how the Memento paper could change the future of AI learning.
“The practical implications are pretty big.”
Reflections on Memory-Centric AI
18:15 to 18:59
Consider the implications of a memory-centric approach for future AI systems.
“Much closer to the idea of a generalist AI that learns as it goes.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive, where we extract the most important nuggets of knowledge from a stack of sources so you can stay well-informed and even surprised. Today, we're plunging into a really fast-moving corner of AI. We're talking large language models, LLMs, and how they're becoming, well, autonomous agents. It's a huge leap, really quite remarkable. Absolutely. But as these agents get smarter, have you ever stopped to think about the hidden challenge? How do you make them continuously adaptable without, you know, spending a fortune? That's the million-dollar question, really. And that's the core problem we're tackling in this deep dive.
0:36We're looking at a pretty groundbreaking paper called Memento, Fine-Tuning LLM Agents Without Fine-Tuning LLMs. Yeah, it offers a fascinating way around this whole cost issue, giving agents real-time learning chops without those massive computational bottlenecks. Because right now, these LLM agents, they're powerful, sure, but they often hit a kind of wall, don't they? They do. They either stick to really rigid preset instructions, which isn't great if the world changes. Which it always does. Exactly. Or the other route is fine-tuning, and that means retraining the entire huge LLM. Which sounds expensive.
1:12Incredibly expensive, both in compute power and data. It's just not practical if you want an agent learning all the time, adapting to new stuff. Right. So Memento isn't just like a small tweak. It sounds like a different way of thinking about it. It really is. It lets these agents learn continuously right there in real time from what they experience. That is crucial. It's crucial. But, and this is the key bit, without that massive cost of fine tuning the main LLM, it potentially changes the game for scaling these agents. Okay. So, our mission today is to unpack how Memento pulls this off. How does it get this cheap, continuous adaptation?
1:49We'll dig into how it borrows ideas from human memory, how the architecture works, the tools it uses, which are pretty cool, and yeah, how it actually performs on some tough AI tests. Sounds good. Let's dive in. Let's do it. Okay, segment one. Let's get clear on LLM agents and what's holding them back currently. We hear about LLMs answering prompts, but agents, that's something else. What exactly are they? Yeah, good distinction. An LLM agent isn't just waiting for your prompt. Think of it more like an autonomous system. It uses LLMs, maybe one, maybe more, to actually do complex tasks. So proactive.
2:23Exactly. Proactive. It interacts with its environment. It reasons. It makes decisions. It can use external tools. It even has a form of memory. It's working towards specific goals. The paper mentions things like deep research agents, right? Systems that go out and find information. Yeah, or agents that use tools to execute complex plans, maybe even code generation agents that write and fix their own software. They're built for active problem solving. Okay. You mentioned limitations. What are the main ways people build these now, and what are the flaws? So there are basically two main camps. The first uses these rigid, handcrafted frameworks.
3:00Kind of like building with Lego instructions, fixed steps. Pretty much. You build a fixed workflow, hard code the reasoning. For really narrow, predictable tasks, they work okay. But the downside? Zero flexibility. Once it's built, it's static. It can't learn from new info online, can't adapt if something unexpected happens. The world changes. The agent doesn't. Doesn't sound very intelligent in the long run. What's the second approach? The second approach tries to fix the flexibility issue by actually fine-tuning the LLM itself. So you update the core model's internal knowledge. That sounds better.
3:33More adaptable. Potentially, yes. More flexible behavior is possible. But the cost is just enormous. Considerable computational cost, as the paper puts it. Meaning... Massive compute, huge data sets needed, especially if you want it to keep learning. And then there's this risk called catastrophic forgetting. Forgetting, like it learns something new and forgets old stuff. Exactly that. It overwrites previous knowledge. So for open-ended situations where you need continuous learning, fine-tuning just isn't practical. It's too expensive, too slow, and too risky. Okay, so that really crystallizes the challenge Memento is tackling.
4:07How do we get agents that learn continuously from a changing world, but without that insane cost of constantly retraining the core AI? That's the core question, yeah. That's the big hurdle they're trying to clear. All right, segment two. This is where Memento gets really interesting, taking inspiration from us, from humans. Memory-based learning. Right. This is what sets it apart. Humans learn constantly, but we don't, you know, rewrite our entire brain every time we learn something new. Hopefully not. Memento tries to mimic that. The paper points to four human memory ideas, storing specific events as episodic traces.
4:44Like remembering that one time you touched a hot stove. Exactly. Then distilling those into abstract rules, maybe during sleep. Then there's reinforcement learning from good or bad outcomes, kind of like a dopamine hit. And crucially, using case-based or analogy reasoning, solving new problems by thinking, hmm, this is like that other time when. That makes a lot of sense. So Memento's idea is don't mess with the LLM's core knowledge, its parametric memory. Use an external memory instead. Precisely. Instead of trying to change the LLM's fixed internal weights, Memento relies on an external memory bank.
5:20It stores past experiences there, what it saw, did it work, did it fail. These are called trajectories. Like a logbook of experiences. Kind of, yeah. So when it faces a new problem, it queries this logbook for similar past situations. That's the case-based reasoning or CBR. It guides the LLM's decision-making based on past successes and failures. Ah, so it learns from its history. Directly. They call it low-cost continual adaptation via memory-based online reinforcement learning. It's a non-parametric, learn-on-the-fly approach. The learning happens in the memory, not by retraining the LLM. Learn on the fly.
5:56I like that. It feels much more dynamic. And they formalize this with something called an MMDP, Memory Augmented Markov Decision Process. Yeah, don't worry too much about the name. Basically, take a standard AI decision model state action reward. MMDP just formally adds a memory space to it. So every experience, state, action, reward gets stored, making memory a core part of future decisions. Got it. So looking at the system, Memento seems to have three main parts working together. That's right. You've got number one, a planner. It handles the strategy, breaking big tasks down. Number two, a tool enabled executor.
6:31This is the part that actually does things using tools, interacting with the world. And number three, the star of the show. The star, exactly. A growing case bank. This is that external memory we talked about, the place where all the experiences, good and bad, get stored. Okay, let's trace how that works in segment three. The memento architecture in action. They call it a planner-executor setup, a plan-and-act loop. How does that cycle work? It's basically a two-stage loop that repeats. Stage one is case-based planning. The planner, which uses a GPT 4.1 LLM acting as a CBR agent, gets a task. First thing it does, it queries the case memory.
7:09It asks, got any relevant past experiences for a task like this? It pulls out relevant cases, task, plan, success failure. So it's literally looking back at its history. Yeah. Yep. Those retrieved cases, plus the current task description, get fed into the planner LLM as part of the prompt. This helps it generate a better plan, broken down into subtasks. Like getting advice from its past self. Pretty much. Then there's a subtask memory that tracks these subtasks and their outcomes. The planner keeps an eye on progress, replans if needed, and importantly, when the task is done, it writes the new experience back into the case memory.
7:45Closing the loop so it learns from this new attempt. Whether it succeeded or failed, yeah. And then stage two is tool-based execution. Where the executor takes over. Right. The executor, maybe a different general purpose LLM, takes each subtask from the plan. It checks the tool memory, which logs tool use per subtest to figure out which tools it needs. Then it uses those tools via something called the Model Context Protocol, or MCP. MCP. What's that? We'll get more into tools next, but think of MCP as a universal adapter. It lets the executor talk to lots of different tools using a standard language.
8:18Okay, so planner figures out what to do based on memory. Executive figures out how to do it using tools. And it all hinges on that case memory. How does that memory actually get written to and read from? Good question. The write operation is pretty straightforward. After each action or step, the agent records the situation, state, what it did, action, and the result, reward, into the case bank. The state gets encoded. Actions and rewards are stored as is. It just keeps growing. Capturing everything wins and losses. Absolutely. Learning from mistakes is key. Now, for reading from memory, retrieving those past cases, there are two main ways Memento does it.
8:57Non-parametric and parametric. Okay, break those down. Non-parametric. Read MP, or non-parametric retrieval, is the simpler one. It's like basic analogical reasoning. You give it the current task or situation. It calculates how semantically similar the situation is to all the past situations stored in its memory. Comparing the current problem to everything it's seen before. Right, and it pulls out the K nearest or most similar past cases. The study found K4 often worked best. It assumes similar problems might have similar solutions. Makes sense. And the parametric one, read P, sounds fancier. It is a bit more sophisticated.
9:32Read P or parametric retrieval. Here, when it writes an experience to memory, it also updates a little helper model called a Q function. A Q function. Yeah. Think of it as a predictor. It learns over time to estimate the utility or the likelihood of success if the agent uses a particular past case to guide its action in the current state. Ah, so it's not just finding similar past cases, but predicting which of those similar cases are likely to lead to a good outcome this time. Exactly. It performs adaptive case selection. When it reads from memory, it uses this learned Q function to pick the K cases with the highest predicted success rate.
10:10And crucially, training this Q function doesn't involve retraining the main LLM. Correct. That's the beauty. It's trained separately, online, typically using a simple classification method to predict success failure. It adds smarts to the memory retrieval without the massive cost of LLM fine-tuning. That's clever. Adding a layer of learned judgment about which memories are most valuable. Okay, let's shift gears to segment four. Tools. Memento isn't just thinking, it's doing, especially for research tasks. Tools are critical. Absolute critical. You can't do deep research just by talking. You need to find info, process different kinds of data, analyze it.
10:48And you mentioned that MCP, the model context protocol. Right. MCP is key here. It's designed as a standardized model agnostic interface. Basically, a universal way for the agent to connect and talk to all sorts of different external tools and data sources. No need for custom integrations for every single tool. Like a universal remote for AI tools. So what's in the toolbox for getting information? For external information acquisition, it's well equipped. It uses Cerexing, which is a meta search engine. It queries Google, Bing, DuckDuckGo, Brave, pulls results together, and re-ranks them. So broad coverage.
11:23Yeah. And if it needs more than just snippets, it has Crawl4AI. That's a web crawler that can go fetch and parse the full content of web pages. Search finds leads. The crawler digs deeper. Makes sense. But research isn't just text. What about images, audio, spreadsheets? That's where the multimodal heterogeneous information processing toolkit comes in. It's designed precisely for that variety. Like what? Okay, it can use a vision language model, VLM, to understand images, generate captions. It has automated speech recognition for audio. It can parse PowerPoint slides, even describing images within them, convert spreadsheets into something readable, unpack zip files, read code, JSON, XML, translate Word docs, even get summaries of videos using VLMs.
12:08Wow, that's a lot. It is, and it has fallbacks for things like PDFs. The goal is a unified interface to handle almost any kind of data you throw at it, just like a human researcher would piece together info from different sources. Incredible versatility. And then there's the analysis part. Reasoning. Right, the reasoning analysis toolkit. This includes a really important code tool. It's a sandboxed environment where the agent can write and run code shell scripts, Python, using libraries like NumPy, Pandas, PyTorch. So it can actually perform calculations, data and out. Yes, and crucially, it maintains state across steps.
12:41So we can do multi-step data analysis or automation tasks. And of course, there's a basic math tool for simple arithmetic. It's a pretty complete workbench for an AI researcher. Okay, the setup sounds powerful. The tools are impressive. But does it actually work? Segment five, let's look at Memento's results on those benchmarks. How do they test it? They really put it through its paces on four very different benchmarks to see if it's truly general purpose. There was GAIA, which tests long-term planning and tool use with different difficulty levels. Then Deep Researcher focused on real-time web research, finding evidence, connecting dots across multiple steps.
13:17Simple QA for just getting facts right, precision. And Humanity's Last Exam, or HLE, which is super challenging test reasoning on niche academic topics, really pushing the limits. A tough gauntlet. So how did Memento do? What are the highlights? The results were pretty stunning, honestly. On GAIA, it hit the number one spot on the validation set, 87.88 % success rate, and fourth on the main test set, beating some well-known frameworks. Very strong. On Deep Researcher, it got around 67 % F1 score and 80 % overall performance. Now get this. That nearly doubled the score of a standard approach, using chain-of-thought prompting with a static knowledge base.
13:57Doubled. Wow. That really shows the power of using live tools and learning from experience. Absolutely. It shows static knowledge just can't keep up for real-time research. Then on HLE, that really tough academic benchmark, it came in second overall, just barely behind a system labeled GPT-5 in the benchmark. That's amazing for this kind of approach, showing CBR helps even in these really specialized long-tail knowledge areas. Incredible. Competing at that level without fine-tuning the base LLM and simple QA. Factual stuff. Nailed it. 95.0 % accuracy. Set a new state-of-the-art. significantly cut down on hallucinations, which is a big deal for factual questions.
14:36Those are genuinely impressive numbers across the board. But digging deeper, did the paper explore why it worked so well, like those ablation studies? Yes, they did, and the findings are really insightful. First, they looked at the number of retrieved cases. K. You might think more examples are always better, right? Yeah, like fuchsia prompting. But here, no. Performance actually peaked with a small number, K4, for deep researcher. It suggests it's not about flooding the LLM with memories, but retrieving a few highly relevant ones. Quality over quantity. Interesting. What about taking components away?
15:05That showed a clear story. Just adding online tools, the online executor, instead of relying only on the LLM's internal knowledge helped reduce hallucinations and boosted performance. Adding explicit planning gave another solid jump across all benchmarks. Okay, so tools help. Planning helps. What about the core idea, CBR? That was the clincher. adding case-based reasoning on top consistently provided extra additive improvements. We're talking another 4.5 % to 7 % gain on HLE and 6.7 % to 8.2 % on Deep Researcher. It wasn't just one thing, it was the combination, but CBR was clearly a vital ingredient.
15:42So memory really matters. It's not just a gimmick. Did they show it learning over time? They did. The continual learning curve showed the full Memento system steadily improving its accuracy as it ran more tasks and accumulated more experiences in its case bank, outperforming versions without CBR or planning. And importantly, it showed good generalization. Meaning it could handle stuff it hadn't seen before. Exactly. They tested it on completely different data tests. Music, Bamboogle, PopQA, and Memento still showed significant improvements, like 4.7 % to 9.6 % absolute gains. This suggests CBR helps it learn skills that actually transfer to new, unseen problems.
16:21That's huge. That's what you want from intelligence. Fine, right? Adaptability. Any other cool operational details how it used tools or costs? Yeah, a couple of things. The tool usage breakdown showed that naturally code search and web crawling were used most. But as tasks got harder, the agent didn't just call tools more often. It spent more effort interpreting and integrating the information from the tools. More thinking, not just more doing. A sign of deeper reasoning. What about the computational side, token costs? Token costs for the input went up a lot with harder tasks. like way up level 3 GIA tasks needed over 120 ,000 input tokens.
16:57That's mostly from feeding the LLM all the detailed info coming back from the tools across multiple steps. So processing the evidence is the main cost driver for complex tasks. Seems like it. The output tokens, the actual answers, stayed pretty stable and efficient. And one last interesting tidbit, fast versus slow think mode. What's that? They compared planners. Fast, concise planners like GPT 4.1 worked much better than slower, more verbose ones. Too much deliberation, too much text in the planning stage, actually seemed to confuse things and hurt performance. So even for AI, sometimes overthinking is bad.
17:33Get to the point. Pretty much. Concise, structured planning was key. Okay, let's pull this all together. This Memento paper feels like a really significant step. So boiling it down, Memento's core achievement is figuring out how to let LLM agents learn continuously, adapt in real time using this external human-like memory, the case bank, and a powerful set of tools. All without that incredibly expensive, slow process of fine-tuning the giant base LLM every time it needs to learn something new. It's a paradigm shift, right? It could be. The practical implications are pretty big. It points towards building AI agents that are much more scalable and efficient.
18:08Agents that can genuinely acquire new skills over time, handle complex research, and adapt to a changing world. Much closer to the idea of a generalist AI that learns as it goes. Exactly. It's a compelling path in our it's more adaptable and ultimately more capable AI systems. And thinking about this approach, mimicking how we learn from experience by reflecting on specific past events stored in memory, not just by constantly rewriting our fundamental understanding, it really makes you wonder, doesn't it? It does. What could this memory-centric approach mean for how we design future intelligent systems?
18:44How we interact with them? Maybe it even gives us a new lens for thinking about our own continuous learning process, the way our experiences shape us over time. It feels like we might be seeing a blueprint for AI that learns a bit more like we do. A fascinating thought to end on. Lots to consider there.
From the publisher
The research introduces Memento, a novel approach for adaptive Large Language Model (LLM) agents that enables continuous learning without requiring fine-tuning of the base LLM parameters. This method leverages a memory-based online reinforcement learning framework, formally defined as a Memory-augmented Markov Decision Process (M-MDP), which stores past experiences in an episodic memory and continually updates a neural case-selection policy. Memento utilizes a planner-executor architecture and a comprehensive suite of tools, demonstrating state-of-the-art performance on various benchmarks, including GAIA, DeepResearcher, and SimpleQA. The ablation studies confirm that both parametric and non-parametric case-based reasoning (CBR) are crucial for significant performance gains and effective generalization to out-of-distribution tasks.




