In short
ReasoningBank, a Google Cloud AI research memory framework for LLM agents that prevents “forgetfulness” by storing structured, transferable reasoning strategies distilled from both successes and failures; combined with MATES (memory-aware test-time scaling) to improve effectiveness and efficiency via deeper exploration.
Guests
No specific guest names or backgrounds are provided in the transcript; it appears to be a host-led episode with researchers referenced generally.
Key claims
Prior agent memory is passive (raw trajectories) or success-only (rare “perfect recipes”), lacking transferable “why” and ignoring failure lessons. ReasoningBank stores memory items as title + context description + distilled strategy. An LLM judge scores outcomes using prompt-provided success criteria, enabling a closed loop without external labels. Failures become counterfactual signals that sharpen boundaries.
Notable examples
Customer order query—baseline returns a “recent orders” date; ReasoningBank recalls “data completeness” and clicks “view all,” finding March 2, 2022. Web Arena shopping subset—success rate rises from ~46.5% (success-only) to ~49.7% (including failures). Efficiency: shopping task drops from 29 to 10 steps; strategies evolve from procedural actions to compositional reasoning (e.g., search+filters+UI cross-checks).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Fatal Flaw of LLM Agents
0:45 to 1:41
Explaining the forgetfulness of LLM agents and its impact on complex tasks.
“So our mission today is to dive into some fascinating work coming out of Google Cloud AI research and partner institutions.”
Introducing Reasoning Bank
1:41 to 3:12
Overview of Google's Reasoning Bank and its innovative memory framework.
“And, you know, why wasn't that approach good enough?”
Old Memory Systems and Their Drawbacks
3:12 to 4:42
Examining the limitations of previous memory systems in agents.
“And the content is the distilled strategy itself.”
How Reasoning Bank Improves Learning
4:42 to 6:26
Describing how the structured memory items in Reasoning Bank enhance learning.
“It uses its reasoning capacity to justify its assessment against those criteria.”
The Role of Self-Critique in Learning
6:26 to 7:58
Discussing the innovative self-judging mechanism of LLM agents.
“Yeah, that 3.2 percentage point jump, simply from learning how to transform failure into a constructive signal, that's significant in these complex tasks.”
Failures as Learning Opportunities
7:58 to 8:10
Understanding the significance of learning from failures.
“So we've got this robust, structured knowledge base now in Reasoning Bank.”
Case Study: Customer Order Query
8:10 to 9:50
Explaining how Reasoning Bank improved accuracy in a specific example.
“And you mentioned this is about scaling through depth, not just throwing more tasks at it.”
Enhancing Learning Through MATES
9:50 to 11:01
Overview of memory-aware test time scaling and its impact on learning.
“The second was sequential scaling, which they termed self-refinement.”
Implementing Deep Exploration
11:01 to 12:34
Breaking down the methods used for implementing deep exploration in agents.
“Did this just mean the agent was getting stuck and giving up faster on hard problems?”
The Evolution of Memory Items
12:34 to 14:00
Describing how memory strategies evolve over time in agents.
“Does the agent actually get smarter in more complex ways over time?”
Show all 11 chapters
Scaling Agent Learning Through Structured Memory
14:00 to 15:13
Learn how structured memory enhances LLM agents' adaptability and reasoning.
“And then this powerful structured memory synergizes really effectively with Matt, that novel scaling approach that focuses on depth of exploration.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. So if you've spent any time marveling at the power of large language model agents, I mean, they can code, they can browse, they can handle these really complex multi-step plans. You know they are smart. Seriously smart. But they have this one kind of fatal flaw. They are fundamentally forgetful. It's incredibly frustrating, isn't it? You set them up on a complex, long-running task, maybe software development or some kind of persistent web admin. Right. And they just discard every valuable lesson they literally just learned. They tackle every new instance. As if it's the first time.
0:36Exactly. They're doomed to repeat the same mistakes again and again, simply because they lack any real form of self-evolving strategic memory. And that inability to learn, to really learn from experience, means they often fail when things get truly complicated or nuanced. So our mission today is to dive into some fascinating work coming out of Google Cloud AI research and partner institutions. they've developed a novel memory framework specifically designed to solve this problem. It's called Reasoning Bank. Yeah. Reasoning Bank is potentially a game changer because it moves way beyond just passive logging.
1:10It doesn't just keep a diary, you know. It actively distills those experiences, the good and the bad, into, well, a strategic playbook, a playbook of generalizable reasoning strategies. Interesting. We'll also explore this really critical synergy between this kind of high-quality memory in a technique they call memory-aware test time scaling, or MATES, that lets the agent really deepen its learning from experience. Okay, so let's maybe unpack the old way first. If agents had memory systems before Reasoning Bank, what were they actually storing? And, you know, why wasn't that approach good enough?
1:45Well, historically, their memory systems were essentially passive record keepers. Pretty basic, really. Some systems stored raw trajectories. Think of it like saving an entire browser session log every single click, every step. Oh, wow. That sounds messy. It gets noisy and just way too large very quickly. Then other systems, they tried storing only common successful routines, like a perfect recipe for a task that happened to go right that one time. So if I ran a thousand tasks, I'd ended up with either this, like mountain of unstructured text logs, or maybe just a little booklet of perfect but maybe rare recipes.
2:20Precisely. Yeah, exactly. The core drawback was abstraction or the lack of it. They couldn't really pull out the high level why. The underlying reason. Yes, the transferable principle that made something work. And maybe even worse, by often focusing only on success, they completely ignored all the valuable lessons hidden in their own failures. Which is where we learn a lot, right, in failure. Absolutely. So Reasoning Bank sounds like it fixes that by forcing some structure onto the experience. So what does a memory item actually look like in this new system? Good question. It's highly organized, which is a big shift from those messy text logs.
2:57Every memory item gets abstracted into three key parts, designed for clarity and, crucially, transferability. Okay, what are they? First, there's a concise title, just identifies the strategy. Then a brief description summarizing the context, the situation, and finally, the content. And the content is the distilled strategy itself. Exactly. The distilled reasoning steps or maybe some operational insights, that structure is what makes the memory truly actionable, not just a record. And how does this become, you know, a continuous learning loop? The agent gets a task, it retrieves relevant memories, it acts, and then what happens?
3:33Right. So after acting, the agent actually analyzes its own outcome. This is pretty neat. An LLM acts as an internal judge. Its own judge. Yes. And this is without needing external ground truth labels. Importantly, it decides if the experience was a success or a failure. Then that new experience gets analyzed, distilled down into one of these new structured memory item. The title description content thing. And then consolidated back into the reasoning bank. So it's this closed loop where the agent literally updates its own strategic knowledge after, well, pretty much every significant interaction.
4:07Wait, hang on a second. LLM judges itself. Isn't trusting an LLM to grade its own homework, especially failure, a bit like asking the fox to guard the hen house? How do they make sure that internal judgment is actually reliable? Yeah, that's a really crucial point. And honestly, it's the key innovation that unlocks learning from failure. The LLM as a judge system isn't just going, hmm, I feel good about this. It's specifically using the task success criteria that were provided in the original prompt. You know, did the agent actually achieve the stated goal? Did the code compile successfully? Things like that.
4:40Ah, so it has objective criteria to check against. Right. It uses its reasoning capacity to justify its assessment against those criteria. And the self-critique, I mean, even if it's not perfectly accurate 100 % of the time, it generates the necessary signal for the next stage, distilling those really useful counterfactual lessons from the failures. Okay, that self-critique mechanism, that leads us right into what feels like the most vital insight here. Actually, using the failures. You said in older systems, failure was basically just noise. It dragged performance down. Right. Why is extracting wisdom from mistakes suddenly so crucial for generalization now?
5:17Because, well, successes validate strategies you already think might work, but failures, failures define the boundary. Failures provide these essential counterfactual signals and pitfalls. Think about it. If you only teach an agent the successful paths through a maze, it doesn't know where the dead ends are until it hits them again. Right. It doesn't learn the don't-go-there lesson. Exactly. Stored failures sharpen the guardrails. They help prevent making the same mistake twice. It's fundamental to robust learning. And the research really backs this up with hard numbers, right? It shows this approach completely flips the script.
5:52When older systems tried to learn from failure data, they often just got confused or their performance dropped. But Reasoning Bank actually thrived on it. Absolutely. It was quite stark. The baselines, the older methods, often saw their performance stagnate or even degrade when failure data was mixed in. It just acted like noise to them. But when Reasoning Bang incorporated failures, take the Web Arena shopping subset, for example, the success rate actually jumped substantially. It went from about 46.5 % with only success traces up to 49.7 % when failures were included. Wow, that's over three percentage points.
6:26Yeah, that 3.2 percentage point jump, simply from learning how to transform failure into a constructive signal, that's significant in these complex tasks. Okay, this is where the value really clicks into place for me. That customer order query task example they gave. Can you give us the rundown on how the memory actually saved the day there? Sure. So the agent was asked a simple question. What is the date of my first purchase? Pretty straightforward, right? Seems like it. The baseline agent, the one with no strategic memory, just acted on the surface level. It looked at the page, saw a table labeled recent orders, checked that.
7:02And gave a recent date, presumably. Exactly. Which was completely wrong because that table only showed, well, recent orders. It completely missed the actual earliest purchase. Huh. So the old agent was basically like that friend who only checks the last five text messages and misses the crucial info from two weeks ago. Exactly. Perfect analogy. Now, the reasoning bank agent recalled a strategy, a strategy likely generated from a past failure or a hint related to finding complete data. And that memory item, that little distilled piece of wisdom, guided it to click the view all link right next to recent orders.
7:39Which revealed the full history. Precisely. That ensured it accessed the full order history, allowing it to correctly identify the true earliest order date, March 2, 2022. in their example. It wasn't just following steps. It applied learned judgment about data completeness. That's the difference. Okay. So we've got this robust, structured knowledge base now in Reasoning Bank. That's step one. Now let's talk about how the researchers use that memory to basically turbocharge the agent's learning process. So where Mattis comes in, memory-aware test time stealing. And you mentioned this is about scaling through depth, not just throwing more tasks at it.
8:16Exactly. You can think of Mattis as the agent instead of just trying one thing quickly. It sort of takes a deep breath and maybe brainstorms five different detailed ways to solve the same problem before committing to the first idea it had. See, traditional scaling often just means throwing more compute power at the task, which might just result in five disorganized, maybe even poor attempts generated really fast. Mattis allocates that extra compute differently. It uses it to generate abundant, diverse trajectories for a single task. Right. But the memory, the reasoning bank, is the key director here, isn't it?
8:49It's not just a brute force exploration. Not at all. That synergy is absolutely essential. The high quality memory from Reasoning Bank guides that initial deep exploration. It pushes the agent towards avenues that are more likely to be promising based on past experience. And then the diverse outcomes generated by that exploration, you know, seeing successful paths right next to dead end paths for the same problem. They create these really rich contrastive signals. Contrastive signals, meaning it learns by comparing. Exactly. It helps the memory extraction process forge even stronger, more generalizable lessons.
9:23It learns better because it can directly compare what works and what doesn't work, side by side for the same specific challenge. Makes sense. So how exactly did they implement this deep exploration you mentioned, Mantis? They actually studied two main implementations. First, they used what they called parallel scaling, or self-contrast. This basically involves generating multiple solutions, multiple trajectories simultaneously. simultaneously. Then the system compares them to filter out flaky or spurious solutions and identify the consistent, reliable patterns. Okay, parallel attempts. What's the second way?
9:58The second was sequential scaling, which they termed self-refinement. This involves iteratively refining the reasoning within a single trajectory. So the agent is constantly reviewing its own steps, its own scratch pad notes, you could say, and treating those intermediate reflection as valuable memory signals to improve its current path. Interesting, like thinking harder about its own process. Kind of, yeah. Using its own work in progress as a form of memory. And the results from the scaling experiments. Did they confirm that memory was the crucial ingredient? They did. MATESC consistently outperformed the sort of vanilla scaling approaches, the ones that tried to just use more compute without this sophisticated memory aggregation.
10:39It really proved that the synergy, the memory guiding the deep search, is essential for converting raw computation into actual cognitive gain for the agent. So we've seen better results, better effectiveness. But Reasoning Bank also made the agents faster, right? You mentioned, what, 16 % fewer interaction steps overall? That's quite a gain. It is. But when we look at that efficiency game, here's the question. Did this just mean the agent was getting stuck and giving up faster on hard problems? Or was it actually finding the solution more quickly when it succeeded? Yeah, that's the really important distinction to make.
11:12And the study clearly shows it was the latter. The savings didn't just come from, you know, cutting off failures prematurely. The reduction in steps was actually most pronounced on successful cases. In some domains, it took up to 2.1 fewer steps on average for a successful run compared to the baseline agent with no memory. So the memory is genuinely acting like a highly effective internal GPS. It's guiding, purposeful decision making, cutting out the wrong turns and the redundant loops. Absolutely. It helps the agent avoid that unnecessary or cyclical exploration that plagues less intelligent systems.
11:47Look at that shopping task example they documented. Right, the one where the baseline struggled? The baseline agent, totally lost in the digital aisles, got stuck in really inefficient navigation, clicking back and forth, searching for filters. It took a massive 29 steps to complete the task. 29 steps. Wow. The reasoning bank agent, leveraging its stored knowledge about category filtering, about how product search usually works on that type of site, completed the exact same task in only 10 steps. 10. That's almost a third of the steps. It's a huge tangible efficiency win. And it's generated entirely by that learned strategic knowledge stored in the memory.
12:25Purposeful action. Okay. So if this continuous updating of high quality structured memory is happening all the time, what kind of more advanced, maybe emergent behavior starts to show up? Does the agent actually get smarter in more complex ways over time? Well, yes. And that's perhaps the most fascinating part. The memory items themselves, the actual strategies being stored, they evolve. They evolve. How so? If you track the strategies that get generated and refined over time, they often start out quite low level, very procedural actions like, you know, find the correct navigation link for X, basic stuff.
12:59Then they tend to progress towards more adaptive self-reflection. Things like always re-verify identifiers before submitting a form, you know, learning to double check to reduce simple errors. Adding checks and balancing. Exactly. And finally, they can mature into quite complex compositional strategies, things like systematically use the search bar and the filters together, then cross-reference the task requirements against the UI elements displayed before attempting to extract the final result. Wow, that's multi-step strategic thinking. It is. The agent genuinely evolves from just performing simple actions to incorporating these higher-level critical reasoning checks and complex plans.
13:38It's learning how to approach problems more intelligently, not just storing raw facts. Hashtag Outtrack Outro. This has been a really compelling look at what seems like a significant next step in agent architecture. Yeah. So to summarize, Reasoning Bank tackles that core problem, the chronic forgetfulness of LLM agents. It does this by creating structured, transferable knowledge, extracted not just from successes, but crucially also from those counterfactual failures. And then this powerful structured memory synergizes really effectively with Matt, that novel scaling approach that focuses on depth of exploration.
14:12Right. The deep dive on a single problem. To achieve new levels of both effectiveness and efficiency. It's faster and better. Yeah. I think the overarching takeaway is pretty clear. This memory-driven experience, scaling, learning deeply from both success and failure, it really positions LLM agents to transition, to go from being just single-task problem solvers towards becoming genuinely adaptive, perhaps even lifelong learning systems. Closer to mimicking how real expertise develops over time. Exactly. They're moving closer step by step. And that idea you mentioned about the strategies themselves evolving from simple actions to complex compositional reasoning checks, that's pretty astounding.
14:48It makes you think. Which leads us to our final provocative thought for you, the listener. Since this kind of high-quality memory seems to enable agents to learn these complex emergent behaviors over time, what fundamental human skill do you think remains the hardest to abstract and distill down into one of these structured, transferable reasoning units for an agent to truly learn and generalize? What's still uniquely human in that learning process?
From the publisher
This paper introduces **ReasoningBank**, a novel memory framework designed to enhance Large Language Model (LLM) agents by distilling and structuring reasoning patterns from both successful and failed task trajectories. Traditional memory systems typically overlook failure experiences and lack the ability to abstract high-level reasoning, a limitation ReasoningBank addresses by creating **structured memory items** (title, description, content) that capture transferable insights. Furthermore, the paper proposes **Memory-aware Test-Time Scaling (MaTTS)**, which leverages this high-quality memory to guide diverse exploration, forming a positive feedback loop where memory improves scaling, and scaling enriches memory. Experimental results across multiple benchmarks, including WebArena and SWE-Bench-Verified, demonstrate that ReasoningBank significantly **improves success rates** and **enhances efficiency** by reducing the average number of steps required to complete tasks compared to existing memory approaches and memory-free agents.




