Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

11 Oct 2025 · 18 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Agentic Context Engineering (ACE) for self-improving LLM agents—how to adapt task context over time without catastrophic forgetting or “context collapse,” using structured, incremental updates and deterministic merging.

Guests

Two researchers/hosts discuss the ACE framework and its results; one leads the explanation of failure modes (brevity bias, context collapse) and the other details ACE’s architecture (generator/reflector/curator) and performance/cost findings.

Key claims

(1) Short, concise prompts cause “brevity bias” and remove crucial tool/domain details. (2) Monolithic context rewrites trigger “context collapse,” shrinking large knowledge to useless summaries. (3) ACE avoids this via delta bullet updates, deterministic non-LLM integration, and deduplication (“Grow and Refine”).

Notable examples

AppWorld adaptation step 60: context shrank from ~18,282 tokens to 122 after rewrite; accuracy dropped 66.7% to 57.1%. ACE used DeepSeek v3.1 to match IBM CUGA (GPT-4.1) on overall accuracy and beat it on harder splits; finance gains on Finer/Formula (~8.6%). ACE learns from execution feedback (errors/success) without labeled supervision; caveat: bad/noisy feedback can pollute the playbook.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Paradox of Context Adaptation

1:16 to 3:08

Explore the paradox of context adaptation and its impact on LLM performance.

“It's called ACE, agentic context engineering.”

Hidden Pitfalls of LLMs

3:08 to 7:26

Delve into the two major pitfalls of LLMs: brevity bias and context collapse.

“Which leads us right into the second problem, which sounds much more dramatic.”

Introducing ACE: Agentic Context Engineering

7:26 to 13:19

Learn about ACE, a framework designed to enhance context adaptation in AI.

“The main innovation here is getting rid of that monolithic rewriting entirely.”

Technical Innovations of ACE

13:19 to 14:00

Discover the technical innovations behind ACE and its efficiency gains.

“yes, that code was good, or no, that API call was wrong.”

Understanding Context Adaptation

14:00 to 14:37

Learn about the advancements in handling long context workloads in AI.

“I think so, especially because the underlying tech trends support it.”

Caveats of Context Engineering

14:37 to 15:21

Explore the limitations and potential pitfalls of ACE in AI systems.

“Despite all these clever innovations, ACE isn't foolproof.”

The Importance of Feedback

15:21 to 16:02

Understand the critical role of rigorous feedback in learning agents.

“Garbage in, garbage out still applies, even here.”

Selective Unlearning in AI

16:02 to 16:35

Discover the concept of selective unlearning and its significance.

“You can actually read the rules and tips it has learned.”

Accountable AI Systems

16:35 to 17:32

Learn how responsible memory management can enhance AI trustworthiness.

“The simple script that does the merging.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00When you think about these really advanced AI systems we're seeing now, you know, the autonomous agents that might code for you, or the tools doing complex financial analysis, they bump up against a problem. that feels, well, almost human memory, or rather the challenge of forgetting. It's actually a massive bottleneck. Yeah. When large language models, LLMs, tackle these complex multi-step jobs, they're relying less on their huge static training data and much more on the immediate context. Okay. So the prompt, the instructions, the memory they build up during the task. Exactly. That's the stuff injected right into their prompt.

0:37In this whole process, we call it context adaptation. It's incredibly valuable because, frankly, it's way faster and more flexible than retraining the entire model every time you wanted to learn something new. Right. Context adaptation seems like our best bet for quick AI improvements. But the systems using it now, they seem pretty brittle. Our sources are highlighting this core conflict when LLMs try to update that context, try to learn over time. The whole system is just prone to, well, catastrophic failure. It kind of loses its mind. That's a good way to put it. The system gathers all these insights.

1:09But the methods we've been using actually end up corrupting the very knowledge they're trying to save. It's a real paradox. So today we're doing a deep dive into a new framework designed to fix this. It's called ACE, agentic context engineering. Yep. And our mission here is to really understand how ACE manages to stop this catastrophic forgetting and how it drives these really impressive performance gains. We're talking on average over 10 % improvement for general agent tasks. Wow. Okay. And nearly 9 % on really specialized benchmarks, like, say, financial analysis where precision is everything.

1:45Okay, this is where it gets really interesting then. Why do the current systems fail so badly? The research points to two major kind of hidden pitfalls, right, that cause LLMs to just shed critical knowledge when they adapt their context. The first one's called brevity bias. Yeah, this one's a bit subtle, and it really tripped up some of the earlier attempts at prompt optimization. If you look back at methods like GPA-generalized expert prompting agents, they often aim for really concise, abstract instructions. The thinking being, shorter is better. Keep it simple. Kind of. The logic was, keep the prompt short.

2:18Maybe the LLM won't get lost or distracted. And, you know, there are practical reasons, too. Like cost. Context window limits. Exactly. Shorter context meant lower costs, fewer tokens. They were optimizing for token count. But here's the kicker. The research now clearly shows LLMs often perform better with detailed, longer contexts, rich contexts filled with high-fidelity information. So when those older optimizers push for brevity, they unintentionally force the system to strip out crucial details. We're talking specific domain insights, detailed instructions on how to use tools, even specific lessons learned from past failures.

2:57All the stuff an agent actually needs to navigate tricky situations. Precisely. All vital for agents working in complex areas. So focusing on being concise made the LLM forget the most important bits. Which leads us right into the second problem, which sounds much more dramatic. Context collapse. What exactly is happening there? Context collapse. Yeah, that's the big one. It happens when the system tries to do what's called a monolithic rewrite of a large context. So imagine the agent learns a new rule or a new piece of information. Okay. Often the task given to the LLM is, okay, rewrite this entire document of accumulated knowledge, but integrate this new piece.

3:32The whole thing from scratch. Pretty much. And the problem is when that context document gets really big, let's say, you know, 15 ,000, maybe 20 ,000 tokens, the LLM, under pressure to rewrite it all, tends to just get lazy. It compresses the whole thing into a short, often completely useless summary. And you've got a really stark example of this from the App World case study they did, right? Can you set the scene for us? Yeah, it's quite striking. So picture this agent. It's navigating a complex software application environment. It's learning, adapting. At Adaptation Step 60, its context playbook had grown huge, about 18 ,282 tokens.

4:08And it was doing well. It was thriving. It hit 66.7 % accuracy. Really good. Yeah. But then because the system forced it to rewrite that massive document, the very next learning step, context collapse hit hard. And the result? You said devastating. The context document just plummeted from over 18 ,000 tokens down to just 122 tokens. Wow, just gone. All that accumulated expertise just vanished. And predictably, performance crashed down to 57.1 % accuracy. Now think about this. The agent started with a 63.7 % accuracy baseline back when it had no adaptive context at all. So it didn't just forget.

4:49It actively made itself worse than if it had never even tried to learn in the first place. Exactly. Which really begs the question, if context is the key to knowledge and improvement, how on earth do we let it grow continuously without it basically self-destructing? Okay, so that tees up the solution perfectly. ACE, agentic context engineering. The big breakthrough seems to be structural, right? Instead of that fragile single document. Right. ACE treats the context completely differently. It's not a summary. It's a comprehensive evolving playbook. Think of it more like a seasoned engineer's field manual, something that's specifically organized to accumulate and index strategies, tips, rules, everything over time.

5:24And that organization is managed through a pretty clever workflow, modular, agentic. Yeah, it's a smart setup. It draws some conceptual inspiration from earlier ideas like dynamic cheat sheet, which also tried to separate up the doing from the reflecting. ACE avoids that cognitive overload, that pressure to rewrite everything, by distributing the learning process across three specialized components. They work together like a little team. Okay, let's break those down. First, you've got the generator. That sounds straightforward. That's your standard LLM, basically. So the one executing the actual task-generating code, producing reasoning steps for a new query.

6:01It does the work. Got it. Then, the critical part, the reflector. This is key. The reflector's only job is to critique what the generator just did. It examines the execution, praises what worked, sure, but crucially, what failed, and it distills concrete, generalized insights. Lessons learned. Okay, so it figures out why something worked or didn't. And then the third piece is the curator, the librarian. Huh, yeah, kind of. The curator is the architect, the organizer. It takes those distilled insights from the reflector, which are now structured, and formally integrates them into the evolving playbook.

6:35It files them correctly, basically. And this playbook isn't just a jumble of notes. It's structured. Super important. Using that app world example again, instead of a vague instruction like try harder next time, the ACE playbook contains specific structured reusable rules. Things like rule ID 12. Always resolve user identities from the correct source application based on credential checks. Much more useful. Or specific troubleshooting tips. Tip ID 45. If authentication fails, systematically check for API errors, list common errors before attempting to clean local credentials. It's specific, actionable, and formatted, so it can be slotted right into the next prompt easily.

7:16And it's that specific structured approach that avoids the collapse disaster we talked about. Okay, let's get into the technical weeds a bit. How does ACE actually ensure this playbook can scale without collapsing? The main innovation here is getting rid of that monolithic rewriting entirely. ACE uses what they call incremental delta updates. The context isn't one big essay anymore. It's maintained as a collection of itemized, structured bullet points. Think of them as little units of knowledge. Each one gets a unique ID, maybe some feedback counters tracking usefulness. So if the agent learns, say, one new troubleshooting tip.

7:52It doesn't rewrite the whole 18 ,000 token manual. No way. The curator just creates a tiny localized edit. They call it the delta context, like adding one new bullet point. Okay, but here's what jumped out at me from the sources. The curator creates the delta insight, but the actual merging into the main playbook is handled by deterministic non-LLM logic. What does that mean and why is that distinction so important? Ah, yes, that's the crucial safety net. Remember, context collapse is an LLM failure mode. It's the LLM under pressure buckling when asked to rewrite something huge. So ACE takes that rewriting function away from the LLM entirely for the final integration step.

8:31The non-LLM logic is just simple deterministic code. Think of it like a small rule-based script that handles structured data manipulation. So like find section X, append this new bullet with IDY. Exactly that kind of thing. It takes the old playbook, it's structured data, and the new structured bullet, the delta context from the curator, and it just merges them based on simple rules like matching IDs, appending to a specific section, maybe updating a counter. It's predictable. So the LLM provides the intelligence, the insight, but a simple, reliable algorithm handles the filing, the organization.

9:06That's really clever. It guarantees that if the playbook is 18 ,000 tokens long, adding a small delta makes it, say, 18 ,050 tokens long. It can't randomly shrink. Precisely. No compression risk at the integration stage. And this ties directly into their second technical innovation, the Grow and Refine Principle. Okay, what's that? It's about balancing growth with quality control. AC allows the playbook to continuously expand, appending those new bullets from the curator. That's the grow part. But it also needs to manage redundancy. So you don't end up with 10 slightly different versions of the same tip.

9:40Right. Right. The refined part uses things like semantic embeddings, mathematical representations of meaning to automatically identify and deduplicate knowledge points that are functionally the same. It keeps the playbook from getting bloated with repetitive info, keeps it manageable and fresh. And crucially for anyone thinking about deploying this stuff, these technical moves translate into massive efficiency gains, don't they? This isn't just about smarter AI. It's making adaptive AI cheaper and faster. Oh, hugely. The source data is pretty stark. ACE cuts down the adaptation latency the time it takes to learn and update by an average of 86.9 % compared to those older adaptive methods we mentioned earlier.

10:1987 % faster learning cycles. That's huge. And the cost implications follow. In offline adaptation, like training the agent initially, ACE needed 75.1 % fewer rollouts, fewer simulation runs than GPA. Less compute, less time. And in online, real-time operation, compared to other modular approaches like dynamic cheat sheet, ACE showed a 91.5 % reduction in latency and an 83.6 % reduction in the actual token dollar cost. Wow. When you're potentially running millions of queries through these agents, slashing your latency and cost by that much. That's the difference between a cool research paper and a system you can actually deploy affordably and scale up.

11:00Okay, so the tech is smart, it's efficient. What does this actually mean for performance in the real world? Let's look at the results, especially in those demanding areas. The agent success story they highlight on that App World benchmark is genuinely stunning. Get this. ACE, running on a significantly smaller open source model, DeepSea v3.1, managed to match the overall average accuracy of the current top ranked production agent. Which is? IBM CUGA, which is powered by the much larger proprietary GPT 4.1 model. Hold on. So a smaller open source model, just by using the smarter ACE context engineering framework, performed as well as a state of the art closed model.

11:37On average, yes. It's a powerful demonstration that sophisticated strategy and context structure can sometimes compensate for raw model size or brute force. And what's more, that ACE-powered deep-seek agent actually surpassed the proprietary agent on the harder test-challenge split of that benchmark. Meaning? Meaning its accumulated knowledge, its playbook, seemed to be more robust and better generalized for the really tough unseen problems. That's impressive. Okay. Beyond general agent tasks, you mentioned finance earlier. That's a field where precision is absolutely critical and mistakes can be incredibly costly.

12:15Absolutely. And we see strong results there, too. An average gain of 8.6 percent across critical financial benchmarks like Finder and Formula. Yeah, to give you some context. Yeah, what are those? Finer and Formula are specific natural language processing data sets designed to test deep understanding of financial concepts, recognizing financial entities, understanding their relationships in text. These tasks require very precise domain knowledge. For example, knowing and correctly applying complex XBRL rules. XBRL, that's the standard for business reporting. Extensible business reporting language, yeah.

12:48It standardizes how financial data is structured and communicated. It's very rule-heavy. ACE's ability to accumulate, structure, and reliably retrieve those precise rules seems to be why it excels on these tasks, where more general models might struggle or hallucinate. And maybe the most exciting piece for building truly autonomous systems, ACE can build these detailed, effective playbooks without needing labeled supervision. That's a huge point. It doesn't need a human sitting there reviewing every single attempt and saying, yes, that code was good, or no, that API call was wrong. It learns from natural execution feedback.

13:25Like, did the code actually run without errors? Exactly. Did it run successfully? Did the API call return the expected data payload? Did the action achieve the intended outcome in the environment? Those kinds of signals, which you often get automatically, are what guide the reflector component. So it learns from trial and error, essentially, but in a very structured way. Right. And that self-correction loop, based on real operational feedback, is absolutely essential if you want to build LLMs that can continuously and independently improve themselves out there in the wild without constant human handholding.

13:59So connecting this back to the bigger picture, ACE really strengthens the case that focusing on adaptive context might be a better investment in many cases than just constantly fine-tuning massive models. I think so, especially because the underlying tech trends support it. Modern machine learning systems, cloud platforms, they're all getting much better at handling long context workloads efficiently. Things like KVCache reuse make retrieving information from large contexts faster and cheaper than ever. So context adaptation becomes not just smarter, but also more practical and cost-effective, and more interpretable than tweaking billions of weights inside a black box.

14:35Definitely more interpretable. But before we wrap up, we do need to touch on the caveat. yet, the limitation. Despite all these clever innovations, ACE isn't foolproof. Okay, where's the catch? Its effectiveness really hinges on the quality of that feedback signal, and crucially, on the reflector's ability to interpret it correctly. Meaning, if the feedback itself is noisy or misleading. Or if the reflector component, the one that's supposed to distill the insights, draws the wrong lessons from the execution traces. If it extracts meaningless correlations, or just plain bad advice. Then the curator will just dutifully file that bad advice into the playbook.

15:12Exactly. The curator trusts the reflector, so the context could gradually become polluted with misleading or even harmful instructions. The learning agent still needs to be, in a sense, a rigorous critic of its own performance and the signals it receives. Garbage in, garbage out still applies, even here. That's a really important point. Garbage in, garbage out. But assuming you have reasonably good feedback, the implications of getting this right, of having structured, evolving context, they seem huge, especially for responsible AI development. They really are. Because this kind of context adaptation, like ACE, feels central to the future of continuous and responsible learning.

15:51The key is that these playbooks ACE creates are detailed, they're structured, and crucially, they are human interpretable. They aren't just opaque numbers inside a neural network. Right. You can actually read the rules and tips it has learned. And that interpretability unlocks something absolutely essential for future AI systems operating under real-world regulations, selective unlearning. Selective unlearning. You mean the ability to go in and specifically remove information? Precisely. The ability to actively remove outdated information, maybe incorrect facts it learned earlier, or critically private or sensitive information that shouldn't be retained.

16:28Think about legal requirements. Like GDPR's right to be forgotten or CCPA mandates in California. Exactly. Those kinds of things. If a model's playbook contains a piece of context, maybe a specific rule derived from data that's now considered private or a strategy that's become obsolete or even harmful, you can potentially use that deterministic non-LLM logic we talked about. The simple script that does the merging. Right. You could use similar logic to find that specific bullet point by its ID, remove it cleanly, and verify that it's actually gone, all without needing to completely retrain the foundational model or mess with the underlying weights in complex ways.

17:07So you get accountable, auditable memory management. Which is arguably an essential building block, maybe even a prerequisite for any sophisticated AI system that we're going to deploy widely and trust in sensitive domains in the future. A system that can not only remember and learn more effectively than ever before, but one that can also responsibly forget when it needs to. That's maybe the final provocative thought for you, the listener, today. The next big frontier in AI might not just be how well these systems learn, but how intelligently and how countably they manage what they know.

From the publisher

This paper introduces Agentic Context Engineering (ACE), a novel framework designed to enhance the performance of Large Language Models (LLMs) in complex applications like agents and domain-specific reasoning by evolving their context, or "playbook." ACE addresses two key limitations of prior context adaptation methods: brevity bias (the loss of detailed domain knowledge for conciseness) and context collapse (where iterative rewriting erodes information). Through a modular process of generation, reflection, and curation, ACE builds contexts that are structured, incremental, and comprehensive, leading to superior performance on benchmarks like AppWorld and financial analysis tasks. Critically, the framework achieves significant improvements, such as a 10.6% gain on agents, while also reducing adaptation latency and cost compared to strong baselines by using localized, delta updates instead of monolithic rewrites.

More from Best AI papers explained

All 475 episodes
Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsBest AI papers explained · 18 min
Listen in VO