Learning to Continually Learn via Meta-learning Agentic Memory Designs

20 Feb 2026 · 20 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The “Groundhog Day” problem in agentic AI—stateless foundation models lose context between sessions—and how ALMA (Automated Meta Learning of Memory Designs for Agentic Systems) uses a meta-agent to automatically write better memory architectures for other agents.

Guest backgrounds

No guest names or bios are provided in the transcript; it’s a single host-led discussion.

Key claims

Humans are “bad librarians” because handcrafted memory (lists/graphs/vector RAG) is brittle and task-specific. ALMA’s meta-agent searches over raw Python code to invent memory systems, evaluates them by running worker agents, and iterates hundreds of times. It beats human-designed baselines across ALFWorld, TextWorld, MiniHack, and Babacup, while costing about nine cents in experiments. It uses sandboxing and code-safety checks.

Notable examples

ALFWorld: replaces inefficient lists with an affordance graph and spatial experience tracking (e.g., apples near dining tables). Baba Is AI: abandons spatial graphs and builds rule-block perception parsing plus a strategy library for dynamic “wall is stop/win” logic.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introducing ALMA: AI's New Memory Architect

2:18 to 2:45

Learn about ALMA, an AI designed to create better memory systems for other AIs.

“And our mission today is to unpack this because the headline here is pretty wild.”

The Bottlenecks of Human-Designed Memory

2:45 to 4:34

Understand the shortcomings of current human-designed memory systems for AI and the need for more adaptable solutions.

“It's AI engaging in actual introspection.”

ALMA's Framework: Meta-Agent Overview

4:34 to 4:59

Discover ALMA's meta-agent framework and its role in improving AI memory design through reflection and coding.

“So the ultimate goal here is continual learning.”

How ALMA Evolves Memory Designs

4:59 to 6:34

Explore the iterative process ALMA uses to create and refine memory structures for AI tasks.

“It's a bit of a mouthful, but the concept is so cool.”

Diverse Memory Designs: The Nature of Exploration

6:34 to 8:52

Examine how ALMA maintains a diverse set of memory designs for exploration and potential innovation.

“It could invent a squiggly line time series hybrid if that mathematically helps solve the problem.”

Case Studies: ALMA in Action

8:52 to 12:17

Review two case studies from ALMA’s applications in household tasks and dynamic logic puzzles.

“It's programming that allows for serendipity.”

Performance and Cost Efficiency of ALMA

12:17 to 13:19

Assess how ALMA performs against human counterparts and its surprising cost efficiency in running designs.

“Okay, so the big question for anyone listening, does it actually work?”

The Historical Context of AI Memory Development

13:19 to 14:00

Contextualize ALMA's advancements within the historical evolution of AI and memory systems in computer science.

“Because the memory it designs is hyper-efficient.”

The Evolution of AI Systems

14:00 to 15:00

Learn about the phases of AI development and the shift to self-designed systems.

“it knew exactly how to leverage that extra data without getting bogged down.”

The Role of ALMA in AI Memory Design

15:00 to 17:00

Discover how ALMA enables AI to design its own memory systems and the implications.

“We are currently handcrafting the systems, the memory, the tools, the retrieval flow.”
Show all 13 chapters

Understanding the Black Box of AI Memory

17:00 to 18:20

Examine the challenges of transparency in AI memory systems designed by the AI itself.

“We know that looking up the word Apple gives us location.”

Personalization of AI Tools

18:20 to 19:20

Learn how ALMA can create personalized AI tools that grow alongside users.

“It moves AI from being a generic, out-of-the-box tool to a deeply specialized expert that literally grows with you.”

AI's Role in Curating Memory and Identity

19:20 to 20:14

Explore how AI may choose what to remember, impacting its identity and functionality.

“So we are moving from AI that just passively records data on a hard drive to AI that actively curates its own biography.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, I was actually messing around with the latest GPT5 update on my phone this morning. Oh, yeah. Trying to get it to write your emails for you. Trying to get some actual work done, yeah. But I realized something while I was doing it, something that has secretly been driving me crazy for years now. It's what I call the Groundhog Day problem. Groundhog Day. Right. Like waking up to the same Sonny and Cher song every single morning. Exactly. But with AI, think about it for a second. I spend an hour working with an agent, right? Let's say it's a really sophisticated coding assistant. We solve a complex problem together.

0:34I teach all the context of my specific project. We get into a serious flow state. And then you close the window? And then I close the window. The next time I open it, poof, gone. It greets me like a total stranger. Hello, how can I help you today? It's like we never even met. It is genuinely maddening, isn't it? I mean, that is the classic statelessness of foundation models. They are brilliant. Absolute geniuses in the moment. Right. They're polymaths who know everything from medieval poetry to Python. but they have the memory span of a goldfish. Every single interaction is a completely clean slate.

1:10And I know right now, you listening to this, you're probably shouting at your device, but wait, we have memory modules now, we have RAG, we have vector databases. Yes, we bolt on these digital notebooks for the AI to write things down in. But honestly, it still feels clunky. It feels like a Band-Aid on a bullet wound. Well, it feels clunky because it essentially is a Band-Aid. And here is the twist that people in the industry rarely talk about. Up until now, humans have been the ones designing those notebooks. We're the ones deciding how it remembers. Exactly. We sit there and say, okay, write your summary here, store the user preferences over there, put the embeddings in this specific vector store.

1:49So we're basically acting like librarians trying to organize their shelves. We're guessing how an alien intelligence wants to store its own books. Precisely. But what if we're just bad librarians? What if the way we organize information, which is based heavily on human logic and human narratives, isn't actually the way an AI needs to organize it to be effective? And that is exactly the question we're tackling today. Welcome to the Deep Dive. We are looking into some really fresh, groundbreaking research today. It introduces a system called ALMA. And our mission today is to unpack this because the headline here is pretty wild.

2:25We aren't designing the memory systems anymore. No, we really aren't. With LMA, researchers have built an AI agent whose sole job is to write code to invent better memory systems for other AIs. And the scary, wonderful part is that it's beating the human design. Hands down. So we basically have an AI architect building the brain space for other AI workers. That is the perfect way to put it. It's AI engaging in actual introspection. It's engineering its own mind. So before we get into the how of all this, because I saw some of the code it actually wrote and my jaw just dropped, let's talk about why we even need this.

2:57Why is the human way of doing things such a bottleneck? You mentioned we're bad librarians. Right. It really comes down to rigidity versus adaptability. Right now, if I want to build an AI agent to play a text adventure game or maybe manage a household robot. Or solve complex logic puzzles. Exactly. I, the human engineer, have to sit down and meticulously craft the memory architecture for that specific task. Right. I might say, okay, use a list to store the robot's inventory. or use a graph structure to map out the rooms. Which sounds perfectly reasonable. I mean, that's just standard programming.

3:33You define your data structures. It is reasonable, but it's entirely handcrafted. And because of that, it's incredibly brittle. A memory system designed for a cooking assistant, where you need to remember dietary restrictions, flavor profiles, what's currently in the fridge, that structure is terrible for a coding assistant. Right, because a coding assistant needs to remember variable definitions, function calls, API documentation. Exactly. It's like trying to use an old school Rolodex to organize your Spotify playlist. The structure just fundamentally does not fit the task. That's a great analogy.

4:05And here's the real kicker. There are millions of potential tasks in the world. We simply cannot handcraft a bespoke memory system for every single one of them. Humans are way too slow. So we just rely on generic templates. Right. Things like Reason Bank or G-Memory. And those are the current state-of-the-art human designs. They're very good generic solutions. But in a world of highly specialized agents, one size definitely does not fit all. So the ultimate goal here is continual learning. We don't just want an agent to solve a task once. We want it to learn how it solved it, store that experience efficiently, and then be faster and smarter the next time you ask it.

4:46Yes. And if the memory structure limits that process, the agent just hits a ceiling. It literally can't get better because it has nowhere to put the new skills. It's like trying to memorize an entire symphony, but you only have a notepad big enough for a grocery list. So enter ALMA, which stands for Automated Meta Learning of Memory Designs for Agentic Systems. It's a bit of a mouthful, but the concept is so cool. Tell us about this, because this isn't just a database. It's a whole framework. It's a framework built entirely around what they call a meta-agent. Think of this meta-agent as a senior software architect.

5:18Its job isn't to do the actual work or solve the end user's problem. Its job is to design the workspace for the worker bees who do solve the problem. Walk us through the loop. How does this architect actually work? Because it's not just tweaking parameters under the hood, right? It's not just turning a dial from 1 to 10. No, and that is the major breakthrough here. The search space it operates in is code, raw Python code. Wow. The meta agent follows a very human-like design process. First, it ideates. It looks at the problem. Let's say it's a game called ALF World where a robot has to navigate a house.

5:51It looks at past memory designs it has tried for this, and it asks, what worked? What failed? Why did the robot get stuck in the bathroom for 20 turns? So it's actively reflecting on its past attempts, just like a human developer looking at a bug report. Yes. Then it moved to the coding phase, and this is where it gets really wild. It writes actual Python classes. It defines brand new data structures from scratch. It's just writing scripts. You might write a completely new class called Affordance Memory Layer or Strategy Recall. It is actively programming the underlying structure of the memory itself.

6:25See, that's the part that blew my mind. Because Python is Turing-complete, ALMA can theoretically invent any kind of memory structure imaginable. It is not limited to some preset drop-down menu of list, graph, or vector. Exactly. It could invent a squiggly line time series hybrid if that mathematically helps solve the problem. That's insane. Theoretically, yes, it can do anything. So once it writes the code, it moves to evaluate. It spins up a worker agent, the agentic system, installs this brand new memory module into its brain and sets it loose on a task. And it just watches. It's watching the crash test dummies.

7:00It really is. It monitors the logs. Did the agent fail to find the apple in the kitchen? Did the memory code crash halfway through because of the syntax error? Oh, right. Because it's writing real code, it can make syntax errors. Absolutely. So it takes all that feedback, reflects on it, puts the result into its archive, and then just starts the loop all over again. And it does this hundreds of times, just evolution at warp speed. Exactly. And speaking of evolution, there's this visual in the research of what they call the memory design archive. And it literally looks like a family tree. Yeah, I saw that.

7:32It branches out. It doesn't just look like a straight linear line of good to better to best. And that is one of the most fascinating aspects of LMA. It uses a technique called open-ended exploration. It doesn't just keep the best design and throw the rest away. It intentionally keeps a wide diversity of designs. But why keep the mediocre ones? I mean, if I'm building a race car, I want the fastest engine. I don't care about the one that sort of just sputters along. Well, think of it like biological evolution. Or even better, imagine you are a pioneer trying to invent the airplane. Okay. If you strictly optimize for staying on the ground safely, you are going to build a really fantastic car.

8:11But you will never build wings. Because early wings look like terrible car parts. That makes total sense. They create drag, they're flimsy, they don't help you drive on the dirt road at all. Exactly. LMA realizes that a specific memory design might have a really low success rate right now, but it might contain a deeply interesting structural idea. Maybe a totally new way of linking two objects together so it keeps it in the archive. It doesn't just toss it in the trash. Right, because later on, it might combine that weird idea with another one to create a massive breakthrough. If you only used what we call greedy search, which is always picking the immediate winner, you get stuck in a local trap.

8:51You optimize the Rolodex perfectly, but you never actually invent the database. I love that. It's programming that allows for serendipity. So let's get concrete here. Because memory is a pretty vague word, what did LMA actually invent? Give us the aha moments from the research. Let's look at two very different domains to show you just how adaptable this framework is. First, let's talk about ALF world. Right. This is the household simulation. So the agent is basically acting as a robot butler. It gets prompts like, put a clean apple in the fridge or heat the mug in the microwave. Exactly. Simple stuff for a human, but surprisingly tricky for stateless AI.

9:28Now, human engineers usually design memory for this as a simple list. I am holding an apple. I am standing in the kitchen. Which works fine until you have 20 items and five different rooms to keep track of. Right. It gets overwhelming. So LMA looked at this environment and quickly realized a list is highly inefficient here. And it invented a completely new structure called an affordance graph. Affordance. That's a design term, right? It means what an object naturally allows you to do, like a chair affords sitting. Precisely. The memory system it wrote started tracking that microwaves are for heating, not washing.

10:03Sinks are for cleaning. But it actually went a step further than that. It built a spatial experience tracker. What does that mean in practice? It started actively remembering that apples are usually found near dining tables, not near toilets. Whoa. So it wasn't just remembering a list of I saw an apple. It was remembering the actual physical logic of the house. And no human ever told it to do that. It figured out entirely on its own that to solve the task of cleaning an apple, it needed a map of relationships, not just an inventory list. It essentially derived common sense physics from trial and error.

10:37That is incredible. Okay, so that's case study one. Let's talk about case study number two. Baba is AI. This is based on that indie puzzle game, Baba is You, right? Yes. It's a really unique logic puzzle game where the rules of reality change dynamically. You see these blocks of text on the screen. If you push a block that says wall is stop, you can't walk through walls. Standard video game stuff. But if you push the blocks around to spell wall is win, suddenly touching the wall instantly wins the level for you. So a spatial map of where the apple is, like the one it invented in the last example, would be completely useless here because the rules of reality keep shifting.

11:14Walls might turn into water on the very next turn. Exactly. And Alibay figured that out. When it initially tried to use those spatial graphs, the very thing that works so perfectly in the kitchen, it failed miserably in the puzzle game. So what did it do? It completely threw that architecture out and coded a strategy library and a specialized module for perception parsing. Perception parsing. It built a system specifically to scan the game environment for those rule blocks, wall, is stop, and synthesize a dynamic plan based strictly on the current rules. It realized, I don't need to know where I've been.

11:50I need to know what the laws of physics are right at this exact second. That is wild. It specialized the memory structure completely. Household tasks got spatial maps, and logic puzzles got dynamic rule books. It automatically specialized. That is the true power of the meta-agent. It tailors the entire cognitive architecture to the specific problem at hand. A human engineer probably would have tried to force the map solution onto the puzzle game and failed. And ALMA just wrote new code instead. Exactly. Okay, so the big question for anyone listening, does it actually work? I mean, does it legitimately beat the stuff the top human engineers have built?

12:26It does. Across every single domain they tested, ALF World, TexWorld, Minihack, Babacup, ALMA's generated designs beat the absolute state-of-the-art human designs. Wow. And we aren't talking about small, marginal improvements here. In some cases, it was significantly better at reusing its past experience to solve completely novel problems. But what about the cost? Because usually when we talk about advanced AI memory or agents autonomously writing code, the first thing I hear is expensive. Lots of tokens, huge context windows, tons of processing power. That's the most surprising fact in the whole research.

13:03You naturally ask, is this expensive? And the answer in AI is almost always yes. But here, it's actually no. The end-to-end memory cost for Elmay's designs was lower. We're talking around nine cents in the experiments, compared to much higher costs for the human baselines. Wait, how is it cheaper if it's writing and running more complex code? Because the memory it designs is hyper-efficient. It doesn't store junk. A bad memory system, like a human-designed list, dumps everything into the context window. Every single failed step. Every redundant detail. And all of that costs money to process on the API side.

13:38Exactly. LMA learned to ruthlessly curate. It learned to store only the distinct signals that actually help solve the task. So it's basically the difference between a student highlighting the entire textbook versus just writing down the three formulas you actually need for the final exam. That's spot on. And because of that, it scales better, too. When they gave the LMA designs more past data to pull from, the performance shot up much faster than the human designs. it knew exactly how to leverage that extra data without getting bogged down. This really feels like a major turning point. We're moving away from the foundation model itself being the whole product.

14:15If we connect this to the bigger picture, it follows a very clear historical trend in computer science. Think about computer vision back in the early 2000s. Back when we were manually trying to teach computers how to see. Right. We handcrafted what we called edge detectors. Human engineers literally wrote code that said, if there is a sharp contrast in pixels right here, that is an edge. Right, very rules-based. Then deep learning came along and said, stop writing the rules. Let the neural network learn the features itself. And it worked way better. That was phase one. And phase two was neural architecture search, right?

14:49Yes. We used to hand design the actual layers of the network, convolutional networks, recurrent networks. We argued endlessly over how many layers to have. Then we just built AI that searched for the best architecture for us. And now we're in phase three. Phase three. We are currently handcrafting the systems, the memory, the tools, the retrieval flow. ALMA is the start of AI learning its own system design. It's what some people are starting to call software 2.0 or system 2.0. So the GPT-5 or whatever massive model is sitting at the center is just the engine block. It's just raw horsepower. But the driver, the thing that actually remembers where it's going and knows how to drive the car, is this custom-built software wrapper that the AI designed for itself.

15:33Precisely. The foundation model itself becomes a commodity. Eventually, everyone has GPT-5. The real competitive advantage is going to be in the cognitive architecture that wraps around it. And Anna Mae proves that AI is simply better at building that architecture than we are. Now, I do have to play devil's advocate for a second here. We have an AI writing raw Python code to modify its own internal memory structure. It's essentially rewriting its own brain. Isn't that the exact plot of every sci-fi movie where things go horribly wrong? It raises a very real safety question. When you let an agent autonomously write code, you absolutely have to have sandboxing.

16:10For the non-coders listening, define sandboxing in this specific context. It means running that newly generated code in a completely secure, isolated environment, A space where it cannot access the internet, it cannot touch your host operating system files, and it cannot escape the simulation it's running in. ALMA implements this strictly. It evaluates the code to ensure it's structurally safe before deploying it to the worker agent. We really don't want the AI writing a memory module that accidentally or intentionally deletes your hard drive just to save space. Yeah. I decided the most efficient way to manage my memory is to delete all other users on the network.

16:48Let's definitely avoid that. Please. But even with all the safety checks in the world, this shift towards black box memory is significant. When a human designs a SQL database, we know exactly how it works. We know that looking up the word Apple gives us location. It's totally transparent. But when Alame designs it, it might create a bizarre vector embedding system combined with some decaying urgency score and a complex retrieval logic that is fundamentally very hard for a human mind to interpret. it. We might see the agent succeeding brilliantly, but not fully understand how it is recalling the information.

17:25We're handing over the keys to the library. And the librarian is starting to speak a language we don't entirely understand anymore. That is thrilling and honestly, slightly terrifying. But I guess it's the ultimate trade-off of automation. We get massive efficiency and new capabilities, but we inherently lose some control and transparency. But think about the incredible upside here. Think about what this means for you, the user. Yeah, if I'm just an everyday user of AI, why should I care that Alame is out there writing Python code in a lab somewhere? Because the AI tools you use every day are about to get infinitely more personalized and capable.

18:00Imagine having a coding assistant that doesn't just remember your code snippets, but actively evolves a specific bespoke memory structure tailored perfectly to your unique coding style and your company's project architecture. Or a legal AI that develops a totally unique way of cross-referencing case law that no human lawyer has ever even thought of. Exactly. It moves AI from being a generic, out-of-the-box tool to a deeply specialized expert that literally grows with you. Yeah. It's the difference between hiring a temp worker who forgets everything at 5 p.m. and finding a dedicated business partner who builds a perfectly customized filing system for your company over the course of years.

18:42But that partner is a machine. A machine that programs itself. Here's where it gets really interesting to me, though, just as a final thought. If ALMA is out there deciding how to remember things, isn't it eventually going to start deciding what is actually worth remembering? That is the big philosophical shift we're heading toward. Right now, we usually dictate to the AI, hey, save this interaction. But a truly optimized, self-designed memory system might look at a prompt and decide, you know what, that conversation wasn't useful at all. I'm deleting it. Or it might decide this tiny, seemingly insignificant detail about the user's tone of voice.

19:17That's actually crucial data for predicting future tasks. I'm locking that in permanently. So we are moving from AI that just passively records data on a hard drive to AI that actively curates its own biography. And its own skills. It is curating what makes it effective. It is defining its own identity through its memory. Because memory is identity, after all. If the AI actively chooses what to keep and what to discard, it is essentially choosing who it becomes. Well, we started this deep dive complaining about AI having amnesia. Now I'm actively worried about an AI that decides my jokes aren't mathematically efficient enough to be worth remembering.

19:54You might want to refine your material before the next software update. The algorithm is a very harsh critic. I'll definitely work on it. This has been a fascinating deep dive into LMA and the entire future of agentic memory. It's a world where the code writes the code and the memory literally builds itself. Something for all of you listening to Mollover today. Keep learning. And maybe write some things down on actual paper before the AI decides to do it for you. Thanks for listening.

From the publisher

This paper details the development of **ALMA**, a meta-learning framework designed to generate and refine **agentic memory structures** for AI systems. This system utilizes a **Meta Agent** to synthesize specialized memory layers, such as **strategy libraries**, **spatial experience graphs**, and **reflex rules**, which help agents navigate complex environments like TextWorld and MiniHack. To ensure operational security, the framework employs **isolated sandbox environments** and human oversight to prevent unintended behaviors or harmful code execution. The research demonstrates that these **autonomous memory designs** significantly improve task success rates while maintaining lower computational costs compared to traditional retrieval methods. Additionally, the sources include extensive **technical documentation** and utility functions for parsing environmental data, managing databases, and distilling past experiences into actionable logic. These components work together to enable agents to **continually learn** and adapt their strategies based on historical performance and environmental feedback.

More from Best AI papers explained

All 475 episodes
Learning to Continually Learn via Meta-learning Agentic Memory DesignsBest AI papers explained · 20 min
Listen in VO