In short
The episode argues today’s large language models are “static” after pretraining because they lack rapid online memory consolidation, leading to an “anterograde amnesia” effect. It introduces nested learning (NL) as a framework where deep architectures are better seen as nested optimization problems with multiple update frequencies (fast vs slow), inspired by brain oscillations and consolidation (synaptic vs systems).
Key claims
knowledge is split between short-term context windows and long-term feed-forward parameters; NL makes internals “white-box” by tracking gradients/context across levels; standard optimizers like momentum/Adam are reframed as associative memory modules compressing gradient history; replacing linear momentum with an MLP yields “deep optimizers” (DMGD).
Notable examples
continuum memory system (CMS) as chains of MLP blocks updated at different frequencies; HOPE (High Order Self-Referential Permutation Encoder) combines CMS with self-modifying attention.
Guests
No specific guests are named in the transcript.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Static Nature of Large Language Models
0:45 to 3:06
Exploring the limitations of large language models and their inability to adapt post-training.
“Are we missing something fundamental, something biological maybe, that allows for ongoing adaptation?”
Understanding Knowledge Storage in LLMs
3:06 to 5:29
Analyzing how LLMs store knowledge and the implications of their memory structure.
“And nested learning is designed fundamentally to bridge that gap.”
Introduction to Nested Learning
5:29 to 7:40
Introducing the concept of nested learning and how it aims to solve LLM limitations.
“That sounds like how long-term memory should work, right?”
Mathematical Framework of Nested Learning
7:40 to 10:08
Examining how nested learning reframes model dynamics through optimization problems.
“Okay, so if the momentum term is just a simple linear memory, the obvious next step is to replace it with something more powerful.”
The Continuum Memory System and HOPE
10:08 to 12:28
Diving into the continuum memory system and the new model HOPE derived from nested learning.
“A standard transformer has that single static FFN block.”
Future Implications of Nested Learning
12:28 to 13:06
Discussing the broader implications of nested learning for AI and cognitive functions.
“And here's a final thought to leave you with.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. You hand us the sources, and we handle the complexity, turning the cutting edge of research into usable insight. Today, we're digging into advanced machine learning, specifically tackling what might be the biggest challenge facing today's AI. Large language models. They're incredible, sure, but they're fundamentally static. Static, yeah. That's a good word for it. Think about it. They do amazing things. But once that huge pre-training phase is done, they're basically frozen in time. masters of what they learned way back when, but they really struggled to pick up new skills or update themselves continuously.
0:35They just stopped learning. It really is like hitting a cognitive wall, isn't it? We've pushed scale, added layers, more parameters, but it isn't fixing this core limitation and the static nature. It makes you wonder, right? Are we missing something fundamental, something biological maybe, that allows for ongoing adaptation? Do we need a whole new learning paradigm, something beyond just stacking components into end. Exactly. And that's our mission for this deep dive. We're unpacking a really fascinating new paradigm from our sources called nested learning, or NL. The researchers behind it argue that our usual picture of deep learning architectures, it's kind of an illusion.
1:11It hides what's really going on computationally. So we'll explore what NL actually is, how it changes our view of everything, even basic optimizers, and we'll introduce the new model build on these ideas. Wood up here. Okay, but before we jump into NL as the solution, let's make sure we really feel the weight of the problem with current LLMs. It boils down to how and where they store knowledge. Right. If you look inside a standard LLM, knowledge lives in basically two spots. You've got the immediate context window, like the short-term memory of the current conversation, and then you have the vast older knowledge baked into the parameters of the feed-forward networks, the FFNs.
1:48That's usually seen as the model's long-term memory, right? Precisely. And this setup leads to a really striking and honestly kind of worrying analogy. Current LLMs behave a lot like someone with anterograde amnesia. That's the condition where you can't form new long-term memories after something happens. You remember the past, but new experiences just don't stick. They stay fleeting in the short term. Wow. So after pre-training, the LLM is caught in that loop. New info comes in through the context window, the short-term buffer. but it never really updates or reshapes those long-term FFN parameters.
2:23It just can't consolidate new knowledge or skills permanently from what it's seeing now. That's the core issue. Compare that to us, to the human brain. We learn continually throughout our lives thanks to neuroplasticity. The sources point out two key consolidation processes we use. There's this rapid online consolidation, synaptic consolidation that happens almost instantly to stabilize new memory traces, makes them less fragile. But then there's also a slower offline consolidation, systems consolidation. That's like reorganization happening over hours, maybe days, often when we sleep. And the key takeaway here is that LLMs are missing that first step, that rapid online consolidation.
2:59So their short-term experiences are basically just discarded, never integrated properly into their deep structure. Exactly. And nested learning is designed fundamentally to bridge that gap. It's about building models that can do that immediate continuous transfer from short-term experience to long-term structural memory. Okay, so if the problem is this lack of immediate memory consolidation, how does nested learning as a framework actually propose fixing it? What is NL, mathematically speaking? Well, NL completely reframes how we even look at a model. It says, forget just thinking about it as a flat sequence of layers.
3:35Instead, picture a model as a set of nested optimization problems. These can be multi-level, maybe even running in parallel. And crucially, each of these little internal optimization problems has its own goal, its own parameters to tweak, and its own specific flow of information or context flow. Okay, that sounds complex. What's the big idea there? The main takeaway. The big idea is about making the internals transparent and understanding why deep learning works. NL suggests architectures learn effectively because they're constantly compressing their own context flow. By seeing the model as these nested optimization tasks, you move away from a black box.
4:10It becomes more of a, well, a white box mathematically where you can actually track how gradients and context influence different parts at different levels. This sounds very much inspired by the brain's timing. You mentioned the fast and slow consolidation in humans. How does NL translate that multi-timescale idea into the model's parameters? Yeah, Yeah, it draws direct inspiration from brain oscillations, you know, delta, theta, gamma waves. These operate at vastly different frequencies, from super slow, like half a cycle per second, to really fast, maybe 100 cycles per second. The brain doesn't update everything at once with one big clock tick.
4:45NL mirrors this. It assigns different parameter sets within the model their own specific update frequency. Let's call it$5. So computationally, what does that look like? What's the difference between parameters with a high update frequency versus a low one? The frequency, IF Lara, basically determines how often a set of parameters gets updated. So parameters in the, let's say, lower levels, the ones dealing with the immediate raw input, they get updated very quickly. High frequency, they consolidate details right away. But parameters in the higher, more abstract levels, the ones integrating, meaning over longer stretches, they update much more slowly.
5:22Low frequency, they need that longer timescale to integrate. I see. So slower frequency means integrating information over a longer period before the parameters actually change. That sounds like how long-term memory should work, right? More stable? Precisely. And that's why the NRL proponents call the traditional view of deep learning a flattened image. You look at a transformer attention, then FFN it looks sequential. Flat. But that view hides the nested optimization happening underneath. It makes it seem like all parameters learn on the same schedule, with the same global objective driving everything at once.
5:53NL makes that hierarchy based on update frequency,$5, explicit. It defines the computational depth in terms of time. That really does shift the perspective. Okay, if we accept this multi-level optimization view, it forces us to look at the individual components differently. And this, I think, is where NL gets really surprising, right? It redefines even the most basic tools we use. Oh, absolutely. This is a key insight from the NL framework. It suggests that those standard gradient-based optimizers we use all the time think gradient descent with momentum. Or even, Adam, they aren't just algorithms.
6:27They are, in fact, associative memory modules. Their actual function is to compress the history of past gradients into their own internal parameters. They remember the optimization trajectory. Whoa, okay. Let's unpack that. Basic gradient descent, GD, is simple. Step away from the gradient. How does adding momentum suddenly make it a two-level memory system? Right. So simple GD is just one level of optimization. But when you add that momentum term, often called MONI plus one, you've actually introduced a second nested optimization process. The main model weights are updated based on this momentum value, but the momentum value itself, it's being learned, it's acting like a parameter.
7:05The inner level, the momentum is essentially a memory module storing a moving average of past gradients. It's compressing that history. So the momentum term isn't just helping us go faster downhill. It's actually remembering the recent path, like a simple form of gradient memory. Exactly. It's a primitive, keyless associative memory for gradient dynamics. But here's the catch. Traditional momentum uses very simple linear math, usually just vector addition and scaling. That means it has very limited capacity to capture complex patterns in the gradients. And that realization opens the door to something called deep optimizers.
7:41Ah, okay. Okay, so if the momentum term is just a simple linear memory, the obvious next step is to replace it with something more powerful. You got it. That's the thinking behind deep momentum gradient descent, DMGD. Instead of that basic linear function for tracking momentum, why not swap it out for something with more capacity, like a multilayer perceptron, an MLP? And what does using a deep MLP memory for momentum buy you, practically speaking? A linear memory can only really capture, well, linear trends or averages in the gradients. It's quite rigid. A non-linear memory, like an MLP, can learn much more complex, subtle patterns in the gradient history.
8:20It could potentially detect things like oscillations or sharp turns in the loss landscape and adapt the update much more intelligently. It gives the optimizer a much richer capacity to understand and react to the dynamics of the learning process itself. You're essentially teaching the optimizer how to optimize better on the fly. So we're moving towards models where the optimization process itself is adaptive, deep, and has its own memory, instead of just being governed by a fixed learning rate schedule. Okay, this focus on memory systems then scales up, leading to what the source is called the continuum memory system, or CMS.
8:50Yes, CMS is basically the architectural embodiment of these NL ideas, generalizing the short-long-term memory distinction using the multi-frequency concept. Formally, you can think of CMS as a chain of MLP blocks. Let's say MLPF1, MLPFK, each block has its own update frequency tied to a specific chunk size. This means the parameters only get updated every step steps. Right, so that enforces the different learning speeds. The large things mean infrequent updates, slow consolidation, behaving like long-term memory. Exactly. And it's interesting, the sources point out that a standard transformer block with its single feed-forward network is just a very simple case of CMS, where you only have one frequency level,$5.01.
9:30NL argues you need Kaler dollars to get that continuous memory consolidation effect. Okay, so building on this multi-frequency memory foundation, the researchers developed a new model called HOPE. It combines this CMS with something self-referential. That's right. HOPE stands for High Order Self-Referential Permutation Encoder. It's designed from the ground up using NL principles. It takes the Continuum Memory System, CMS, and combines it with a self-modifying component, kind of like in the Titans model, allowing it to adjust its own attention mechanism based on the immediate context. So how does HOPE look different from a standard transformer architecturally?
10:07The difference is quite clear if you visualize them. A standard transformer has that single static FFN block. HOPE, reflecting the NL structure, explicitly uses multiple FFNs operating at different frequencies. You might have a low-frequency FFN, a mid-frequency FFN, and a high-frequency FFN, all layered on top of that self-modifying attention part. So the architecture directly mirrors the multi-level, multi-speed optimization idea from NOEL. It's a full commitment to that concept, yeah. And the payoff seems to be there in the results. The authors tested hope against strong baselines Transformer Plus plus MetNet on standard benchmarks like language modeling perplexity and common sense reasoning tasks.
10:46And the results. Did this multi-level memory actually help? It seems so, quite significantly. Looking at their 1.3 billion parameter models, for example, Hope outperformed Transformer++ and the Titans model on average accuracy across several reasoning tasks. Hope got around 57.2 % average accuracy compared to about 52.3 % for Transformer++ and 56.8 % for Titans, which, remember, has self-modification but lacks that deep multi-frequency memory structure. It's a noticeable jump, almost five percentage points over the standard Transformer++ approach at the same scale. What's driving that improvement?
11:22Is it mostly the memory or the self-modifying part? The authors suggest it's really the combination, the synergy between the two. The deep, multi-frequency CMS provides the foundation for stable, continuous learning. And the self-modifying attention allows it to dynamically adapt its processing the keys, values, queries based on the current context or task. So it's learning the task on the fly while also properly consolidating that learning into its structure. Tying this all together then, nested learning looks like much more than just a new way to draw architecture diagrams. It's a mathematical framework aiming to make the internal workings of these complex AI models transparent.
11:58It suggests that if we want AI with more sophisticated in-context learning, maybe even continual learning, we have to move beyond static layers and start designing these hierarchical memory systems defined by their update speeds. So the big takeaway for AI development seems to be we need to cure this amnesia in LLMs. And the path forward might involve building models that mimic the brain's multi-speed learning structure more closely. It's almost a shift from just optimizing the output of the model, like getting the prediction right, to optimizing the model's internal learning process itself, a kind of meta-learning built right into the architecture.
12:31Exactly. And here's a final thought to leave you with. If Enala helps us see even simple optimizers like atom or momentum as forms of associative memory that compress gradient history, what else could this nested optimization perspective reveal? What about higher level human cognitive functions? Things like creativity or abstract reasoning, maybe even managing emotions. Could they also be understood fundamentally as complex nested optimization problems with different memory systems operating at multiple timescales? It certainly gives you something to think about regarding the real depth of learning.
From the publisher
This paper introduces Nested Learning (NL), a new paradigm that addresses fundamental challenges in AI self-improvement, continual learning and memory for models like Large Language Models (LLMs). NL suggests that existing deep learning methods compress their "context flow" and explains how in-context learning emerges in large models. The authors propose the HOPE architecture, a self-referential learning module with a Continuum Memory System (CMS), which is built on the NL insights that traditional optimizers are fundamentally associative memory modules. Experiments demonstrate that HOPE, using this novel framework, shows promising results across language modeling and common-sense reasoning tasks, often outperforming modern recurrent neural networks and Transformers.




