Deep sequence models tend to memorize geometrically; it is unclear why.

8 Jan 2026 · 13 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Deep sequence models (transformers, Mamba) store knowledge in two competing ways: fast local associative memory (next-token/neighbor lookups) and slower “geometric memory” (embeddings whose dot products approximate multi-hop distances). Geometric memory enables strong generalization on unseen multi-step reasoning, but why it emerges is unclear.

Guest backgrounds

No guests are mentioned in the transcript.

Key claims

Associative memory fails on long-path composition; geometric memory reaches 100% accuracy on unseen 6–10 hop paths on large graphs. Geometry forms in weights (parametric memory) but fails when the graph is provided in-context. Training on both forward and reverse edges (DEDGE) avoids a “reversal curse.” The emergence puzzle is attributed to low-rank spectral bias: gradient descent implicitly aligns embeddings with top Fiedler-like eigenvectors, yielding global structure.

Notable examples

Path-star graphs; dot-product distance vs multi-hop search; Node2Vec embeddings converging to Fiedler vectors; implications for editing/unlearning and for parametric vs prompt memory.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Associative vs. Geometric Memory

1:06 to 2:11

Contrast between associative memory and geometric memory in AI models.

“Let's start with that contrast, because they really are fundamentally different ways of knowing something.”

The Mechanisms of Geometric Memory

2:11 to 3:59

How geometric memory simplifies complex reasoning tasks in AI.

“If the model has never seen Entity of a Law explicitly linked to Entity's Largaller, how on earth does it figure out the relationship between them?”

Research Insights on Memory Structures

3:59 to 4:25

Discussion on experiments revealing strengths of geometric memory.

“This was one of the biggest success cases for this kind of implicit reasoning we've seen.”

Challenges in Proving Geometric Memory

4:25 to 7:17

Challenges faced in proving the advantages of geometric memory over associative memory.

“This amazing reasoning ability only happened when the graph structure was stored in weights, meaning it was baked into the model's parameters during training.”

Emergence of Geometric Memory Structures

7:17 to 10:39

Investigation of how geometric memory emerges during training.

“Maybe the geometric map is just, I don't know, a more mathematically simple way to store the data.”

Broader Implications for AI Knowledge Management

10:39 to 12:26

Exploring the implications of geometric memory on AI knowledge management.

“It points to what they call a visible headroom for practitioners.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, when you ask a modern AI, let's say one of the big language models, a really simple question, something like, what is the capital of France? The common assumption is that it just, you know, pulls that answer instantly, like it's doing a super high speed database lookup. Exactly. And for years, that's pretty much how we all thought about AI knowledge. We called it associative memory. It's basically a giant brute force lookup table. If it sees dualers next of another enough times, the model just memorizes that connection. But the reality, and this is what new research is making so clear, is it's far more sophisticated than that.

0:36we're finding that these deep sequence models, you know, the transformers and mampas that power everything, they don't just use that one simple trick. They actually store information in two completely different ways. And they're often competing with each other. It's like an internal war over how knowledge should be represented. So you have the traditional simple lookup, the associative memory, and then what? And then you have this other, frankly, revolutionary mechanism that we're now calling geometric memory. So today we're going to do a deep dive into that internal competition. Our mission is to figure out why this geometric structure, which is actually harder to learn, seems to emerge at all, and how it transforms these incredibly hard reasoning tasks into something almost simple and what that really means for the future of AI.

1:22Let's start with that contrast, because they really are fundamentally different ways of knowing something. Associative memory is just that. It's local. It's a mechanism that remembers the word Paris is often seen next to the word France in the training data. Okay, so it's a direct link. If you imagine all knowledge as this giant graph, associative memory only cares about the nodes that are right next to each other. Right. It's storing the DeLarder-Next-Dollar relationship. And it's great for just recalling things you've seen a million times. But it's limited. Very. Now, picture geometric memory.

1:54This is a completely different way of thinking. Instead of just memorizing local links, the model creates these rich numerical representations or embeddings that encode the global relationships between everything, even between things that never actually appeared together during training. Whoa, hold on. That's a huge leap. If the model has never seen Entity of a Law explicitly linked to Entity's Largaller, how on earth does it figure out the relationship between them? That is the magic of the geometry. It transforms what would otherwise be a grueling multi-step search problem, say finding a path that takes 10 steps through that knowledge graph.

2:30It turns that into a single, easy-to-learn navigation task in space. So instead of thinking, hey, first I look up A, then I look up B, then C, the model just sort of places A and C on a map and calculates the straight line distance between them. That's a perfect analogy. Exactly. Mathematically, it's using the dot product of their geometric embeddings to approximate that multi-hop distance. It's turning a hard, discrete problem into a simple, continuous one. That sounds incredibly powerful, but I mean, how do you prove that's what's actually happening inside the model? How did the researchers even design an experiment for this?

3:04They needed a task that was basically a trap for the associative approach. So they used something called a path star graph. Imagine a central hub, a root node, with dozens of really long, uniform paths branching out from it. It's a tree. And the task is to predict the end of a long path, say, 10 steps away, just given the start. Right. And in a setup like that, associative memory is just, it's set up to fail. To find a 10-hop path, it needs to do 10 separate lookups. You have to recall step one, then use that to find step two, then step three, and on and on. And that kind of discrete composition is just a nightmare for a neural network.

3:40Computationally, we're talking exponentially hard. It's like searching for a needle in a haystack. But at every step, the haystack gets bigger. But the geometric memory, it just sidesteps the haystack. Completely. It approximates the entire 10-step journey as just one smooth navigation task between the start and end points on its internal map. And did it work? It worked spectacularly. This was one of the biggest success cases for this kind of implicit reasoning we've seen. Transformer and Mamba models got up to 100 % accuracy on paths they'd never seen before. Paths that were 6, even 10 hops long.

4:12On massive graphs, too. So they were solving these really complex reasoning problems just by, what, navigating their own internal geometry? Yes. And that ability was directly tied to where the knowledge was stored. This is the crucial part. This amazing reasoning ability only happened when the graph structure was stored in weights, meaning it was baked into the model's parameters during training. So what happened when they tried the other way? What if they gave the models the exact same task, but just pasted the knowledge, the whole graph, in context, right into the prompt? Total failure. The models failed completely.

4:47And that's the smoking gun, right? It proves the problem isn't the task. It's the storage mechanism. Parametric memory in weights memory allows this powerful geometry to form. In context, memory is just a temporary scratch pad. It doesn't work for this. There was one other little detail that I found fascinating. For this to work, the models had to be trained on both the forward links and the reverse links. That's right. They called it DEDGE. It was the trick to avoid something known as the reversal curse. If you only show a model that A leads to B, it often has no idea that this implies a structural path back from B to A.

5:22Training on both directions lets it build a stable, coherent map. Okay, so let's get to the biggest puzzle here. We've established that geometric memory is just vastly superior for this kind of complex reasoning. But here's the thing that's so counterintuitive. It turns out that the simple brute force lookup table, the associated memory, is actually faster for the model to learn. It completely defies our intuition. Gradient descent, the algorithm that trains the AI, it can find that simple associative structure incredibly fast, sometimes in just two steps on a small graph. It's the path of least resistance.

5:56And yet over time, after hundreds of steps, the model actively rejects that easy solution. It reorganizes its knowledge into the geometric form, which takes way longer to build. Why? Why would the optimizer choose the slow, hard path? Okay, so this is where we put on our detective hats and just eliminate the obvious suspects one by one. Suspect number one, the model only builds the geometry because it's being forced to. You mean you wouldn't build a map unless someone asked you to find a long-distance road? Precisely. That was the hypothesis. But the research completely refutes it. Yeah. This geometric structure emerged even when the models were trained only on local supervision.

6:35All they did was ask it to memorize simple one-step edges. There was zero pressure to reason or find bigger structures. The map just, it built itself. Organically. That's actually stunning. The model is anticipating a need it hasn't even been tested on. Okay, so if it's not explicit pressure, then what about a capacity issue? Is it forced into geometry because it's just running out of space for this simple lookup table? Another good suspect, the capacity constraint argument. But that's also refuted. Associative memory is actually pretty easy to represent, and the researchers used really large embedding dimensions.

7:10There were no bottlenecks. The model had plenty of room for the easy lookup tables. But it still chose the geometric route. It still chose the geometric route just more slowly. Okay, so it's not a space restriction. What about elegance, simplicity? Maybe the geometric map is just, I don't know, a more mathematically simple way to store the data. That's a great thought, and it's a huge idea in generalization theory. But again, counterintuitively, no. For the simple graphs they used, the mathematical complexity of the two storage methods, whether you measure it by bit complexity or L2 norms, was about the same.

7:44In some cases, the associative one was actually a tiny bit simpler. Wow. Okay, so we've ruled out everything. The model is choosing a memory structure that is slower to find, isn't required by the task, isn't forced by the architecture, and isn't even simpler. The optimization process itself is actively fighting over time to create this structure. The puzzle holds up. The emergence of this geometry. It can't be explained by our usual ideas about efficiency or generalization. It suggests the optimization process itself has this deep, almost hidden preference for structure. So if it's not external pressure or internal efficiency, where on earth does this map actually come from?

8:24This is the core inside of the work. The geometry comes from something called a low-rank spectral bias, and it just arises naturally from the way these models are trained. Okay, low-rank spectral bias. You're going to have to translate that for us because that sounds straight out of a textbook. It does, I know. But let's use an analogy. Think of all the raw training data, all those local A to B connections, as a really noisy radio broadcast. It's got hundreds of different signals, different frequencies, all mixed up, and each frequency corresponds to some part of the graph's structure. And the model is trying to tune in to the right station.

8:58Exactly. The training process, guided by the loss function, it acts like a tuning dial, and it naturally, implicitly, favors certain frequencies. Specifically, the model's learned embeddings start to align with what are called the top Fiedler-like eigenvectors of the graph. Okay, Fiedler vectors, graph laplation. For those of us who haven't picked up a graph theory textbook this week, what do these vectors actually do? They're the good stuff. The Fiedler vectors are the mathematical directions that encode the global structure of the graph. They're the blueprints. If the graph is your world map, the Fiedler vectors are the underlying lines of latitude and longitude that define the whole coordinate system.

9:38So the spectral bias is just the natural tendency of the training process to filter out all the noise, all the local irrelevant stuff, and tune in directly to the structural blueprints. Precisely. The geometry is the model performing what's called a low-rank factorization of the whole adjacency matrix. It's softly filtering out the less important eigenvectors, the noise, and just keeping the top ones that define the global structure. This creates that clean internal map. It's an implicit preference for global structure that just stabilizes over time. So the AI isn't just storing facts. It's performing this incredibly advanced mathematical compression to prioritize global structure even when it's harder to do.

10:18And this holds up. Researchers saw the same dynamic in simpler models, like the old Node2Vec algorithm. The embeddings naturally converge to span these Seedler vectors. It's just a side effect of the optimization function itself. What's really wild is that those simpler Node2Vec models often ended up with more strongly geometric embeddings than the big, powerful transformers. And that's a huge tell. It points to what they call a visible headroom for practitioners. It means our current transformer models are probably stuck in a kind of messy, suboptimal middle ground, a mix of associative and geometric storage.

10:51But now that we understand this bias, you could actually design train methods that deliberately push the models to become more geometric. Yes. That could be a huge leap for implicit reasoning in things like natural language processing, where you're constantly having to synthesize connections between disparate facts. So let's zoom out. What are the broader implications here? Well, first, this isn't some quirk of transformers. They saw similar results in Mamba models and on harder graph types. This seems to be a fundamental property of how these deep sequence models learn. It has to change how we think about managing knowledge in AI.

11:27Completely. It reshapes everything from knowledge acquisition to a really big one. Unlearning. Think about it. If a fact is stored associatively, you just delete the entry in the lookup table. Easy. But if that fact is part of a geometric map... It's connected to everything else. It's like trying to remove one small town from a globe. You can't just erase the dot. You'd mess up all the roads, the borders, the distances to dozens of other places. Exactly. The interdependencies make editing or forgetting a specific fact much, much harder. It also seems to be a huge vote of confidence for parametric memory for storing knowledge in the weights over just feeding it in the prompt.

12:06It's very strong evidence for that. The geometry that enables this powerful reasoning just seems to form best when knowledge is learned deeply, parametrically over time. Which brings us all back to you, the listener. The next time you ask an AI a really complex, multi-step question, something that requires it to connect a few different ideas, its success isn't coming from just, you know, checking facts off a giant list. Not at all. It's succeeding because it is navigating an internal, complex, geometric world model that it built all by itself. just from the patterns in its training data. The geometry is what lets it synthesize new connections.

12:43So what this all means is we are really shifting our view of AI memory. We're moving away from the idea of a static list of facts, the associative memory, to this dynamic internal spatial map, the geometric memory, that can solve problems in linear time that should be exponentially complex. And the final really provocative thought to leave you with is this. Given that this incredibly sophisticated geometric structure arises on its own, even when it's harder and slower for the model to find and isn't even explicitly required by the task. It suggests that gradient descent, the very heart of machine learning, is optimizing for something much deeper than just accuracy or efficiency.

13:19There seems to be an inherent, fundamental, and still somewhat mysterious preference for abstract global structure that we are only just beginning to understand.

From the publisher

This research introduces the concept of geometric memory to explain how deep sequence models store and reason over atomic facts. Unlike traditional associative memory, which functions as a simple lookup table for co-occurring entities, geometric memory synthesizes global relationships that enable models to solve complex multi-hop reasoning tasks. The authors demonstrate that models can learn to navigate large, unseen graphs by organizing node embeddings into a spatial geometry that reflects the graph's overall structure. Surprisingly, this geometric bias emerges even without specific architectural pressures, capacity limits, or reasoning-based supervision. By comparing Transformers to Node2Vec, the study reveals a spectral bias that naturally directs models toward these powerful, structured representations. Ultimately, these findings challenge the intuition that parametric memory is strictly local, suggesting new ways to improve implicit reasoning and knowledge discovery in language models.

More from Best AI papers explained

All 475 episodes
Deep sequence models tend to memorize geometrically; it is unclear why.Best AI papers explained · 13 min
Listen in VO