Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

16 Jan 2026 · 14 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Proposes “conditional memory” to complement Mixture-of-Experts (MoE). Instead of using transformer layers to reconstruct static facts, add Enneagram: a fast memory bank for knowledge retrieval, leaving MoE for compositional reasoning.

Guest backgrounds

No guest names or biographies appear in the transcript.

Key claims

Standard transformers lack a native lookup primitive, so early layers waste compute simulating retrieval. Enneagram uses tokenizer compression (23% vocab reduction), multi-head hashing (mitigates collisions), and context-aware gating (suppresses contradictions). Optimal sparse budget allocation is ~20–25% to Enneagram; pure MoE or pure memory is worse (U-shaped scaling). Enneagram improves reasoning via “effective depth” (earlier high-confidence representations).

Notable examples

Diana, Princess of Wales; “Apple” polysemy; benchmark gains: MMLU +3, Big-Bench-Hard +5, Code Generation HumanEval +3; long-context: multi-query NIAH 84.2%→97.0%. System benefit: deterministic indices enable DRAM/SSD offload with ~2.8% throughput penalty.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Inefficiency of Current LLMs

0:42 to 2:30

Understand the inefficiencies in current language models regarding knowledge retrieval.

“We're basically asking, what happens when you stop forcing a calculator to reconstruct every common fact from scratch and just let it remember them instantly?”

Understanding Linguistic Duality

2:30 to 6:08

Learn about linguistic duality and its implications for AI architecture.

“This is the deep, dynamic computation, the heavy lifting.”

Enneagram's Approach to Knowledge Retrieval

6:08 to 9:40

Discover how the Enneagram module addresses knowledge retrieval through innovative techniques.

“So to mitigate that, they use multiple hash heads to query different memory locations at the same time, which just increases the probability of getting a clean hit.”

Impacts of Conditional Memory on Performance

9:40 to 12:12

Examine how conditional memory affects performance in reasoning and memory tasks.

“Enneagram effectively gives the model more sequential depth, more layers to focus entirely on complex global thinking, and it can do it earlier.”

Balancing Memory and Computation

12:12 to 14:00

Find out how to allocate resources between memory and computation for optimal performance.

“The system knows exactly what it needs ahead of time, so it can asynchronously retrieve those embeddings and critically overlap that memory transfer time with the computation that's already happening.”

Exploring Conditional Memory Efficiency

14:00 to 14:18

The discussion revolves around the implications of scaling conditional memory in AI models.

“and mere knowledge retrieval, which can be outsourced?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So if you're tracking the cutting edge of massive AI models, you already know the reigning champion of scaling is a mixture of experts, Moe. It's really the breakthrough architecture that solves this huge parameter budget problem by introducing conditional computation. Right. It lets models get huge without, you know, bankrupting the data centers. It works. It works beautifully for dynamic, complex logic. It is undeniably the current standard. It is. But what if Moe and maybe the whole transformer architecture is forcing our AI to work way too hard on the simplest tasks? I think that's the core question here.

0:36Today, we're doing a deep dive into a pretty radical idea for structural optimization. It's about splitting the model's brain into two distinct parts, a dynamic thinking engine and an ultra-fast static memory bank. Right. We're basically asking, what happens when you stop forcing a calculator to reconstruct every common fact from scratch and just let it remember them instantly? And that gets right to the core inefficiency we're looking at. So our mission in this deep dive is to analyze this concept of conditional memory. It's instantiated by a new module called Enneagram. Yep. And we'll show exactly how Enneagram doesn't compete with Moe, but actually perfectly complements it.

1:15Compliments it. Yeah. We're going to uncover this fundamental flaw in how standard LLMs handle simple information and reveal how this fix leads to massive and honestly pretty counterintuitive performance gains. Not just in facts, right? No, not just in retrieving facts, but crucially in general reasoning, complex coding, and even long context tasks. Okay, let's put that inefficiency into context because it's a great analogy. Yeah. Imagine you're a highly specialized corporate lawyer. Okay. Every single time you need to cite a common legal precedent, I mean a basic clause about contract liability, you are forced to open up the entire legal code, dynamically reread it, and reconstruct the argument from scratch.

1:57Right. You can't just have it memorized. You couldn't just have a bullet point list of common citations ready to go. It's ridiculously expensive in terms of time and cognitive load. Yeah. That's what current LLMs are doing for basic facts. Precisely. And the research behind this pinpoints what they call linguistic duality. Linguistic duality. It identifies that language processing isn't one single operation. It naturally breaks down into two very distinct subtasks that need different kinds of machinery. Okay. So what are the two tasks? The first is compositional reasoning. This is the deep, dynamic computation, the heavy lifting.

2:33The thinking part. Yeah, the sequential, high-level thought you need to solve new problems, like a multi-step math problem or generating some creative text. Makes sense. The second task is knowledge retrieval, and this is totally different. It deals with local, static, stereotyped information. Like what? Named entities, common idioms, formulaic patterns, fixed facts. These are basically lookups. They don't require deep reasoning. They just require instant recall. And here's the architectural paradox, right? The standard transformer models, they don't have a native lookup primitive. There's no dedicated mechanism for that instant recall.

3:10Exactly. So they're forced to simulate that knowledge retrieval through computation. So when the model sees a common phrase like, say, Diana, Princess of Wales, it doesn't just retrieve it from some knowledge bank. No, it has to consume multiple early layers of its attention mechanisms and its feedforward networks. The very layers that are designed for deep thought. Just to piece together that static concept. It is an incredibly expensive runtime reconstruction of something that's basically a static lookup table. You're spending your most valuable resource, that sequential depth, the layers of thought on trivial operations.

3:46It's like using a supercomputer to calculate that 2 times 5 equals 10. Instead of just having that fact available instantly. Exactly. So the first few layers of our very expensive Moe network are just acting as these inefficient lookup tables instead of starting the complex problem solving they were built for. That feels like a huge bottleneck. It is. Now, if we accept that, how does it change how we approach scaling? Because now we're talking about two very different kinds of sparsity. The first one we know well, conditional computation or Moe. Right. Conditional computation is all about selectively activating these specialized teams of experts, the parameters, to handle dynamic logic.

4:27Specialization of thought. That's your team of highly specialized on-call consultants. Expensive, dynamic, and they only activate when the problem is in their niche. A perfect analogy. So the complementary axis is conditional memory, which is what Enneagram brings to the table. This is completely different. It relies on sparse lookup operations to retrieve static embeddings for fixed knowledge. It's not dynamic computation. It's memory retrieval that operates in constant time. So this is a massive pre-index library or an ultra-fast index card catalog. Exactly. It doesn't think. It provides immediate static recall.

5:04But how does Enneagram practically manage that? Because if you're dealing with every possible combination of tokens and grams, that combinatorial space just seems intractable. you'd need an infinite memory table. You would. And that's why the Enneagram module is more than just a simple key value store. It's a really modern adaptation of classic Enneagram embeddings to get that 01 constant time lookup. It uses three clever tricks to get there. All right, what's the first one? First is tokenizer compression. So standard tokenizers often create separate IDs for words that are semantically the same, like apple with a space before it versus apple without.

5:39Oh, right. Enneagram collapses these into single canonical IDs, and this was a huge win, a 23 % reduction in the effective vocabulary size, which makes the memory bank much more semantically dense. Okay, so that's step one. What's two? The second is multi-head hashing. Since we can't build a table for every Enneagram, Enneagram uses deterministic hashing to map a sequence of tokens to a memory index. But simple hashing can have collisions. Right, where different things map to the same spot. Exactly. So to mitigate that, they use multiple hash heads to query different memory locations at the same time, which just increases the probability of getting a clean hit.

6:16So the hashing gives you speed, but multi-head hashing is like a redundancy measure. It's like having three different phone books indexed differently just in case the name is misspelled in the first one. That's a great way to put it. But if you're just pulling static facts out, there's a risk they could be noisy or just wrong for the context, especially with those hash collisions. Does the model just blindly trust the memory? Absolutely not. And that's the third and maybe the most critical piece, context-aware gating. Okay. The memory that's retrieved has to be dynamically modulated by the current hidden state of the transformer.

6:50If the static fact contradicts the global context, say, the classic polysemy of Apple the company versus the fruit of dynamic gate just suppresses that memory. So the gate is like a security checkpoint. The index card gives you a fast fact. but the gate validates it against the actual conversation. Precisely. If I'm talking about a computer and memory suggests red fruit, the gate says, nope, abort. And that prevents the whole reasoning system from getting polluted by fast but wrong information. This brings us to a really critical engineering question, then. The sparsity allocation problem. You have a fixed budget for sparse parameters.

7:27How do you split that capacity between Moe experts for computation and Enneagram memory for knowledge? And the experiments on this revealed a very clear and stable U-shaped scaling law. A U-shape. Yeah. Pure Moe 100 % in computation is suboptimal. It wastes cycles. Pure Enneagram 100 % in memory is even worse. It loses all its reasoning ability. Right. You do both. The optimum performance was achieved by reallocating roughly 20 to 25 % of that total sparse parameter budget to Enneagram. That hybrid point gave the best overall validation loss. That 20 to 25 percent window is fascinating. Is it always in that range?

8:05Or does the sweet spot maybe shift if you train on different kinds of data? It seems pretty stable. It's this delicate balance, right? If you push to, say, 30 or 40 percent memory, the U-curve starts climbing again because you've starved the model of the computational power it needs to actually use the memory correctly. So that 20, 25 percent range seems to satisfy most of the knowledge retrieval needs without kneecapping the model's ability to actually think. Exactly. So we found the perfect mix for our construction crew. You need plenty of dynamic engineers, the MOWE experts, but you dedicate a precise portion of your resources to giving them instantaneously available reference manuals.

8:40That's Enneagram. And that complementarity is what delivers this really surprising performance payoff. We compared a hybrid Engram 27B model against a MOWE 27B baseline. ISO parameter, ISO FLOPs. So the computational cost was identical. And the results? Significant gains in knowledge tasks. Like MMLU, it was up by three points. Okay, that makes sense. But here's the crucial insight, the thing that really validates the architecture. The gains were even larger in domains that rely purely on dynamic thought. Advanced reasoning. Wait, really? Yeah. Big Bench Hard went up by an astounding five points.

9:16Code Generation Human Evil up by three points. Hold on. If you introduce a module for static memory, why did the model get profoundly better at dynamic novel reasoning? That feels like a paradox. Static facts should help fact retrieval, not general intelligence. And that's the aha moment. The answer is what we call effective depth. Effective depth. By relieving the early layers of the transformer from that boring task of knowledge reconstruction, Enneagram effectively gives the model more sequential depth, more layers to focus entirely on complex global thinking, and it can do it earlier. So how do you prove that?

9:50Mechanistically. Researchers use a technique called LogitLens, which tracks how information evolves layer by layer. And they found that the Enneagram models show a systematically smaller KL divergence in their early layers. Oh, translate KL divergence for us. It basically means the model reaches a high confidence, highly structured representation of the input much, much faster. It's like a runner skipping the slow start and immediately being at a full sprint. Because the basic knowledge is just instantly assimilated by the memory bank. Right. And the analysis of the representations themselves is even more compelling.

10:22using something called CKA analysis. Which measures similarity between layers. Exactly. We confirmed the depth advantage. The representations at layer 5 of the Enneagram model align most closely with the feature maps at layer 12 of the baseline MoE model. Layer 5 of the hybrid is functionally equivalent to layer 12 of the pure MoE. So Enneagram is saving the model seven layers of computation just by offloading memorization. Precisely. It's like clearing all the administrative work off a senior executive's desk. They can now dedicate their time entirely to high-level strategy and complex problem solving.

10:57And I bet this extends beyond just reasoning and code. If you free up that attention capacity, you must get huge leverage in long contexts. Absolutely. Delegating local dependencies to instant lookups frees a Tempen to focus on the global dependencies across a long sequence. And that showed up in the benchmarks. Big time. Long context retrieval benchmarks like multi-query NIH saw accuracy jump from 84.2 % to 97.0%. The model can focus on the forest because the memory module is handling the trees. Wow. And finally, we have to look at the system advantage, which is often where MoE models can struggle.

11:33MoE uses dynamic routing, which makes memory management really challenging on accelerators. Because it's unpredictable. You have to keep all those large expert parameters close to the compute unit, which is expensive. Correct. But Enneagram provides this critical architectural relief because its activations use deterministic IDs. The indices are fixed just by the input tokens. Predictability. And that predictability is the game changer. It lets you decouple storage from compute. Which enables the prefetch and overlap strategy. Since the Enneagram tables are huge but fixed, you can offload them to cheaper, abundant, host memory-like standard DRAM or even an SSD.

12:12Exactly. Away from the expensive GPU memory. The system knows exactly what it needs ahead of time, so it can asynchronously retrieve those embeddings and critically overlap that memory transfer time with the computation that's already happening. So by the time the block needs the memory, it's already been loaded. And the performance hit. Negligible. Offloading a massive 100 billion parameter table to host memory had a throughput penalty that peaked at only 2.8%. The cost savings are astronomical, while the latency is virtually zero. It's the ultimate logistics chain. You know exactly which manuals you need, so you send an intern to the archive to get them while you're busy with the intro.

12:49The materials arrive exactly when you need them. Zero downtime. A perfect summary. So the core takeaway here is that the new paradigm for LLMs isn't just about size. It's about optimizing the architecture for purpose. Right. Conditional computation, MOE, is the engine for dynamic logic. Conditional memory, Enneagram, is the engine for static knowledge. Together, they create this optimal hybrid division of labor. And we can actually quantify that division of labor. In a sensitivity analysis, when Enneagram was disabled, performance on factual knowledge tasks just collapsed. It retained only 29 % to 44 % of the original scores.

13:27So Enneagram is clearly the knowledge repository. And ambiguously. But conversely, performance on reading comprehension tasks, things that rely on attention and high-level structure, remained remarkably resilient. How resilient? It retained between 81 % and 93 % of the original scores, even with the entire memory bank switched off. That is the ultimate test of the hypothesis, and it gives us our final thought for you. If a model can perform complex reading comprehension almost perfectly without its explicit memory module, but fails completely at recalling basic facts without it, what does that tell you about the fundamental difference between true reasoning handled by the Moe-e backbone and mere knowledge retrieval, which can be outsourced?

14:08Right. If a 25 % memory budget gives us these games now, what happens when we scale that conditional memory component even further, maybe making the Moe Backbones job even easier?

From the publisher

The researchers introduce Engram, a novel conditional memory module that enhances Large Language Models by integrating a scalable lookup mechanism for static knowledge. While modern models rely on Mixture-of-Experts (MoE) for sparse computation, Engram uses N-gram embeddings to retrieve formulaic or factual information in constant time. This architectural shift creates a U-shaped scaling law that balances neural processing with static memory, allowing the model to offload simple retrieval tasks to early layers. By delegating local patterns to these lookups, the transformer's attention capacity is preserved for complex reasoning and long-context processing. Experiments show that an Engram-augmented 27B model significantly outperforms standard MoE baselines in math, coding, and general reasoning. Furthermore, the system supports offloading massive parameter tables to host memory, ensuring high efficiency with minimal computational overhead.

More from Best AI papers explained

All 475 episodes
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsBest AI papers explained · 14 min
Listen in VO