In short
“Inert map” problem in language models—models can encode a representation learned from in-context data but fail to use it for rule-based navigation and adaptive world modeling.
Guest backgrounds
No guest names or bios appear in the transcript; it’s a two-host discussion.
Key claims
(1) Linear probing shows models reconstruct a 4x4 grid structure from a random walk in latent space (low Dirichlet energy). (2) Despite this, prompting style matters: instruction-based “predict the next word” fails, while pre-filled generation succeeds. (3) For a two-step jump navigation rule, models fail near random chance; reasoning modules can even hurt (ablation: disabling chain-of-thought improved Gemini Flash on the grid task). (4) Explicitly providing the grid lets models navigate, so the failure is bridging implicit learned maps to explicit tasks.
Notable examples
Toy/Ink/City checkerboard; “predict next word” vs pre-filled cursor continuation; “move exactly two steps south” (expected Jam); Dirichlet energy spiking when reasoning begins.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Inert Map Concept
1:19 to 2:16
Discussion on AI's ability to learn but struggle with practical application of that knowledge, leading to the concept of the inert map.
“And honestly, this one, it kind of broke my brain a little bit.”
Understanding AI's Internal Representation
2:16 to 3:29
Explains how AI models process information within a complex mathematical structure known as latent space.
“Yeah, and I mean, metaphorically speaking, the researchers involved in this deep dive use a technique called linear probing.”
Building a Custom Environment for AI Testing
3:29 to 4:21
Describes a unique experimental setup where AI learns from a custom, blind random walk through a grid without visual guidance.
“There you give every single square a completely random name.”
AI's Performance in Understanding Spatial Relations
4:21 to 5:41
Examines how AI models create a mental map from a random walk and the implications of this ability.
“If you, a human listener, were forced to listen to that list for an hour, what would happen?”
The Navigation Challenge for AI
5:41 to 7:17
Details the challenge of having AI not just learn a map but effectively use it to navigate and respond accurately.
“If the AI's map was messy, the energy score would be super high.”
Instruction vs. Pre-Filled Prompts
7:17 to 8:11
Discusses how different prompting techniques influence AI's ability to utilize its internal knowledge, leading to varying success rates.
“However, when they used a pre-filled prompt, it worked flawlessly.”
The Frustration of Inert Knowledge
8:11 to 10:36
Explores the paradox of AI having knowledge yet failing to apply it effectively, showing limitations in reasoning capabilities.
“It's tied exclusively to the specific flow of text generation.”
Reasoning Models and Their Limitations
10:36 to 13:48
Analyzes reasoning models in AI and their struggles with navigation tasks, even with additional processing time.
“That is a very specific and frankly very frustrating bottleneck for developers.”
Word Models vs. World Models
13:48 to 14:01
Clarifies the distinction between AI's ability to predict text versus its understanding of physical and logical systems, highlighting the inert map problem.
“It's not a clean pipe connecting the two.”
Understanding Word and World Models
14:01 to 16:43
Explore the difference between word models and world models in AI.
“They just prompted it with, tell me what the grid looks like.”
Show all 12 chapters
The Reality of AGI and Model Limitations
16:43 to 17:20
Discuss the challenges and limitations faced by AI in practical applications.
“So does this mean the whole AGI is here party is officially over?”
Philosophical Implications of AI Knowledge
17:20 to 18:06
Examine the philosophical questions surrounding knowledge in AI systems.
“If the information is inside the machine, like it is mathematically proven to be right there in the vectors, but the machine cannot act on it or express it, is it actually knowledge?”
Transcript
Automatic transcript. May contain errors.0:00Let's play a little game to start us off. imagine you hire a new personal assistant okay and they are an absolute genius right like roads scholar level photographic memory of the works so you hand them the employee handbook you give them the floor plan of your office building and i don't know a stack of your last five years of emails sounds like a dream hire so far i'm gonna take one it does just so they sit there and they They read absolutely everything in about 30 seconds and you can quiz them on it. You ask what's on page 42 and they quote it back verbatim. Wow. You ask, where's the break room?
0:37And they point right to it on the map. OK, so they definitely know their stuff. Well, that is what you think. Right. Because then you say, great, I need a coffee. Go to the break room and get me one. And they stand up. They walk out of your office and they immediately walk straight into a wall. Wait, they walk into a wall? Literally. Yeah. Into a wall. They have the map in their head. They memorized it perfectly. but they cannot figure out how to physically move from point A to point B within the actual physical building. They have the knowledge, but it's completely frozen. That is a highly terrifying image.
1:07But I know exactly where you're going with this. You're describing the current state of artificial intelligence. I am indeed. We are doing a deep dive today into a mystery that goes by the name of the inert map. And honestly, this one, it kind of broke my brain a little bit. Yeah. Because we all just assume that since AI is smart and since it can answer our questions, it actually understands the world. Right. That is the holy grail assumption in the industry. It would just assume adaptability. Like if I drop an AI into a totally new environment, say a new code base or a complex new legal framework, and I let it read the rules, it should just be able to operate within that system.
1:45But the research we are unpacking today suggests there is this massive invisible wall between learning a map and actually using a map. It turns out our AI models might be building perfect mental images of the world, and yet they are functionally blind when we ask them to navigate them. It's basically the difference between having a map folded up in your backpack and actually taking it out to see where you are. And today we're going to figure out why the zipper on that backpack seems to be completely stuck. So let's set the stage here for everyone listening. How do we even know what an AI is thinking?
2:16Because you can't exactly slide a large language model into an MRI machine. Well, actually, you kind of can. Yeah, and I mean, metaphorically speaking, the researchers involved in this deep dive use a technique called linear probing. Which sounds like an alien abduction term, but go on. It really does. But think of it this way. An AI model, like the ones powering chat GPT or Claude, processes information in what we call latent space. It's this massive, multidimensional mathematical cloud. Every single concept, every word, turns into a set of numbers, a vector. Okay, so the word apple is a bunch of numbers, and the word orange is a bunch of numbers, and usually those numbers are grouped close to each other in the cloud because they're both fruits.
2:58Exactly. So if we want to know if the AI actually understands a new concept, we can just look at those numbers. We can see if the actual geometry of that mathematical cloud changes. Got it. So to test this inert map theory, they didn't just use Shakespeare or Wikipedia data. They built a custom city of words. Right, because if you use real-world data, the AI might just be cheating. I mean, it already knows Paris is in France. So they had to create a fake world. Imagine a grid. Let's keep it super simple. A 4x4 checkerboard. 16 squares. Okay, I'm visualizing a checkerboard. There you give every single square a completely random name.
3:36One square is Toy. The square directly north of it is Ink. The square to the east is City. Toy, Ink, City. Okay. Making sense so far. But here's the catch. The AI doesn't get to see the grid, doesn't get a nice JPEG image, it doesn't even get a text description saying, hey, this is a 4x4 grid. So it's basically flying blind. Totally blind. All it gets is what they call a random walk. Which is just a list, right? It's a stream of consciousness journey through this invisible city. Imagine you are blindfolded and someone is just dragging you around the checkerboard, shouting out the names of the spots you land on.
4:09Toy Ink City, Wing City Ink Toy. Okay, I get it. It's just a raw sequence. Toy, Ink, City, Wing Cat, Jam, just over and over again. Exactly. Thousands and thousands of steps. Now, here is the crucial part. If you, a human listener, were forced to listen to that list for an hour, what would happen? I mean, I'd start to notice the patterns. I'd realize, hey, every time we leave Toy, we almost always hit Ink or City. But we never, ever go straight from Toy to Jam. Exactly. You would logically deduce the adjacency. you would start to build a mental map in your head, you'd realize toy and ink are neighbors.
4:46So the big question is, did the AI do that? This is where that alien abduction linear probing comes in. They fed this random walk into some standard open weights models. Think of your llamas, your olmos, the foundational models that run a lot of today's tech. Then they essentially paused the model and looked at the raw math inside. And did the numbers actually move? It was beautiful. The internal mathematical representation of the word toy had physically shifted in that vector space to sit right next to ink. The AI had literally reconstructed the two-dimensional grid structure inside its brain purely from reading that random list of words.
5:21That is wild to me. It didn't know it was a grid. It didn't know it was a map. But the math just organized itself into a map anyway. They even used a specific mathematical metric to prove it called Dirichlet energy. Which I am sure is a massive hit at cocktail parties. Oh, huge hit. Everyone loves talking about derricklet energy, but essentially it measures how clumped or smooth data is. If the AI's map was messy, the energy score would be super high. But the energy here was incredibly low. The map was completely solid. The AI knew the city. Okay, wait, stop. This sounds like a massive success story.
5:56We dropped the AI in a totally new environment, gave it the random walk, and it learned the layout perfectly. Why are we calling this deep dive the inert map? Because the resurfers didn't just want to know if the map existed in the map, they needed to know if the AI could actually use it. Ah, the navigation test. Right. Because it is one thing to subconsciously know the layout of a building, it's a completely different thing to actually give someone directions. So they gave the AI a basic task. They said, okay, you've seen the walk, now predict the next word. Seems incredibly easy. If I'm standing at city and the map says wing is next door, I just say wing.
6:32And here lies the very first paradox. It turns out the AI's ability to answer that seemingly simple question depended entirely on how you asked it. Is this the instruction versus pre-filled thing? Because honestly, I was looking at that part of the research and I'm still struggling to wrap my head around it. It is a bit of a mind bender. So there are generally two distinct ways to prompt an AI. The first is instruction. You say to the bot directly, Here is a list of words representing a walk toy ink city. Now I want you to predict the next word. That sounds like every standard prompt I've ever written.
7:07Here's the data. Now do the work. And it failed. Wait, it failed. But the map is right there. We literally just saw it in the math. It failed significantly. It could not leverage the map it had just mathematically built to answer the user's specific question. However, when they used a pre-filled prompt, it worked flawlessly. Okay, let's role play this for a second. I want to make sure you're listening or getting this too. What exactly does a pre-filled prompt look like? In a pre-filled prompt, you don't ask the AI a question at all. You trick the AI into thinking it has already started speaking.
7:37You feed in the text saying the walk continues to Yank City, and then you just leave a blinking cursor. And then it just keeps going on its own. Yes. If it thinks it is generating the list, it accesses the internal map perfectly. It predicts wing. But if you step in and give it the list and demand an answer, it plays totally dumb. That is bizarre. It's like if I ask my genius assistant, where's the coffee? They stare at me completely blankly. But if they're just muttering to themselves, they say, I'm going to get coffee and then walk straight there. That is a phenomenal analogy. It suggests the knowledge is incredibly fragile.
8:11It's tied exclusively to the specific flow of text generation. And it isn't store as a solid fact that you can query on demand. But wait, it gets worse, right? Because simply predicting the next word is basically baby steps. The real ultimate goal of AI is what they call adaptive world modeling. AWM. Yeah, this is the big leagues. This asks the question, can you take the map you learned and actively apply a new rule to it? Give me an example of the rule. So the AI has read the random walk. It has that grid locked in its head. Now the researchers step in and say, we are playing a game. The rule is a two-step jump.
8:48From your current position, move exactly two steps south. Where are you? Okay, let's trace this mentally. I'm at Citi, I know Wing is south, and Jam is south of Wing. So two steps down, the answer has to be Jam. Right, it's easy for you. You visualize the grid, you count two steps, and you're there. And the AI. The OpenWeights models, those exact same models that built that beautiful low-energy map, failed completely. Completely, like not even close. It was near random chance. They could not solve the navigation puzzle to save their lives. I just want to grab the model by the shoulders and scream, it's in your head, you know exactly where Jam is, just go two steps down.
9:20And that frustration is the exact definition of the inert map. The representation is totally inert, it sits there safely encoded in the latent space, but the actual mechanism that does the logical reasoning, the part that processes the verbal rule, move two steps, cannot access that map whatsoever. It's like the left brain simply isn't talking to the right brain. Precisely. The bridge is completely out. Now, I want to play devil's advocate here for a second. Are we absolutely sure the AI isn't just bad at math? Maybe the concept of moving two steps is just too complicated for it. A very fair question, and the researchers definitely thought of that.
9:57So they ran a control experiment. They took away the random walk entirely. They didn't make the AI magically learn the map from a list. Instead, they just explicitly wrote out the grid. So they basically handed it the cheat sheet. They wrote out, this is a 4x4 grid. City is at coordinate 0, 0. Wing is directly south of city. And how did it do? It solved it easily. Unbelievable. So if you explicitly tell it the map, it can navigate fine. If it learns the map implicitly through the walk, it can't navigate at all. Exactly. It's that specific combination that causes the failure. It fundamentally cannot bridge the gap between an implicit map learned from experience and an explicit task driven by a rule.
10:36That is a very specific and frankly very frustrating bottleneck for developers. It really is. And here is the big so what for the researchers. They dug into the brain of the AI while it was actually trying to solve this specific puzzle. They really wanted to see what happens to that beautiful grid structure when the AI actively tries to think about the rule. Like doing a real-time fMRI scan while someone takes a math test. Exactly. And they found that the exact moment the AI switched modes from passively reading to actively reasoning the map just degraded. Wait, it dissolved. It got completely fuzzy.
11:12The Dirichlet energy spiked way up. It seems that the sheer cognitive load of trying to parse the new instruction actually overwrites or totally obscures the mental image it just spent all that time building. So it literally can't walk and chew gum at the same time. It cannot hold the implicit context in its working memory while simultaneously processing a brand new logical rule. It just drops the map entirely to read the instructions. Okay, so that covers the open weights model. But I know what everyone listening right now is thinking. They're saying, I don't use llama. I use the big guys. I use the massive models that think.
11:44Ah, yes, the reasoning models. Yes, the ones that do chain of thought. The ones that write out. Let me think about this step by step. Surely, if the AI takes a minute to scribble down its thoughts on a digital notepad, it can figure out the map right. Enter the giants. The researchers tested the absolute state-of-the-art reasoning models. We're talking the Gemini series, the GPT-5 class models. Did they save the day? Well, on a simple one-dimensional line, just a straight list of stops, the reasoning models did okay. But as soon as they moved back to that two-dimensional grid, the 4x4 squares performance completely collapsed again.
12:27Even with the extra thinking time? Even with the thinking time. And honestly, this is probably my favorite part of the entire deep dive. They did what's called an ablation study. And an ablation study is where they deliberately break part of the model to see what happens, right? Right. They took a model like Gemini Flash, which is specifically designed to think before it answers, and they forcefully stopped it from thinking. They completely turned off the chain of thought. They just told it to guess. Don't overthink it. Just give me the answer. And its performance actually improved. You are kidding me.
12:55I am completely serious. For Gemini Flash, turning off the reasoning module made it measurably better at the navigation task. That is hilarious. It's exactly like when you're taking a multiple choice test and your gut tells you the right answer, but then you think, wait, that seems too easy. And you actively talk yourself into bubbling in the wrong answer. That is exactly what is happening under the hood. The reasoning chain was literally hallucinating. It would write out, I am moving south, so I must be at ink. But ink is north. The reasoning engine was actively making up a path that directly contradicted the perfect map it already held in its own head.
13:31So the gut instinct, the latent map, was right. But the conscious brain, the reasoning part, was completely wrong. In that specific model, yes. Now, to be totally fair, in the much bigger model Gemini Pro, the reasoning did help a little bit. But it still wasn't a silver bullet. What it really shows us is that we don't fully understand how the reasoning chain actually accesses the underlying knowledge. It's not a clean pipe connecting the two. Didn't the researchers even try just asking the models to describe the map in plain text? They did. They just prompted it with, tell me what the grid looks like.
14:04And they very often just couldn't do it. That is so eerie. It's like the AI is saying, I know the way, but I literally cannot tell you the way. It really highlights a fundamental distinction that I think is the biggest takeaway for us today. We have built incredibly powerful word models, but we are still really struggling to build functional world models. Let's unpack that distinction for a second because it's super important. Word model versus world model. Right. So a word model is essentially a massive statistical engine. It knows that the word royal is highly likely to be followed by the word family.
14:38It just predicts text very, very well. But a world model understands the underlying mechanics of a physical or logical system. It understands cause and effect. It gets spatial geometry, object permanence. And this whole inert map problem is basically the glaring symptom of that gap. It really is. We are seeing real proof that just because a model can flawlessly predict the next word in a sequence, like our random walk, doesn't actually mean it has built a functional world model that it can navigate. So taking a step back, why does this matter to the listener? Why should a CTO who is driving to work right now care about Durek Energy and Inert Maps?
15:17Primarily because of the current massive hype cycle around agents. Ah, yes. The mythical AI agent. The autonomous thing that's going to run my entire life and do my taxes. Exactly. We are actively being sold a vision right now where you can drop an AI into a totally new context. your specific business, your weird code base, your messy daily schedule, and you expect it to just learn the ropes simply by reading your documents. That's in-context learning. And then you expect it to autonomously solve complex problems. Right. The pitch is, here is the company handbook. Now go negotiate this vendor contract for me.
15:51Exactly. But that requires two distinct steps. Step one is build the map, read the handbook. Step two is navigate the map, actually apply those rules to the new contract negotiation. And what we're saying today is that step two is completely broken. It is at the very least highly unreliable. The AI might memorize your company handbook perfectly, but when a weird edge case situation comes up that requires combining three different rules from three different pages, the AI might just freeze. It might have the inert map of your company sitting right there in its latent space, but be completely unable to actually use it.
16:25It's the difference between hiring a lawyer who memorized the entire law library and hiring a lawyer who can actually stand up in court and argue a complex case. Ideally, you want both. But right now, the tech industry has produced a lot of students with photographic memories who completely panic the second you ask them a practical applied question. So does this mean the whole AGI is here party is officially over? I don't think it's over at all. But I do think the music has been turned down a little bit. This research really grounds the hype in cold, hard engineering reality. It shows us that simply making the models bigger, throwing more parameters and more data at them, might not actually solve this specific problem.
17:04Because even the absolute giants failed the grid test. Exactly. This strongly suggests we might need a totally different architecture moving forward. We need to figure out how to wire the backpack to the hands. How do we make the latent space reliably accessible to the reasoning engine? You know, sitting here thinking about it, it really makes you question the fundamental definition of knowledge. How so? Well, think about it. If the information is inside the machine, like it is mathematically proven to be right there in the vectors, but the machine cannot act on it or express it, is it actually knowledge?
17:38That is a very deep philosophical cut for a Tuesday. I try my best. Yeah. But seriously, are we just looking at a digital mirror? Like we see our own human data reflected back at us, organized incredibly neatly by the math, and we naturally trick ourselves into thinking there is a real mind behind it that understands. But maybe it is just a really complex reflection. That is the ultimate danger of anthropomorphism. We see the map, so we naturally assume there's a navigator. But right now what we actually have is a map without a hiker. A map without a hiker. That is a fantastic image to end on.
18:10Because until we solve the inert map problem, our AI assistants are effectively incredibly well-read, genius-level intellects who are just stuck standing in the break room staring blankly at the coffee machine, completely unable to find your office. Well, on that highly cheerful note, hopefully you, the listener, can find your way to your next meeting without needing a chain of thought prompt. We can certainly hope. Thanks for joining us for this deep dive. We will see you next time.
From the publisher
This paper investigates whether large language models (LLMs) can effectively utilize novel information they learn during a single conversation. While models successfully create internal representations of new concepts from context, the study reveals these representations are often inert and cannot be applied to downstream tasks. Experiments on next-token prediction and adaptive world modeling show that both open-weights and closed-source reasoning models struggle to use this in-context knowledge reliably. Even when a model's latent states reflect a new logical structure, it frequently fails to deploy that structure to solve problems. These findings suggest a significant gap between encoding information and the flexible deployment required for truly adaptable artificial intelligence. Overall, the authors conclude that current architectures require new training methods to move beyond simple data recognition toward functional in-context reasoning.




