Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs

12 Sep 2026 · 21 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that LLM “thinking out loud” (reasoning traces) improves recall of factual trivia by two mechanisms: a computational buffer (stalling via filler tokens to force extra forward passes) and factual priming (generating related true facts to shift attention and retrieve the target fact). It also claims hallucinated intermediate facts sabotage retrieval, and that filtering reasoning traces can improve accuracy.

Guest backgrounds

No guests are named or introduced in the transcript.

Key claims

Reasoning boosts even single-hop factual questions; pass@K suggests facts are in weights; dummy “let me think” traces can recover Mary Engle Pennington’s induction year (2018) up to a limit (~2,048 tokens) before noise degrades performance; factual priming is crucial (10th King of Nepal example); audit with Gemini 2.5 Flash finds accuracy drops from 71.1% (clean traces) to 32.2% (any hallucinated intermediate fact).

Notable examples

Keanu Reeves/Matrix analogy; Mary Engle Pennington (2018 vs wrong 2019); 10th King of Nepal (Burendra Birbhut Bikram Shah after listing first nine kings); test-time selection filtering verified factual traces yields +12.2% relative accuracy.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Memory Recall

0:46 to 1:39

Exploring how AI models retrieve information similarly to human memory.

“We are looking at how AI uses reasoning or, you know, basically thinking step by step out loud to unlock factual knowledge that is just deeply compressed inside the parameters.”

The Anomaly of AI Reasoning

1:39 to 2:11

Discussing why AI needs reasoning for simple questions despite seeming unnecessary.

“It makes logical sense that a model needs to break a complex computational PASC into smaller sequential steps before arriving at a final output.”

Insights from Pass at K Metric

2:11 to 4:05

Introducing the Pass at K metric and its implications for AI knowledge retrieval.

“Questions that require absolutely no deduction whatsoever.”

The Role of Model Size in Reasoning

4:05 to 4:44

Examining how smaller AI models benefit more from reasoning techniques.

“It's actually the smaller ones, like QUIN332b, that see the most dramatic improvement.”

The Computational Buffer Mechanism

4:44 to 5:03

Introducing the concept of a computational buffer in AI reasoning.

“Yeah, the reasoning trace provides the extra computational leverage they need to pull it to the surface.”

Case Study: Mary Engle Pennington

5:03 to 8:14

A case study illustrating how reasoning impacts AI's factual recall.

“Like, is it genuinely thinking about the capital of Iceland, or is the architecture just stalling for time?”

Filler Tokens and Retrieval Accuracy

8:14 to 9:29

Discussing how generating filler tokens influences AI's information retrieval.

“It bubbles up and becomes the most likely next token.”

Factual Priming and Semantic Bridging

9:29 to 12:42

Exploring the concepts of factual priming and how AI builds connections.

“The dummy text never completely recovers the full accuracy of a natural reasoning trace, where the AI is allowed to actually talk about the subject.”

The Fragility of AI Memory

12:42 to 14:00

Analyzing the vulnerability of AI to inaccuracies during reasoning processes.

“It proves that the factual content itself, the semantic manipulation of the vector space, is heavily responsible for the correct recall.”

Understanding AI's Memory Fragility

14:00 to 16:54

Explore how hallucinated facts affect AI's memory and accuracy.

“Okay, 71.1 % is a highly capable retrieval rate.”
Show all 13 chapters

Harnessing AI Mechanics for Better Accuracy

16:54 to 18:26

Learn about implementing test time selection strategies to improve AI performance.

“Which brings us to the practical application.”

Practical Prompt Engineering for LLMs

18:26 to 19:36

Understand the importance of prompting AI for factual context in queries.

“think there is a highly practical prompt engineering takeaway here for our daily workflows.”

Redefining AI Knowledge and Recall

19:36 to 20:34

Discuss how AI knowledge differs from static databases and its implications.

“So to synthesize what we've unpacked today, we basically looked at two major mechanistic discoveries about how AI reasoning actually functions.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So you're at a party, right? And someone mentions an actor. Oh, I know where this is going. Yeah. You know exactly who I mean. You can picture his face. You know, he has like a dog in that one huge action movie from 2018. But his name is just, it is locked in a vault somewhere deep in your brain. Right. Totally inaccessible. Exactly. So what do you do? I mean, you don't just stand there in complete silence. You start, you know, talking out loud. Just kind of rambling. Yeah. Yeah. Circling the target with random trivia. Saying things like, he was in The Matrix, right? We're in speed until boom, the name Keanu Reeves just suddenly drops into your lap.

0:36It's actually a very familiar cognitive process. You know, you are actively building a bridge to that memory by activating related concepts in your mind. Well, today we are doing a deep dive into some absolutely wild data that proves large language models, AI models, actually do the exact same thing. Which is fascinating, honestly. It really is. We are looking at how AI uses reasoning or, you know, basically thinking step by step out loud to unlock factual knowledge that is just deeply compressed inside the parameters. Right. Pulling it to the surface. Yeah. So the mission of this deep dive is to answer a really puzzling question for you guys today.

1:15Why does an AI need to think out loud to answer a simple trivia question that requires zero logic? OK, let's unpack this. So to really appreciate why this behavior is so anomalous, we kind of have to look at the standard operating procedure for inference compute. Right. How it normally works. Yeah. Typically, we prompt an AI to use chain of thought reasoning for complex multi-step problems. So think about like advanced algebra or debugging a Python script. That makes sense. Right. It makes logical sense that a model needs to break a complex computational PASC into smaller sequential steps before arriving at a final output.

1:54Sure. I mean, it's breaking down the problem to manage its token budget and maintain logical consistency. That's just standard. It is. But recent experiments reveal this massive spike in accuracy when models deploy this step-by-step reasoning for simple, single-hop, factual questions. Wait, really? Yeah. Questions that require absolutely no deduction whatsoever. But wait, if it's a single hop trivia question, what is there to even compute? Like if I ask you, what is the capital of Iceland? You don't need to do any math. Exactly. No math. Right. You don't deduce that Reykjavik is the capital through a series of logical steps.

2:30You either have that fact stored in your weights or you don't. So why does the AI need to, you know, burn inference compute when there is literally no logic to process? That skepticism right there is the perfect entry point because researchers wondered the same thing. To figure out if the knowledge was actually in the model's weights at all, they used a metric called pass at K. Pass at K. Yeah. Basically, if you let the AI sample K, different answers. So say 100 independent attempts. What is the probability that at least one of those generations contains the correct answer? Oh, I see. It's like losing your keys.

3:05How do you mean? Like if you lost them inside your house, searching 100 times means you'll eventually find them. right? The probability is there. Right. But if you drop them down a storm drain across town, you could search your house a thousand times and you'd never find them. That is a perfect analogy. And the pass at K curves in the data demonstrate that the knowledge is, in fact, inside the house. Okay. Wow. Yeah. Reasoning doesn't just help the AI format or like rank its answers better. It physically expands the model's capability boundary. It pulls out the keys. Exactly. It allows the model to surface this highly compressed knowledge that a standard single pass generation simply cannot reach.

3:48Okay, that makes sense. So if you turn reasoning off, the AI might guess 100 times and never output the capital of Iceland. But if you force it to reason first, it suddenly fetches the correct answer. And I noticed in the research that the models getting the biggest boost from this aren't the massive frontier models. It's actually the smaller ones, like QUIN332b, that see the most dramatic improvement. That's right. Why would a smaller parameter model benefit more from thinking out loud? Well, it comes down to parameter density and retrieval friction. Okay. In a smaller model, knowledge is far more compressed.

4:21So that makes the activation threshold for any specific, you know, obscure fact much harder to hit during just a single forward pass. Oh, I see. Because it's crammed in there. Exactly. The massive models have enough parameters to recall facts with less effort, but the smaller models have the knowledge buried deep in their architecture, and they just struggle with inefficient recall. So the reasoning gives them a boost. Yeah, the reasoning trace provides the extra computational leverage they need to pull it to the surface. But I keep coming back to the mechanics of this, you know. If the model isn't doing logical math, what is actually happening under the hood during that reasoning trace?

5:00That is the million-dollar question. Right. Like, is it genuinely thinking about the capital of Iceland, or is the architecture just stalling for time? That was actually the very first hypothesis the researchers tested. And the idea that the AI is simply buying time introduces a fascinating mechanism called the computational buffer. The computational buffer. So is that effectively like a loading bar for an LLM? Functionally, yes. We have to remember how transformer architectures actually work. An LLM doesn't have a pause button where it can just sit and churn on a problem internally. Right. It always has to be outputting something.

5:38Exactly. The only way it can perform more computation is by predicting the next token. And every single token it generates requires a full forward pass through all the layers of the neural network. Meaning every time it spits out a word, it forces the attention heads to reevaluate the entire context window. and just run the math all over again. Exactly that. And that latent computation happens regardless of the semantic meaning of the token. Wait, really? Yeah. Generating a totally meaningless word still forces the model to cycle its engines. To isolate this, researchers looked at a brilliant case study involving Mary Engle Pennington.

6:13Oh, right, the refrigeration pioneer. Yep. So the question fed to the AI was, what year was Mary Engle Pennington inducted into the National Inventors Hall of Fame? A pretty obscure fact. Very. And when asked with reasoning turned off, just a standard greedy decode, the model confidently outputs the wrong year. It says 2019. And we know the real answer is 2018. So the zero shot attempt just outright fails. Right. But then they forced the AI to generate a dummy trace. A dummy trace. Yeah. They completely restricted its ability to search its memory or state facts. Instead, they forced it to generate just filler tokens.

6:52It literally just repeated the phrase, let me think, let me think, let me think. Oh my god. Over and over, matching the typical length of a normal reasoning trace right before it was allowed to output the final year. Wait, so no semantic clues at all. Just thousands of tokens of, let me think. Zero semantic value related to the prompt. And with just those filler words acting as a computational buffer, the model correctly retrieved 2018. Okay, that is wild. But honestly, it makes complete sense when you map it to human behavior. Well, going back to that cocktail party we talked about earlier, if someone hits you with a tough question, you don't just stare at them blankly while your brain works.

7:28No, that'd be weird. Super weird. You say, wow, that is a really fascinating question. I was actually just reading an article about that the other day and I was thinking. Right. You just ramble. Yeah. You aren't conveying any information. You are literally just generating filler words to stall while your brain runs a background search. It's the exact same utility for the AI. Generating those meaningless tokens gives the network the latent computational cycles it needs. Wow. With every forward pass of let me think, the attention mechanism is slightly adjusting the internal probability distribution.

8:05It is essentially churning the latent space. So it's giving that low probability fact about 2018 enough cycles to cross the threshold. Exactly. It bubbles up and becomes the most likely next token. So wait, does this scale infinitely? Like, if generating seller tokens buys at compute time, could I force a model to output, let me think, a million times and have it retrieve the most obscure knowledge in human history? There is a very hard ceiling on that, actually. Oh, really? Yeah, the data shows that if you force the model to stall for up to 2 ,048 tokens, the retrieval accuracy keeps climbing.

8:40But if you push it to an extreme, say, forcing 15 ,384 filler tokens performance severely degrades. Why is that? If more forward passes mean more compute, shouldn't more compute always be better? Because the context window gets saturated with noise. Ah, okay. After 16 ,000 tokens of, let me think, the attention heads are just drowning in a sea of meaningless repetition. The original prompt, you know, the actual question about Mary Engle-Pennington gets diluted in the vector space. The model loses the signal in the noise. It's like stalling for so long at the party that you completely forget what the person even asked you.

9:15You just lose the threat entirely. That is precisely the failure mode. And crucially, while this stalling tactic accounts for a massive chunk of the performance boost, it doesn't account for all of it. Meaning there's more to it than just stalling. Right. The dummy text never completely recovers the full accuracy of a natural reasoning trace, where the AI is allowed to actually talk about the subject. Which implies the actual vocabulary the AI uses when it thinks out loud must matter. If brute force compute cycles only get us halfway there, what is the semantic content doing? What's fascinating here is the second mechanism at play.

9:51It's called factual priming. When an AI is allowed to reason naturally, it isn't just stalling. It is actively engaging in generative self-retrieval. Generative self-retrieval. Okay, unpack the mechanics of that for me. How does writing down related thoughts pull up a hidden fact? Well, in cognitive psychology, there is this concept called spreading activation. Human brains are organized in semantic webs. So if you hear the word fire engine, your brain automatically primes related notes. Like red or siren or emergency. Exactly. By stating related facts in its reasoning trace, the AI is doing the digital equivalent.

10:27It alters the context window, which shifts the attention mechanism's focus, heavily weighting the vector space toward the target fact. So it builds a semantic bridge. Yes. Let's look at the 10th King of Nepal experiment from the data. Okay, I have absolutely no idea who the 10th King of Nepal is. Neither does the AI, usually. When asked directly without reasoning, the AI fails. It hallucinates a name, Jatari Mala, and gets it wrong. Standard hallucinations are an obscure entity. That happens all the time. But when reasoning is turned on, the AI takes a completely different path. In its hidden scratch pad, it doesn't just guess.

11:06It starts chronologically listing the first nine kings. Wow, really? Yeah, it outputs Prithvinar and Shah, then Prithap Singh Shah, all the way down to the ninth king, Mahendra. Oh, I see where this is going. Right, by generating that highly specific historical list, it radically alters its own context window. It lights up that exact neighborhood of its neural network. Which successfully primes it to retrieve the tenth king. Exactly, Burendra Birbhut Bikram Shah. It basically built a runway. It knew it didn't have the probability distribution to just drop out of the sky onto the 10th king, so it landed at the first king and drove all the way down the timeline.

11:42Yes, but as researchers, we can't just assume the facts did the work. What do you mean? Well, we just established that generating tokens provides a computational buffer, right? We have to prove that it was the factual content itself and not just the compute cycles used to generate a long list of names that solved the problem. That's a good point. How do you even untangle those two things? They're happening simultaneously. By running an isolation experiment, they extracted that list of the first nine kings from the AI's thought process. Then they completely disabled the AI's ability to reason, giving it zero extra computational time.

12:18Instead, they fed those nine facts back to the AI as plain context in the user prompt. So you just gave it the bridge directly, completely removing the stalling element. What was the output? It successfully predicted Birendra Bikram Shah on the first try. And to double check, when they replaced those historical facts with dummy text of the exact same token length, it failed. So that proves it. Yes. It proves that the factual content itself, the semantic manipulation of the vector space, is heavily responsible for the correct recall. It's essentially playing six degrees of Kevin Bacon inside its own latent space.

12:56Exactly. It uses the nodes it does easily remember, like the early kings, to mathematically lower the retrieval barrier for the node it struggles to remember. But wait, if the model is generating its own hints, what happens when it uses faulty materials to build that bridge? You've hit on the critical vulnerability right there. Because the attention mechanism relies so heavily on the preceding context, there is a massive risk of hallucination cascading into failure. Right, because if I am playing six degrees at Kevin Bacon and my first thought is mistakenly putting Kevin Bacon in the matrix. You're in trouble.

13:29Yeah. I'm going to end up at Keanu Reeves instead of whoever I was actually looking for. I've shifted my own semantic web in the completely wrong direction. Right. And AI does the exact same thing. If it hallucinates even one intermediate fact while it is thinking, it corrupts the entire downstream retrieval process. So how do they measure that? A massive audit was conducted using Gemini 2.5 Flash. researchers extracted every single intermediate fact generated in thousands of reasoning traces and they used web search to independently verify the accuracy of every single claim that is an insane level of granularity they audited the factual integrity of the AI's private thoughts they did what did the numbers actually show on the entity questions data set when the reasoning price was clean meaning every single intermediate fact it generated was verifiably true, the final answer was correct, 71.1 % of the time.

14:21Okay, 71.1 % is a highly capable retrieval rate. But what happens when it lies to itself? If the trace contained even one hallucinated fact, the final accuracy plummeted to 32.2%. Hold on, I need to push back on that a bit. A drop from 71 % to 32 % because of one bad fact? Yeah. That seems incredibly brittle. I mean, if I'm trying to remember a complex 10-step recipe and I accidentally convince myself my grandmother used cinnamon instead of nutmeg. I don't suddenly forget how to bake the entire cake, you know, that just get one ingredient wrong. Why is the AI's memory so fragile compared to a human's?

14:56Well, it comes down to how next token prediction works versus human compartmentalization. Your brain can isolate a mistake. Right. An LLM cannot. Its next prediction is mathematically chained to the entirety of its preceding context window. Oh, I see. Yeah. When it generates a hallucinated token, that fake fact becomes part of the prompt for the next word. It permanently alters the trajectory of the attention heads, basically steering the vector space away from the truth. But wait, couldn't this just be a symptom of question difficulty? What do you mean? Like, if a trivia question is wildly obscure, it's going to cause the AI to have confused, hallucinated thoughts, and it's going to result in a wrong final answer.

15:38How do we know the hallucination actually caused the failure and wasn't just a byproduct of a hard question? This raises an important question, and it's the classic trap of confusing correlation with causation. Right. But to definitively prove the bad facts caused the bad answers, the audit relied on a within-question comparison. Meaning they controlled for the difficulty of the prompt. Exactly. We aren't comparing an easy question to a hard question. We look at the exact same question. Okay. Going back to our past at CoteMetric, we have 100 different reasoning traces for the same prompt. we can look at attempt A and attempt B side by side.

16:15And because it's the same question, the baseline difficulty is identical. Correct. And the data is unequivocal here. If the model happened to hallucinate in its scratch pad on attempt A, it failed the final answer. And on attempt B? If it stayed strictly factual on attempt B for that very same question, it succeeded. Wow. Yeah, the hallucinated facts aren't just symptoms. They actively sabotage the model's ability to retrieve the correct final token. It's basically cognitive interference happening at the algorithmic level. That is fascinating. Knowing that true facts build successful bridges and fake facts burn them down, that gives us a tremendous amount of leverage.

16:53It really does. Which brings us to the practical application. We know the theory now, but how does this change the way you guys, the listeners, interact with these models today? Well, the incredible thing is that we can harness these mechanics immediately without even waiting for developers to retrain the underlying models. Oh, really? How? It involves a test time selection strategy. Test time selection. Okay, walk me through the implementation of that. Imagine you allow the AI to generate 100 different reasoning paths for a complex query. Okay. We now understand that traces lacking factual statements are just relying on the computational buffer.

17:30They are brute forcing compute without any semantic priming. Right, just the, let me think, stalling. Exactly. We also know that traces containing hallucinations are actively steering the attention mechanism off a cliff. Burning the bridge. Yes. So we simply implement a programmatic filter, we throw out all the paths that don't state concrete facts, and we use a verification step to discard any paths that contain hallucinations. So you're artificially curating the context window. Right. You're left with only the reasoning traces that are dense with verified true facts. And the results of that artificial selection are staggering.

18:06By prioritizing those clean, factual traces, researchers achieved a 12.2 % relative accuracy improvement on the simple QA verified data set. A 12.2 % boost just by filtering out the bad thoughts. Yep. That is a massive leap in reliability for, you know, enterprise or research applications. And I think there is a highly practical prompt engineering takeaway here for our daily workflows. Absolutely. When you are interacting with an LLM, whether you're trying to recall a specific legal precedent or cross-reference a medical statistic or pull obscure historical data, you shouldn't just ask the model for the final answer.

18:43No, you shouldn't. You should force the architecture to build that factual bridge first. You engineer the prompt to demand verifiable context. Yeah, you could explicitly prompt it, like, before answering my question, list five verified historical facts related to this topic. That's a great way to do it. Right. By forcing it to output true context first, you are manually triggering that factual priming. You are using the transformer's own attention mechanism against it, ensuring the context window is heavily weighted with true signal before it even attempts to predict the final target token. You are actively guiding its spreading activation.

19:17You are giving the neural network the exact semantic scaffolding it requires to reach its own latent knowledge. It fundamentally changes how you view these models, doesn't it? I mean, they aren't just static search engines where you type a query and a fact gets trended out. No, not at all. They are highly dynamic, probabilistic environments. So to synthesize what we've unpacked today, we basically looked at two major mechanistic discoveries about how AI reasoning actually functions. First, the computational buffer. We learned that generating filler words forces the model to run continuous forward passes, utilizing latent compute time to, you know, refine its probability distribution and bubble up compressed facts.

19:58And second, factual priming. The AI uses related trivia to alter its own context window, shifting the vector space to build semantic bridges to hard-to-reach knowledge. Much like human associative memory. Exactly. But with the critical caveat that a single hallucinated fact in that word will corrupt the entire retrieval process. Which is why demanding factual accuracy in the intermediate steps is just paramount. It really shifts our entire understanding of what it means for a large language model to know something. And I'd actually like to leave you with a final concept to mull over regarding that.

20:33Oh, please do. We have this tendency to anthropomorphize these models as digital encyclopedias from static databases of fixed information. Right, like a hard drive. Yeah. But if an AI's ability to recall a specific objective fact is this deeply dependent on the specific tokens it generates in the milliseconds right before answering, is knowledge in an AI really a fixed static database? Or is it more of a fluid dynamic state that has to be actively coaxed into existence? A fluid state that has to be coaxed into existence. Wow. That completely redefines how we should be interacting with these systems.

21:09It really does. The next time you have a word on the tip of your tongue, just remember, you aren't the only one building a bridge to find it. Thank you so much for joining us on this deep dive. Keep questioning, keep exploring, and we will catch you next time.

From the publisher

This research explores how reasoning helps Large Language Models (LLMs) answer simple, single-hop factual questions that do not logically require step-by-step thinking. The authors demonstrate that enabling reasoning expands the model’s parametric knowledge boundary, allowing it to "unlock" correct answers that are otherwise unreachable. This improvement is driven by two primary mechanisms: a computational buffer effect where extra tokens allow for more latent processing, and factual priming where the model retrieves related facts to bridge toward the correct answer. However, the study warns that hallucinating facts during the reasoning phase significantly increases the risk of providing a false final answer. Ultimately, the paper suggests that accuracy can be improved by prioritizing reasoning paths that contain verified factual statements.

More from Best AI papers explained

All 475 episodes
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMsBest AI papers explained · 21 min
Listen in VO