How do LLMs use their depth?

27 Oct 2025 · 12 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How LLMs use their depth during inference, arguing for a “guess then refine” workflow plus “complexity-aware depth use” (more layers for harder parts).

Guest backgrounds

No guests are mentioned; it’s a solo “Deep Dive” host discussing research.

Key claims

Early layers make high-frequency statistical guesses; most early top predictions are later overturned. Later layers perform massive contextual refinement, with a high decision flip rate. Depth allocation is task-dependent rather than uniform.

Notable examples

Pythia 6.9 layer-1 top tokens: over 75% from the top 1–10 most frequent words; only ~33% of final predictions come from that bucket. Decision flip rate: nearly 80% of layer-1 top-110 predictions change by the final layer (almost 100% for rarer early tokens). POS case: function words peak around layer 5; content words around layer 20. Multi-token fact recall: first token of “New York City” appears around layer 27; later tokens earlier (around layers 20 and 12). Constrained tasks: early layers filter valid options; deeper layers (e.g., up to MMLU) alternate top choices until late. Efficiency warning: early exiting may harm accuracy by cutting off crucial refinement.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding LLMs: Guess Then Refine

0:45 to 2:26

Exploration of the internal processes of LLMs during word prediction.

“And the findings seem really structured, quite revealing, actually.”

The Role of Early Layers in Predictions

2:26 to 3:55

Discussion on how early layers make statistical guesses and their implications.

“So with that technical grounding, let's jump into the core behavior, this guess then refine idea.”

Massive Contextual Refinement in LLMs

3:55 to 5:10

Examination of how initial predictions are revised significantly by deeper layers.

“It's the starting block for the real process, which is overturning that guess.”

Complexity-Aware Depth Use

5:10 to 6:40

Insights into how LLMs use layers flexibly based on task complexity.

“The model isn't just refining things uniformly.”

Case Study: Multi-Token Fact Recall

6:40 to 7:55

Analysis of how the model handles multi-token predictions and their computational depth.

“Now let's look at case study two, multi-token fact recall.”

Constrained Downstream Tasks in LLMs

7:55 to 9:48

Review of how LLMs adapt their depth for tasks with limited answer choices.

“That really implies the serious cognitive load is all front-loaded onto that first token.”

Implications for Building Efficient LLMs

9:48 to 12:04

Discussion on the findings' implications for LLM efficiency and potential pitfalls.

“So putting it all together, if we try to visualize the whole workflow, the LLM starts with a quick, high-frequency statistical guess.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. If you work with large language models, you know they're defined by their depth dozens, maybe over 100 layers stacked up. But here's this kind of internal mystery. When an LLM predicts the next word, is that decision basically made right away, like in the first layer, and then just cast along? Or does the model, you know, truly use all those layers to kind of think its way to the best answer? That internal journey, that process, is exactly what we're tracing today. For a long time, we've sort of treated LLMs like black boxes, haven't we? But now, thanks to some really insightful new research, we can start to peel back those layers.

0:35Our mission today is pretty straightforward. Let's try to understand the actual strategic workflow these models use during inference. How do they leverage all that computational depth? Right. And the findings seem really structured, quite revealing, actually. The research we're diving into today points to this kind of dual strategic framework for how LLMs operate. That's right. First, there's this strategy they call guess then refine. That describes the overall approach to making a prediction. And second, they use something called complexity-aware depth use, which basically shows they use more layers for the harder bits of the problem dynamically.

1:13Okay, but before we really unpack that whole strategy, we need to talk about how this kind of analysis is even possible. I mean, figuring out what's happening inside layer 5 of, say, a 40-layer model sounds incredibly difficult. Historically, we had tools like LogitLens, I think. But the research seems to say that for really robust insights, especially way down in those early layers, you need something more reliable. Yeah, that's exactly the methodological challenge. Older methods like LogitLens were sometimes a bit unreliable. It's like the model speaks a different language in its early stages.

1:46Those early feature representations, they operate in completely different mathematical spaces compared to the final output. So if LogitLens was trying to listen in on those early layers, it was kind of misinterpreting, like guessing what the model was saying. Pretty much, yeah. Like an inaccurate translation or eavesdropping without understanding the dialect. That's where this newer tool, TuneLens, really changes the game. It uses a learned transformation to think of it like a very specialized, reliable translator, specifically trained to facefully decode those intermediate layer outputs. It makes sure that what we think we see the model predicting at layer one is a genuine reflection of its internal state at that depth.

2:24much more accurate. Okay, that makes sense. So with that technical grounding, let's jump into the core behavior, this guess then refine idea. It all kicks off with what the researchers call early statistical guessers. Exactly. In the very first layers, the LOM is working with really limited information. It hasn't fully pulled in the context through attention yet, and it hasn't accessed the factual knowledge that's typically stored deeper down. In those middle MLP layers, the feed-forward networks where a lot of the world knowledge seems to live. So if context is low, what's your best bet? Go with the odds.

2:57Statistical probability, right? The most frequent stuff. Precisely. The predictions in those earliest layers are just dominated by high-frequency tokens. Super common words. For instance, with Pythia 6.9, they found over 75 % of the top-ranked token predictions at layer 1 belong to the top 1 to 10 most frequent words in the training data. Wow, 75%. Wait, if layer one is mostly just guessing V or A or punctuation, doesn't that kind of make those early layers seem almost useless? Why not just skip them and start computation where the context actually starts getting processed? That's a really good question, and it gets to the heart of why the guess is necessary, even if it's often wrong initially.

3:37You see, while only about, say, 33 % of the final predictions end up being from that super high frequency bucket, Starting there is actually the optimal strategy when you have minimal information. It maximizes your chance of being right initially, and it gives the later layers a starting point to work from to refine. I see. So the guess isn't the point. It's the starting block for the real process, which is overturning that guess. So let's talk about that refinement. How much does the model actually revise those initial guesses as it goes deeper? Oh, a massive amount. The paper calls it massive contextual refinement.

4:09And we can actually quantify it using something called the decision flip rate. That basically measures how often the top prediction from an early layer gets changed by the final layer. For models like Pythia 6.9b and Llama 3.8b, they found nearly 80 % of those top 110 predictions made way back in layer one are eventually modified or completely replaced by the end. 80%. That's huge. That's not just a minor adjustment. That's like a total rethink four out of five times. It means layer one is almost always wrong, fundamentally, even when it picks the statistically safest word like the. Exactly. And that flip rate goes up even higher if the token picked in layer one was rarer, say, something outside the top 1 ,000 most frequent tokens.

4:50In those cases, the flip rate is almost 100%. The prediction is incredibly temporary until context really kicks in. The permanence of a prediction, its likelihood of sticking, only increases as you go deeper and integrate more context. Okay, that distinction, the statistical guess versus the contextual refinement, leads us right into the second major point. The model isn't just refining things uniformly. It's intelligently scaling its effort, right? This is the complexity-aware depth use. It doesn't just grind through every layer the same way for every task. Precisely. It uses its layers flexibly.

5:24To show this, the researchers looked at three really interesting case studies. Let's start with the first one. Prediction by part of speech category or POS. Ah, so comparing easy structure words versus the harder meaning words. Exactly. For the easy category, you know, function words like determiners, the ad positions, punctuation marks, these are generally high frequency structural elements. The model figures out the correct token for these much earlier. It hits rank one, typically around layer five for both Pythia and Lama 3. Layer five. That's incredibly shallow. It's like the model just goes, okay, syntactically I need a small connecting word here.

5:59Bam. Done. Task complete, basically. Right. But then contrast that with the hard content words. Nouns, verbs, adjectives. These carry the core meaning. Predicting the right one requires much deeper semantic understanding and integrating the surrounding context. These words only emerge as the correct prediction much, much later, often closer to layer 20. So the first five layers handle the basic sentence structure, the syntactic scaffolding, you could say. And then the next 15 layers or so are really dedicated to the heavy lifting, integrating the meaning needed to pick the correct noun or verb for that specific context.

6:37That's a clear division of labor. It really is. Now let's look at case study two, multi-token fact recall. This is inherently harder, right? It involves retrieving knowledge stored within the model. Yeah, I'd imagine the computational effort ramps up significantly here. It does. Tokens related to fact recall generally only start appearing correctly after layer 15 or so. That's already much deeper than the layer 5 for function words. But what's really fascinating is how the model treats different tokens within a multi-token answer, like, say, New York City. Okay, so the first token, new, feels like the critical one.

7:09The model has to commit to the whole concept right there. That must be the hardest part computationally. You anticipated the core finding perfectly. For a multi-token answer, it's the first token that demands the greatest computational depth. It emerges really late around layer 27 for the first token of a three-token fact in the Pythia model they studied. Layer 27. Wow. The model is nearly finished computing before it feels confident enough just to start uttering the first word of the fact. But here's the kicker. The subsequent tokens, the second, York, and third city, they emerge much, much sooner.

7:45Often around layer 20 and layer 12, respectively. Once the model has successfully navigated the prediction of the start of that sequence, the rest becomes conditionally easier. The path is set. That really implies the serious cognitive load is all front-loaded onto that first token. It almost feels like predicting new requires some kind of look-ahead planning to make sure the whole sequence, New York City, makes sense. Is that how the researchers interpreted it, that the model needs to anticipate the full sequence? Yes, that's strongly suggested by the data. The model seems to reserve its deepest, most context-integrated computation for initiating that complex sequence.

8:19It's like it needs to commit to a multi-step plan before it even generates the first part. Fascinating. Okay, for our final case study, constrained downstream tasks. Things like multiple choice questions or maybe sentiment analysis where the answer is limited, A, B, C, D, or maybe just positive-negative. How does the model use its depth when the possible answers are restricted like that? It shows a really clear two-step process, again, split across the model's depth. Step one happens in the earlier layers, maybe the first half. And it's mostly about identification and filtering. Filtering what, exactly?

8:54Filtering for the valid options. The model rapidly identifies the allowed choices, say A, B, C, D, in a multiple choice question, and brings them into the top prediction ranks. That's the relatively easy subtask, just recognizing the valid vocabulary for the answer format. This gets done pretty early on. Okay, so the first half of the layers basically screen the candidates, identify the players on the field, and the second half, the deeper layers, that's where the real decision-making happens. Exactly. Those remaining deeper layers are reserved for the actual reasoning and deliberation. When looking at benchmarks like MMLU, that massive multiple-choice test suite, You can see the model using the second half of its depth to really weigh the evidence between those top-ranked valid options.

9:35It often switches its top prediction back and forth, sometimes right up until the very last layer, before settling on the final answer. It shows that deep contextual refinement is crucial, even when the choices are limited. So putting it all together, if we try to visualize the whole workflow, the LLM starts with a quick, high-frequency statistical guess. Then it subjects that guess to massive refinement based on context flowing in from the attention layers. And then it strategically allocates its deepest, most computationally intensive layers to tackle the hardest parts, whether that's picking the very first word of a factual recall or making that final reasoned choice in a complex task.

10:14That sums up the duality really well. They're early statistical guessers, yes. But crucially, they're also late contextual integrators. And they scale that integration effort intelligently based on the specific demands of the task, be it simple syntax or complex knowledge retrieval. So what does this understanding actually mean for people out there trying to build, you know, faster, more efficient LLMs? Well, these findings offer some direct clues for efficiency, but they also raise a bit of a red flag, I think. We now have a much clearer picture of where the computational heavy lifting happens.

10:47But this understanding might clash with some popular efficiency techniques. For example, the research suggests that methods like early exiting, where you stop the computation early if the model seems confident, might fundamentally conflict with this natural guess-then-refine dynamic we've been discussing. Ah, that makes perfect sense. If the model is inherently built to revise, to overturn its initial high confidence through lots of deeper refinement, then stopping it early, even if the current top prediction looks good at that shallow layer, means you're potentially short-circuiting the most critical part of its reasoning process.

11:22You're cutting off that essential contextual analysis. Precisely. The potential cost of exiting too early could be a significant drop in accuracy, maybe in unpredictable ways. So the final provocative thought we want to leave you with is this. If MLMs have essentially evolved through training to use their full depth to correct those initial, often superficial statistical guesses, what's the true, perhaps hidden cost, maybe in terms of unexpected or catastrophic errors, of forcing them to stop early, right when that crucial, massive contextual refinement is meant to be happening? It definitely raises a serious question about whether some efficiency hacks might be undermining the very mechanism that makes these models so powerful in the first place.

12:04Thank you for helping us dive into the internal life of these fascinating models today. We'll see you next time on the Deep Dive.

From the publisher

The research paper explores how Large Language Models (LLMs) utilize their depth during inference, proposing a "Guess-then-Refine" framework to explain layer-wise prediction dynamics. The authors use the TunedLens method to trace intermediate representations, revealing that early layers function as "statistical guessers" by promoting high-frequency tokens as initial predictions due to limited contextual information. As processing continues through deeper layers, these initial guesses undergo "massive contextual refinement" to become contextually appropriate tokens. Furthermore, the study demonstrates "Complexity-Aware Depth Use," where LLMs intelligently dedicate shallower layers to simpler tasks, such as predicting function words, while reserving deeper layers for more complex computations like recalling multi-token facts or reasoning through constrained-choice tasks.

More from Best AI papers explained

All 475 episodes
How do LLMs use their depth?Best AI papers explained · 12 min
Listen in VO