Lost in the Middle (The Agents Season, Episode 3)

4 May 2026 · 20 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

“Lost in the Middle” explains a technical “U-shaped” attention/memory effect in LLMs with long context windows: models perform best when relevant information is at the beginning or end of the provided context, and worst when it’s in the middle (“dead zone”), even when the context window is large enough.

Guest backgrounds

No guests mentioned; this is a solo host episode.

Key claims

Context windows are finite working memory (tokens). Adding more context doesn’t reliably help; relevant middle content can be effectively ignored. The effect persists across context sizes and saturates quickly with more documents.

Notable examples

A Stanford/ Berkeley 2023 paper (“Lost in the Middle, How Language Models Use Long Contexts”) varied the position of the correct document among lists (e.g., 20 vs 50). For GPT-3.5 Turbo, performance could be worse with documents than with no documents. Follow-up MIT 2025 work (“On the Emergence of Position Bias in Transformers”) attributes the bias to causal masking and residual-connection “anchoring” favoring early/late tokens. Implication for agents: front-load critical instructions; longer context alone won’t prevent losing key step-3 info buried mid-history.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Context Windows in LLMs

1:16 to 3:37

Explore how context windows affect LLM performance during conversations.

“If you want to start at the beginning, you can go back a couple of episodes.”

The Lost in the Middle Phenomenon

3:37 to 6:00

Examine the findings of a Stanford paper about context location affecting performance.

“actually the name of the paper that we're going to dive into today.”

Implications of Context Position on Performance

6:00 to 8:02

Discuss how context position impacts the ability of LLMs to recall information.

“And it introduces this idea of a U-shaped curve.”

Human Memory Analogies and Context Saturation

8:02 to 10:04

Learn about the parallels between human memory and LLM performance under context saturation.

“And it was like latching onto maybe the first or the last incorrect document and using that to give an incorrect answer.”

Architectural Factors Affecting LLM Attention

10:04 to 14:00

Delve into the architectural reasons behind LLMs favoring information at the beginning and end of sequences.

“an analog in human memory and psychology.”

Understanding Context in AI Agents

14:00 to 18:10

Learn how context architecture affects the performance of AI agents.

“Those are skip connections that allow information to bypass attention.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Katie Malone:When I was a young scientist and I was learning how to give a good talk, I got some advice and it was have a good intro and have a good ending and don't worry too much about what happens in the middle. Like it matters, don't get me wrong, but if you're really going to focus on someplace, make it the beginning and the end because that's what people are going to remember. And the stuff in the middle, you can kind of just get away with being a little more loosey-goosey. And I was thinking about that as I was preparing for what we're going to talk about today, because LLMs are kind of the same way.

0:36Katie Malone:In this episode, we're going to explore a really interesting phenomenon on how attention works with long context windows. This is the third in our season-long arc about AI agents, so we are going to tie it back to the specific implications that it has for AI agents. But this is an idea and a topic that I think even without that agentic motivation, it's just really interesting and cool. And so I'm excited to talk to you about it. Today's episode is about Lost in the Middle, how models do and don't use context. You're listening to Linear Digressions. So just as a quick recap, this is part of a season where we're unpacking AI agents piece by piece.

1:24Katie Malone:If you want to start at the beginning, you can go back a couple of episodes. The first one is what even is an AI agent. And then in the last episode, we talked about the React algorithm and tool use. Now this is the third one about lost in the middle. But if you haven't heard those two and you just want to dive right in, this is going to be a good episode no matter what. And so I want to take that concept from the intro of nailing the beginning and the end, but then what's in the middle is a little bit of a gray zone. This is something that if you've worked with AI a lot, might start to ring true with you in some of the conversations you've had with LLMs.

2:03Katie Malone:There's this common phenomenon where you start out a conversation with an LLM and it's crisp, it's tight, it's responding to everything you say, it feels really good. And then as you start to go, as the conversation deepens, you're working through something complicated, maybe going back and forth, you notice that the conversation quality starts to degrade. Maybe it's forgetting things that you told it early on. Maybe it gave you some advice that contradicted what it said 20 minutes ago. Maybe it's just something feels a little bit off. And this is not just you. This is an actual technical phenomenon.

2:39Katie Malone:It has a name in the research literature, and it's a consequence of the context window. Context window, probably something that you've heard of before. This is the finite amount of information that a model can hold in its working memory at any given moment. So as far as you're concerned, if you're an LLM, everything outside of that window simply doesn't exist for the model when it's generating its response to you. For a chatbot, this can be annoying. It means that conversations as they go on and on tend to degrade. At a certain point, you kind of have to just start over. But in the context of AI agents, which might be working on a task across dozens or even hundreds of steps.

3:17Katie Malone:And as it's going, it's accumulating observations, it's following threads, it's maintaining this running understanding of where it is and where it's been. This is one of the fundamental constraints of the field. So we're going to talk today about the scientific exploration of this lost in the middle phenomenon. And that was actually the name of the paper that we're going to dive into today. It's a 2023 Stanford paper. and what they did in that paper is they looked very carefully at how models actually use the context they're given. The title of that paper is Lost in the Middle, How Language Models Use Long Contexts.

3:54Katie Malone:But first, let's take a step back and talk about the context window just a little bit more. Probably something that you've heard of, but a quick level set on what the context window actually is. It's a term that gets thrown around a lot, but it's worth the moment to take a clear explanation for anyone who hasn't bumped into it directly or so we're all thinking about it in exactly the same way in the context of this particular conversation. So when you're interacting with a language model, whether it's an AI chatbot or there's an agent, everything the model can see, which will include your instructions, the conversation history, any documents you've fed it, tools that it's used, and what they've returned, all of that is living inside the context window.

4:36Katie Malone:It's not a stretch to say that this is the model's working memory. And context windows are finite. They're measured in tokens, usually about a word is about 0.75 tokens. And as models have gone on, the context windows have gotten bigger. Some of the early models had context windows of a few thousand tokens, so you could have inputs or conversations of a length of a few pages or a few dozen pages. more recent models as the field has progressed those have had context windows of hundreds of thousands of tokens and even up to millions and there's a natural assumption which is bigger is better when it comes to context windows so the more context that the model can have access to the more that it can hold in its memory it can handle longer tasks it can maintain coherence across more steps and this is true to a certain extent and it's certainly the way that mostly people think about context windows in the context, sorry, of a new model.

5:37Katie Malone:They'll say new model comes out, context window is the biggest we've ever had, or what have you. And this is taken as this kind of raw capability of the model, the bigger is always better. But that's not really the case, or it's at least a little bit more complicated than that. So that's where this paper from Stanford and Berkeley comes in. Again, Lost in the Middle, How Language Models Use Long contexts. And it introduces this idea of a U-shaped curve. So here's the setup. In this paper, the researchers ran a series of experiments where they were studying how language models solved a task, namely reasoning over input documents.

6:18Katie Malone:So for example, here are 20 documents. One of them contains the answer to a question, find the answer. So pretty straightforward. But what they did in this paper that was clever was they systematically varied where the relevant document appeared. Sometimes it was first in the list, sometimes it was at the end of the list, and sometimes it was somewhere in the middle. And then they studied what happens. If the models were using their context windows robustly, namely they were actually reading and paying attention to all of the information they've been given equally, then the position of the relevant document shouldn't matter much.

6:51Katie Malone:Like it's in there somewhere, the LLM is going to find it, doesn't matter if it's the beginning, the end, or anywhere in the middle. Wherever it sits, models should find it. But as you may be guessing, at this point, that is not what they found empirically. They found something that was closer to a U-shaped performance curve, like the letter U. The models would do the best when the relevant information was at the very beginning of the context or at the very end. But if it's somewhere in the middle, the performance was degrading significantly. And this was true even when the context window was well within the model's technical capacity.

7:27Katie Malone:At the time that this was done, this is 2023, so the models they were looking at were the GPT-3 series. And in one case, they even found that for one of the models, GPT-3.5 Turbo, the performance of the model was actually worse when the answer was in the middle of the document list than when there were no documents at all that were supplied and the model was just producing the answer to the question from memory. So it's like it did worse on the question by having access to all of these documents available to it than it would have done by just relying on its training data. So it's kind of like this middle position document with the correct answer was effectively invisible to it.

8:06Katie Malone:And it was like latching onto maybe the first or the last incorrect document and using that to give an incorrect answer. So this is pretty crazy. What this implies is that just having more context is not going to make a model better at performing. Because if the information that you actually need to perform the task ends up buried in the middle of the context, that extra context on either side of you can actually make the performance worse. And just to be clear, this research finding doesn't go away when you have larger context windows. So they tried it on context windows of different sizes and found that performance as a function of position was nearly identical for all the cases that they studied.

8:50Katie Malone:So you can have a giant context window, 100 ,000 token context window, and that's not going to help you if the model doesn't actually pay attention to the middle of the information that you've given to it. In this paper, they did another experiment too, where they were studying what this effect looks like as a function of the number of documents that are available to the model. So they looked at this effect when they gave it 20 documents. They also looked at this effect when they gave it 50 documents. So you may think that if you give it more information, it has more material to reason over. Maybe you raise the probability that you have the correct answer in there somewhere.

9:26Katie Malone:So maybe if you just feed it more information, that raises the probability that it can latch on to the thing that's going to get the right answer for it. Finds that this isn't really the case. It saturates really quickly. when they go from 20 retrieved documents to 50, it improved performance in this study by about 1.5 % for GPT 3.5 Turbo. So even as you're giving it more opportunities to find the correct answer, you're giving it more information, the model seems to be just kind of saturated and overwhelmed. It didn't actually use that extra information effectively, even when there was relevant information somewhere in there.

10:03Katie Malone:And like we kind of alluded to at the beginning, this has an analog in human memory and psychology. So in that context, it's called the serial position effect. So in memory research, humans tend to best recall things that happen at the beginning and end of a sequence or a list or something like that. So primacy and recency, if you want to think of it that way, and stuff that's in the middle is the least remembered. But it's kind of surprising that this is echoed so clearly in the transformer architecture, which technically can pay attention to everything in its context equally. But what this paper showed in 2023 was that in practice, empirically, that is not the case.

10:46Katie Malone:So this was the point at which, as I was learning about this, I said, hold the phone. Why? I understand enough about transformers and the attention mechanism to understand that this is not supposed to be happening. The LLM is supposed to be able to look over the entire context that it has and pay attention to zero in on the part that's going to be most relevant. So what's going wrong in the middle? And it turns out there's been a follow-up research to try to unpack this, because this original 2023 paper, it was just reporting an empirical result. There were some hypotheses about why this might be happening, but as it is now, in 2026, there's been follow-up research that might have a clearer explanation for this phenomenon.

11:29Katie Malone:And it has to do with a couple of different architectural effects that are in play. The first thing that you might be thinking, if you're really intuiting about this strongly, is that this might be an artifact of the training data. So if the model is trained on data where important information usually appears at the beginning or end of documents, then maybe it learned that. So maybe it realizes that when humans are talking about things, they put important stuff at the beginning and end, just kind of like I did in my example in the intro. So maybe they're just learning that as a general pattern.

12:02Katie Malone:And that's partially true as a story about training data, but it's not the full picture, and crucially, it's not the primary explanation. The primary explanation has to do with the structure of LLMs. This was proven mathematically by MIT researchers in 2025. So that MIT paper is by Jin Yi Wu and collaborators, and that's called On the Emergence of Position Bias in Transformers. These models are using something that's called causal masking, which is a design choice that means that each token in the sequence can only attend, can only have attention that's going to tokens that came before it, not tokens that are coming after.

12:46Katie Malone:In other words, it's reading from left to right. And at any given point, you know what's been said up to that point, but you can't see into the future, so to speak. And this is a deliberate architectural decision. This is not like an accident. And that sets up an asymmetry where token number one, the very first token in the context, gets attended to by every single subsequent token. Token number two gets attended to by everything from token number three onward. Token number 500, which is sitting in the middle, gets attended to by only everything after it, which might be half the sequence. So what that means is that that first token is on the most computational pathways.

13:27Katie Malone:It's being paid attention to along the most channels in the LLM kind of doing the reading of the input text. But then as you get into the middle and the end of the text, you're on far fewer pathways. So you may be wondering, well, what's going on with the part where it remembers stuff at the end? So maybe this explains a little bit of intuition for like why stuff at the beginning is favored. Why is stuff at the end favored a little bit too? There's a second architectural feature in play here, which is what's called residual connections. Those are skip connections that allow information to bypass attention.

14:06Katie Malone:And this starts to get a little bit beyond some of my deep technical knowledge, to be totally honest with you. But as I understand it, basically what happens is that there's a separate anchoring mechanism for the stuff that's right at the end. So that can maintain a stronger signal across layers in a way that the middle can't. So if you don't have the advantage of being at the beginning and being on more computational paths, if you're not at the end where you have this additional structural anchor through those residual connections, if you're in the middle, then you're in this dead zone. And the fact that this has to do mostly with the architecture of LLMs themselves, it has to do with the way that transformers are working or layers are anchored to each other kind of helps explain why just making the context window larger doesn't necessarily solve this problem.

14:57Katie Malone:It's in the architecture of the model itself. So with all of that, what does this mean for agents? So if you're an agent and you're working on a long task, you are constantly accumulating context. So you might start out with the goal that you've been given, some beginning instructions, directions about what tools you have available to you, And then you go out and you start interacting with the world. And remember, our definition of an AI agent for this series is that something that can observe, reason, and act. And it repeats those three steps in a loop. And so as it's observing, it's getting information back from the environment.

15:37Katie Malone:It's reasoning, which tends to add tokens to the context as well. And so as it's going through and tearing out those tasks, it's accumulating context, sometimes quite quickly. And then if there's critical information from step three that ends up kind of buried under 20 steps of subsequent context by the time you get to step 25, then the model may have functionally lost that information. It didn't technically fall out of the context window, but it's in that middle dead zone. So if you're building an agent, this nudges you towards something that you might have already intuited anyway from experiments or just from your own sense of management logic maybe, which is that if you can, try to put important stuff at the beginning of the context window.

16:23Katie Malone:So front-loading critical constraints and instructions, not a bad idea. It's not fail-safe because agents accumulate context dynamically and you don't always control where information ends up, but do what you can to put the information that's important up front. And it's not necessarily going to make your life worse to have a big context window, but if you're having issues with memory management or with your agent seeming to struggle with long running tasks that you're giving it, just giving it a longer context window does not necessarily solve for all of this. So with all of this in mind, may be wondering, wow, it seems like context management is actually really important and how should I think about this as an agent builder or just as someone who wants to understand how agents work?

17:12Katie Malone:Well, that is what we are going to talk about in the next episode, context management. If you're not going to solve this problem by expanding the context, instead, real solutions here have to be smarter than that. You have to be actively managing the context. Then that motivates why context management is such an active territory right now. So that's where we're going to be going next week. So where does this leave us now? Well, the original Lost in the Middle paper was published in 2023. And since then, context windows have grown dramatically. And the mechanism has been proven mathematically. So if you're wondering, like, is this a solved problem?

17:50Katie Malone:It's partially solved. It's better. Models have gotten better at long context tasks through various training interventions, but there's still this geometric structure of causal masking that hasn't fundamentally changed. And the problem hasn't gone away. It's just become more manageable. So next time we will talk about what people have done to actually manage this constraint. We'll talk about a system that was inspired by operating system memory management. And we'll talk a little bit about compaction, which is a mechanism that's running quietly inside every long coding session you've ever had to manage this exact challenge.

18:24Katie Malone:With that, I want to thank you for listening. I want to remind you to come over to Substack and subscribe to our newsletter. If you haven't already, look for Linear Digressions and hit subscribe. I'd also like to, if you are not officially subscribed to this podcast, I'd like to invite you, whether it's through Spotify or iTunes, your podcast platform of choice, hit subscribe, leave a review. It really helps people find it. Thank you so much for joining us this week, and we'll talk to you next week all about context management.

19:21Ben J. Johnson:LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

From the publisher

Just like a memorable talk lives or dies by its opening and closing, LLMs have a surprisingly similar quirk: they pay close attention to what's at the beginning and end of their context window — and kind of zone out in the middle. This "lost in the middle" phenomenon has real consequences for anyone building AI agents that rely on long-context reasoning. In this episode we dig into the research behind how (and how poorly) models actually use the information you feed them, and what it means for the agentic systems we're all trying to build.

More from Linear Digressions

All 35 episodes
Lost in the Middle (The Agents Season, Episode 3)Linear Digressions · 20 min
Listen in VO