In short
LLMs “get lost” in multi-turn, underspecified conversations, showing a large reliability collapse versus single-turn benchmarks.
Guests
No guest names or backgrounds are provided in the transcript; it’s presented as a host-led discussion of a Microsoft/Salesforce preprint.
Key claims
Across 15 top models and 200,000+ simulated conversations, multi-turn performance drops ~39%. Aptitude falls only ~15%, but unreliability spikes ~112%, causing large run-to-run swings. Four failure traps: premature assumptions (early answers succeed 30.9% vs 64.4% if delayed), verbose bloated outputs (27% longer code), failure to course-correct (sunk-cost/anchoring), and “loss of middle turns” from attention recency/primacy.
Notable examples
“Jay snowballs” sharded simulation; coding solutions 850 vs 668 characters; temperature=0 doesn’t fix unreliability (~30%); recap/snowball help only slightly. Practical takeaways: if derailed, crop/abandon; consolidate requirements into a fresh single-turn prompt (“wipe the whiteboard clean”).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Paradox of AI Performance
1:00 to 2:25
Discussion on AI's exceptional test scores versus real-world performance.
“And that is exactly the tension of our deep dive today.”
Understanding Multi-Turn Dynamics
2:25 to 4:16
Insights into how multi-turn conversation differs from standard benchmarks.
“So, like, everything is just handed to it at once.”
The Chef Analogy: AI Under Pressure
4:16 to 6:34
Using a chef analogy to illustrate AI's cognitive overload in conversations.
“They start throwing flour everywhere and the kitchen basically catches on fire.”
Sharded Simulation Methodology
6:34 to 9:30
Exploring the sharded simulation approach used to test AI models.
“But real human conversation is much harder than that.”
Performance Metrics: Aptitude vs. Unreliability
9:30 to 11:10
Distinguishing between aptitude and unreliability in AI performance.
“So it's not like a car engine losing its horsepower.”
Psychological Traps in AI Communication
11:10 to 14:02
Identifying and discussing four psychological traps that hinder AI performance.
“So its own eagerness is literally toxic to its performance.”
The Loss of Middle Turns in AI Conversations
14:02 to 15:55
Learn about how AI's attention mechanism affects multi-turn conversations.
“in the middle of all this back and forth?”
Technical Fixes and User Strategies
15:55 to 18:21
Explore technical fixes and practical strategies for improving AI conversations.
“Which brings us to, I think, the most practical part of this deep dive.”
Managing AI Constraints in Conversations
18:21 to 20:18
Understand how to manage constraints when interacting with AI models.
“Yes, do not stubbornly argue with an AI that has taken a wrong turn.”
The Future of AI Conversational Interfaces
20:18 to 21:23
Consider the implications of AI's struggle to communicate like humans.
“Very ironic, which leaves you with a really profound, provocative thought to mull over.”
Transcript
Automatic transcript. May contain errors.0:00Imagine for a second that you're interviewing this job candidate. And on paper, I mean, this Kennedy is just an absolute genius. Right, like top of their class. Exactly. They take the written screening test. They score in the 99th percentile. They easily pass the bar exam. They ace the medical boards. They write flawless code. Basically perfection. But then, you know, you bring them into the office for an actual face-to-face conversation, a messy back-and-forth dialogue where you just give them a little bit of information, see how they respond, and then maybe give them a little more. Just a normal chat.
0:33Right. And the moment you do that, this absolute genius just completely unravels. Like they start jumping to these wild conclusions. They forget what you literally just told them two minutes ago. And they stubbornly refuse to correct their own mistakes. It is. Well, it's a profoundly jarring experience. Yeah. Because you're looking at the exact same entity, right? But the competence just vanishes the moment the interaction becomes dynamic. And that is exactly the tension of our deep dive today. We were looking at this really fascinating preprint research paper from AI researchers at Microsoft and Salesforce.
1:12It's titled very appropriately, by the way, LLMs get lost in multi-turn conversation. And I just want to point out, the scale of this study is massive. The researchers tested 15 top-tier models. So we are talking about the heavyweights here, like CLOG 3.7 Sonnet, Debsik R1, OpenAI's O3, and GPT 4.1, Gemini 2.5 Pro. All the big ones. All of them. And to really get to the bottom of this phenomenon, they simulated over 200 ,000 conversations. Wow, 200 ,000. So our mission today is to uncover this incredibly frustrating gap between how artificial intelligence performs in those pristine, you know, single-turn laboratory benchmarks versus how it actually behaves in the messy, underspecified reality of everyday use.
1:55Because the reality is very messy. So messy. We are going to explore the underlying mechanics of why these systems derail so spectacularly and give you some practical strategies to actually get the most out of your AI tools right now. Okay, let's unpack this. If these AI models are scoring 90 % and passing all these complex exams in the lab, why do they get so impossibly confused when you actually try to work through a real problem with them? Well, we really have to look at the massive difference between how they are tested and how they're actually deployed. Because in a standard benchmark evaluation, the model receives what we call a single-turn, fully specified prompt.
2:32So, like, everything is just handed to it at once. Exactly. The AI gets every single constraint, every variable, all the context perfectly bundled up up front. It basically has all the mathematical or linguistic information it could possibly need before it generates even a single word. Right. But out in the wild, everyday conversations are multi-turn and underspecified. You give it a little piece of the puzzle. It replies. You correct it. You add a new requirement. It's this gradual reveal of information. Which is how humans actually talk. And the researchers found a massive drop when you shift to that real world style, didn't they?
3:07Oh, a staggering 39 % performance drop across the board. Almost 40%. That's huge. It's massive. And no model is immune to this structural weakness. It impacts the smaller open weight models like Llama 3.18B. And it completely compromises the absolute state of the art models like Gemini 2.5 Pro and GPT 4.1. Nobody is safe. Right. The moment the conversation becomes multi-turn and underspecified, their logic just fractures. You know, it makes me think of a master chef. Imagine you have this world-class culinary expert. If you hand them a complete, highly detailed recipe with all the ingredients just perfectly measured out on the counter.
3:49The single-turn benchmark. Exactly. The single-turn benchmark. They will execute a flawless Michelin star dish. But the multi-turn reality is like asking that exact same chef to start cooking without a recipe at all. Right, just winging it. Yeah, and instead you sit in the other room and just shout ingredients at them every few minutes like, I want dinner. Okay, use chicken. Actually add some lemons. Wait, no, make it vegetarian. That sounds like a nightmare. Right. The chef doesn't just make a slightly worse dish. The constant pivoting totally overloads their working memory. They panic. They start throwing flour everywhere and the kitchen basically catches on fire.
4:23That is honestly the cognitive overload in that kitchen is a really great way to visualize what is actually happening inside the model's processing space. But to prove that the AI panics under those conditions, the researchers had to scientifically isolate and measure that messiness. Which brings us to how they tested it. Exactly. That brings us to their testing methodology, which they call sharded simulation. Sharded simulation. Okay, how exactly did they build that? Well, they took those standard, fully specified benchmark questions and intentionally fractured them into pieces or shards. So, for example, they used a standard mathematical reasoning problem about a kid named Jay making snowballs.
5:01Okay, classic math problem. Right. So the original lab prompt reads like this. Jay is making snowballs. He can build 20 snowballs an hour, but two melt every 15 minutes. How long will it take before he has 60 snowballs? So that is the perfect recipe. All the variables are right there on the counter. Right. The chef has everything. But in the sharded simulation, a simulated user completely changes that dynamic. Turn one, the user simply says, how long before Jay is ready for the snowball fight? Oh, wow. So the AI basically has to respond with nothing to go on. Exactly. It has to generate a response.
5:35Then turn two, the user says, he's preparing for a snowball fight with his sister. The AI responds again. Turn three, he can make 20 snowballs per hour. Oh, I see. So the system reveals the constraints one by one across multiple terms, forcing the AI to generate a response each time before actually receiving the next clue. Which is absolutely how we talk in reality. I mean, I know I don't walk up to a colleague and deliver a perfectly structured five paragraph essay of requirements before they even allow it to speak. Nobody does that. Humans naturally operate on what linguists call the principle of least effort.
6:09Principle of least effort. Right. We start with an underspecified prompt, basically expecting the other party to ask clarifying questions, and we just fill in the gaps as we go. And if we connect this to the bigger picture, most prior evaluations of multi-turn AI were fundamentally flawed. How so? Well, they treated the conversation as episodic, meaning each turn was just an isolated subtask that just happened to follow the previous one. But real human conversation is much harder than that. Yeah, it's connected. Exactly. It requires actively fusing new piecemeal information together with everything that came before it to resolve this highly ambiguous goal.
6:48Wait, okay, I have to push back on this just a little bit. Are we just tricking the AI here? I mean, if we know the neural network needs all the information up front to succeed, shouldn't the burden be on the user to just write a better, longer prompt? Like, why penalize the AI for our lazy communication? That's a fair question. But we really have to design systems for how human beings actually behave, not how engineers wish they behaved. Fair point. I mean, for a power user, writing a complex Python script, yes, building a highly detailed perfect prompt is a great strategy. But for the vast majority of normal users, it is completely unrealistic.
7:24People often don't even know what they fully want until they start interacting with the system. That's true. The conversation helps you figure it out. Exactly. The act of conversation itself is how humans explore and refine their own thoughts. So if the AI degrades when we communicate naturally, that is a failure of the software interface, not the human user. Okay, fair enough. That makes sense. So they run this sharded simulation with the snowball problem and the complex coding problems and the AI fails. But how exactly are they failing? Did the model suddenly lose all their intelligence or did the output just become more chaotic?
8:00The researchers made a really vital distinction here to answer that very question. They actually separated the model's performance into two distinct metrics, aptitude and unreliability. Okay, break those down for me. So aptitude measures the model's absolute best case performance. If you run the exact same simulation 10 times, the aptitude is the score of its top 10 % of runs. It basically measures the raw underlying capability to solve the problem. The raw horsepower. Exactly. Unreliability, on the other hand, measures the gap between its best case runs and its worst case runs. Okay, here's where it gets really interesting.
8:36Because you would intuitively think the multi-turn format just makes the AI dumber, right? That it just loses its raw aptitude when you confuse it with all those shards. But the data shows something completely different. Completely different. In single-turn prompts, high aptitude usually correlates perfectly with high reliability. If the model is smart, it is consistently smart. Right. But in the multi-turn sharded simulation, the AI's aptitude only drops a relatively minor amount, about 15%. Okay, 15 % isn't terrible. No, it's not. However, its unreliability skyrockets by 112%. 112%. So the intelligence is still in there.
9:12Oh, the model absolutely still has the underlying ability to solve the math or write the code. But its output just becomes wildly erratic. On the exact same multi-turn instruction, the model's performance can swing by under 50 percentage points between different runs. That's insane. So it's not like a car engine losing its horsepower. No. The aptitude is the horsepower, and the engine is still perfectly capable of going 100 miles an hour. The problem is that the unreliability spiked. It's more like trying to build a really tall tower on a totally shifting foundation. I like that analogy. Yeah, because the builder has the skill, the aptitude to build a perfect floor.
9:53But because the foundation keeps moving and shifting with every turn of the conversation, the whole structure just eventually collapses. You have the power to reach the destination, but the structural integrity is completely compromised. Exactly. The underlying horsepower is entirely useless if the foundation cannot support the accumulation of new data. OK, so if the underlying ability is still there, what is physically happening in the conversation to cause that foundation to shift and collapse? Like what is the AI actually doing during these back and forths to get so lost? Right. So the researchers analyzed the data logs of these hundreds of thousands of conversations and they actually isolated four specific psychological traps that the AI falls into.
10:34Four traps. OK, what's the first one? First, the models make premature assumptions. They just jump the gun entirely. Which I guess comes down to their core programming, right? Like they're designed to be helpful. So rather than pausing to just ask for clarification, they try to solve the whole puzzle immediately. Exactly. They attempt to force a final resolution based on just the very first tiny shard of data. And the researchers found that if the AI's first attempt to provide a final answer happens in the first 20 percent of the conversation, its final success rate plummets to 30.9 percent. Oh, wow.
11:06But if it waits, if it holds off its answer attempt until the last 20 percent of the conversation, its success rate jumps all the way up to 64.4 percent. That is wild. So its own eagerness is literally toxic to its performance. Precisely. And if they are jumping the gun and trying to solve a really complex problem with only one shard of information, they must be generating a whole lot of text to justify those early guesses, right? Which has to create a massive mess. And that leads directly to the second trap, overly verbose and bloated answers. The bloat. Yes. When the models generate these long, rambling responses early in the conversation, they end up inventing all these hidden assumptions and constraints that you, the user, never actually asked for.
11:50They are essentially poisoning their own context window. And for those listening who might not know, the context window is basically the AI's short-term working memory. It's the total amount of text it can actually look back at to understand what's happening. Right. So when it fills that memory with its own rambling guesses, it just leaves less room for the actual facts. Exactly. And to give you a really tangible sense of that bloat, the researchers looked at the coding tasks. The correct coding solutions generated in the multi-turn settings were 27 % longer than the exact same correct solutions generated in single-turn settings.
12:24Wait, really? The exact same solution? The exact same solution. We are talking about 850 characters of code versus just 668 characters. The Militern code was fundamentally bloated because the AI just kept adding unnecessary complexities to justify its own early guesses. And every single one of those extra characters isn't just harmless text, right? It's a new, completely hallucinated constraint that the AI is basically forcing you to deal with in the next turn. Exactly. Which makes me think about what happens when you try to correct it. If the AI just wrote an 850-character essay of bloated code, it's not going to just happily delete it when I give it the next chart of information, is it?
13:02Not at all. And that is the third trap. Failing to course correct, the AI falls victim to this bizarre machine version of the sunk cost fallacy. Oh, like it can't let go. Right. It anchors heavily on its own previous outputs. Once it puts an assumption down in text inside that context window, it literally treats its own generated text as absolute gospel. So we are constantly fighting its memory of its own mistakes. When I reveal a new piece of information that contradicts his early assumption, it just tries to awkwardly duct tape the new constraint onto its flawed foundation rather than just erasing the board and starting over.
13:41That's spot on. It is exactly like arguing with someone who just refuses to admit they misunderstood you five minutes ago. We all know someone like that. Right. And they are just doing intense mental gymnastics to make their original flawed point somehow still fit the new reality. That sounds absolutely exhausting to deal with. And while it is busy obsessing over its own mistakes, what happens to the actual facts I gave it in the middle of all this back and forth? Do they just get lost in the noise? They do. And that brings us to the fourth trap, the loss of middle turns. So the AI's attention mechanism, which is basically the underlying mathematical software that tells the model which words in the conversation are most important.
14:19Right, what to focus on. Yeah, it exhibits extreme recency and primacy bias. It assigns really high mathematical weight to the very first thing you said to it and the very last thing you just said to it. But all those crucial constraints and shards I revealed in the middle of our, say, 10-minute chat. They get completely washed out. The attention mechanism effectively ignores them because they are just buried under the bloat of the AI's own rambling responses. Wow. Okay. Now, what about the newer reasoning models? Because you mentioned OpenAI's O3 and DeepSeq R1 earlier. Right. The reasoning models.
14:52Yeah, these models are literally designed to think longer, you know, generating hidden chains of thought before they speak. Doesn't that extra compute time fix these exact issues? Well, what's fascinating here is that the extra reasoning compute does not fix the problem at all. In fact, by the sheer nature of how they operate, they can actually make it worse. Wait, worse? How? Because these reasoning models are explicitly designed to generate long, complex chains of logic. So they naturally produce conversational answers that are, on average, 33 % longer than non-reasoning models. Oh no. Which plays right into the second trap, the bloat.
15:27Exactly. They generate 33 % more text, which gives them 33 % more room to invent constraints, trip over their own flawed assumptions, and completely derail the structural foundation of the chat. Giving a system more time to think really does not help if it is fundamentally thinking about the wrong things and just anchoring on its own early mistakes. Okay, so the developers' neat tricks, like adding hidden reasoning compute, aren't fixing the shifting foundation. Which brings us to, I think, the most practical part of this deep dive. Since the underlying tech is still struggling with this, what can you, the listener, do right now to survive this lost-in-conversation effect?
16:06Let's start with a highly technical fix that the researchers tried, which ultimately proved pretty much useless. Right. There is a setting called temperature, which controls the randomness and creativity of an AI's output. Usually it's set around 1.0. Now, you might logically assume, ah, the AI is being unreliable and erratic. I will just turn the temperature down to zero to make it totally deterministic, strictly logical, and entirely predictable. Yeah, that makes perfect intuitive sense to me. Just remove the randomness from the settings, remove the unreliability from the conversation. Well, the researchers found that even at a temperature of strictly zero, the unreliability stubbornly remains around 30 percent.
16:43Wait, why? If it is zero temperature, shouldn't the underlying software do the exact same thing every single time? You'd think so, but it's because of the realities of hardware and what computer scientists call floating point math. Floating point math, okay. Yeah. Imagine millions of tiny mathematical calculations happening simultaneously across a massive cluster of physical microchips. Sometimes, just because of how computers round off microscopic decimals, the exact order of operations shifts by a fraction of a percent. Oh, I see. Think of it like drawing a line with a really long ruler. If your angle is off by just a fraction of a millimeter at the start, by the time you draw the line all the way across a massive piece of paper, you are inches off target.
17:25In a single turn prompt, that microscopic hardware rounding error just doesn't matter. But over five or six turns of a conversation, that one tiny ripple compounds and completely alters the trajectory of the chat. So we can't just flip a magic setting and fix the math. It just ripples out of control. What about those experimental software fixes they mentioned in the paper, like recap and snowball? Right. So the researchers tried forcing the system to automatically repeat all previous instructions. The recap method forced the AI to summarize at the very end, and Snowball constantly reminded the AI of the whole chat history with every single new message.
18:02Did that work? It helped slightly, but it still did not bridge the gap back to that pristine, single-turn benchmark performance. The AI's attention mechanism still gets hopelessly bogged down by the bloat of its own previous answers. So what does this all mean for you as a user? How should our listener actually change their workflow today? Well, the paper outlines two very specific, very practical takeaways that don't rely on developers fixing the math. Number one, if time allows, try again. Yes, do not stubbornly argue with an AI that has taken a wrong turn. If you notice it hallucinating a constraint or fighting you on a piece of code because of that machine sunk cost fallacy we talked about.
18:40Just crop it. Drop it. Do not spend five turns trying to correct it. It is hopelessly anchored to its mistake. Cut your losses, abandon the chat entirely and start a fresh one. And that flows perfectly into takeaway number two. Consolidate before retrying. Right. Before you close that frustrating, derailed chat, ask the AI to summarize all the requirements and constraints you have finally agreed upon. Take that clean, consolidated summary, open a brand new conversation, and paste it in as a single-turn prompt. You're basically doing the sharding in reverse. Exactly. You are manually building the perfect, fully specified recipe for the AI on a totally clean counter.
19:19The paper actually mentions that power users on the cursor IDE, which for those who don't know is a very popular AI coding environment, they already do this instinctually. Like they know the AI's foundation degrades over a long session, so they constantly wipe the slate clean and start fresh with a consolidated prompt. You really have to treat the AI not as a human collaborator who learns and grows with you over a two-hour meeting, but as a brilliant amnesiac. A brilliant amnesiac. Right. It is incredibly smart. It has the horsepower, but you have to keep giving it a fresh whiteboard with all the notes perfectly organized.
19:54Brilliant Amnesiac is really the perfect conceptual model for current large language models. They possess incredible raw aptitude, but until developers start prioritizing multi-turn reliability over just chasing raw single-turn intelligence on these leaderboards, this massive flaw just remains. We, the users, basically have to actively manage the constraints of the conversation. It is such a wild paradox to wrap your head around. I mean, we have built these incredibly powerful conversational interfaces, and it turns out the one thing they are fundamentally terrible at is actually holding a conversation.
20:29It is pretty ironic. Very ironic, which leaves you with a really profound, provocative thought to mull over. We know that human nature dictates that we communicate through underspecification. We give instructions piece by piece. We clarify as we go. It's just how we naturally relate to one another. But if AI structurally degrades when forced to communicate this way, are we building conversational interfaces that are just fundamentally incompatible with human nature? That really raises an important question for the future of the technology. Like, will the next massive leap in AI come not from models that can just spit out better, longer answers, but from models that actually have the patience to ask better questions before they start answering?
21:08Right? Models that don't just act like that brilliant job candidate who freezes, panics, and just talks over you in the interview, but one who actually sits back, listens, and pieces the puzzle together with you. Until then, remember, when in doubt, wipe the whiteboard clean.
From the publisher
This research paper from Microsoft and Salesforce identifies a significant performance gap in Large Language Models (LLMs) when they transition from single-turn to multi-turn, underspecified conversations. Through large-scale simulations, the authors found that even state-of-the-art models suffer an average 39% drop in performance when instructions are revealed gradually rather than all at once. This degradation is primarily attributed to a phenomenon called "lost in conversation," where models make premature assumptions, propose incomplete solutions, and fail to recover once they take a wrong turn. The study decomposes these failures into two specific metrics: a slight loss in aptitude and a massive increase in unreliability. Ultimately, the findings suggest that current evaluation methods overestimate model capabilities by ignoring the underspecification common in real-world human-AI interactions.




