In short
Research on “Turnwise” shows multi-turn chat can degrade a language model’s performance versus its own single-turn ability (“conversational amnesia” / “turn decay”). It argues common benchmarks are saturated and mis-measure failures caused by conversational history confusion.
Guests
No guest names or backgrounds are provided in the transcript; it’s a host/deep-dive style conversation.
Key claims
Longer interaction can make models “dumber” due to context acting as an “active anchor.” MT-Bench can’t separate baseline-skill failure from history-induced distraction. A new benchmark, Turnwise Evil, uses pairwise self-comparisons (multi-turn vs identical single-turn question) and finds win rates well below 50%.
Notable examples
Organizing a small closet (tips → drawer question). Sports-commentary seed prompt (e.g., Jordan’s last shot, “Miracle on Ice”) used to synthesize prior turns. Training Olmo 3-7b Instruct with ~10,000 Turnwise data examples improved Turnwise Evil self-score up to 12%; SFT hurt single-turn ability, while DPO helped and flattened turn-decay from turn 1 to turn 8.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Illusion of AI Understanding
1:00 to 3:08
Learn about the misconception that more conversation leads to better AI understanding.
“And we're going to look at how engineers are finally measuring this hidden flaw and the incredibly clever backward engineered trick they're using to fix it.”
Testing Flaws in AI Models
3:09 to 5:59
Understand the flaws in existing AI testing methods and their implications.
“And that is exactly why a new benchmark called Turnwise Evil was created.”
The Turnwise Evil Benchmark
6:00 to 8:00
Explore how the Turnwise Evil benchmark isolates conversational stamina from intelligence.
“Even Frontier state-of-the-art models, we're talking about massive models like GPT-5-chat, they also underperform on this multi-turn test compared to their own single-turn abilities.”
Synthetic Data Methodology in AI Training
8:01 to 10:29
Uncover the synthetic methods used to train AI and the impact on performance.
“And placing that seed prompt at the end is crucial.”
Implications for Future AI Conversations
10:30 to 11:16
Consider the future of AI and its potential for long-term conversational memory.
“And, you know, if an AI's ability to maintain context can be drastically improved just by exposing it to synthetic simulated histories of eight turns, well, what does this mean for the future of digital companions?”
Transcript
Automatic transcript. May contain errors.0:00You know, you probably assume that the longer you talk to an AI, the better it understands you. Right. I mean, that's the logical assumption. Yeah, because you're giving it more context, right? You're refining your prompts, building a relationship, basically. Exactly. But, well, today we are looking at some fascinating new research that proves the exact opposite is true, which is wild. It really is. Mathematically, according to the data, the more you talk to an AI, the dumber it actually gets. It's a phenomenon that feels a lot like conversational amnesia. It really does. And it completely breaks the illusion of what these systems are supposed to be.
0:39Yeah. Because we interact with them through this chat interface, right? Yeah. Which inherently promises a fluid, ongoing conversation. But it's kind of a trick. Exactly. That interface is essentially a mask. It's hiding a fundamental limitation in how these models actually operate behind the scenes. Which is exactly the mission of our deep dive today. We are exploring recently released findings that reveal the mechanics of why AI models struggle so much with these multi-turn conversations. Right. And we're going to look at how engineers are finally measuring this hidden flaw and the incredibly clever backward engineered trick they're using to fix it.
1:15Yeah. And to really grasp the problem, it requires a shift in how we think about an AI's education. OK. How so? Well, while we expect the AI to be this dynamic conversational partner, the data shows that historically it has been trained much more like a student aggressively cramming with flashcards. Well, flashcards. That makes sense. Right. It is built and optimized for a single prompt and a single response. Look at the front of the card, get the answer on the back. Boom. Done. Exactly. So when you force that system into a fluid, unpredictable, multi-turn interaction, you are dragging it completely outside of its core developmental environment.
1:55Right. So to understand why it gets confused, we first have to understand how we actually test its intelligence. Yes. Which brings us to the testing flaw, because most AI models, as you said, are trained and evaluated on single-turn data when prompt one responds. Right. A flashcard. But we do have multi-turn tests, right? Like an empty bench. We do, but the new analysis explains the multi-turn versus single-turn gap. Those existing tests like MT-Bench are, well, they're essentially saturated. Meaning they're too easy now. Yeah. The models score so highly on them that the tests have lost their diagnostic value.
2:32But more importantly, they're inherently flawed in what they actually measure. How so? They cannot distinguish between a model failing, because a task is genuinely too hard, versus a model failing because it just got confused by the conversational history. Okay, let's unpack this for a second. It's like failing a driving test. Okay. Did the student fail because they genuinely don't know the mechanical steps to parallel park? Like they just lack the baseline skill. Right. Or did they fail simply because they got distracted by the examiner talking to them in the passenger seat? Oh, I love that analogy.
3:04That's a perfect way to look at it. You have to isolate the distraction from the baseline skill. Exactly. And that is exactly why a new benchmark called Turnwise Evil was created. Turnwise Evil. Yes. Its entire architectural purpose is to isolate that conversational stamina from general baseline intelligence. Okay, but how exactly does this new test work without getting clouded by other variables? Because untangling baseline smarts from distraction in an AI sounds, well, it sounds impossible. It does, but they do it through this highly controlled pairwise comparison. Yeah. And the setup is just incredibly elegant.
3:43Lay it on me. So let's use a practical example from the research. Imagine a multi-turn conversation about organizing a home. Okay. In turn one, the human user asks, what are some tips for organizing clothes in a small closet? Very standard. So the AI gives its normal advice, right? Decluttering, using vertical space, that sort of thing. Right. The AI does its job. Then in turn two, the user asks a follow-up. what are different drawers I should have for clothes? Okay, so now it has to answer the drawer question while remembering the small closet context. Exactly. Now, to evaluate this properly, Turnwise Evil takes that exact same final question, what are different drawers I should have for clothes, and asks it to the model in a single-turn setting.
4:23Completely fresh. A blank slate with no prior conversation whatsoever. Wait, so you have the model answering the exact same question in two totally different environments, one with the baggage of the previous conversation and one totally cold. Yes. And here is the most brilliant part. They use a metric called turn wise evil self. It compares an AI's multi-turn response directly to its own single turn response. To itself. To itself. Wait, wait. So you're telling me an AI can actually perform worse than itself on the exact same question just because we gave it the prior context of the conversation.
4:59What's fascinating here is that yes, absolutely. That is wild. The data shows win rates well below 50 % in this setting, meaning it loses to itself most of the time. So the context isn't helping it. The context is an active anchor dragging it down. Exactly. It mathematically proves that conversational history actually degrades the model's baseline ability. It really is the distracted driver failing to parallel park. Precisely. We know it has the intelligence to answer perfectly because it did it in a single turn setting. The conversation is what tripped it up. So if this test is explicitly designed to catch this degradation, who is actually failing it?
5:36Because I'm assuming it's just the older clunky models. You would think so, but the data is eye-opening. Open models like ULMO-3 and LAMA 3.18B have massive performance gaps. And just to clarify for everyone, LAMA 3.18B, that's an 8 billion parameter model, right? So it's relatively small and efficient compared to the Giants. Yes, smaller models are definitely more prone to context overload. but here is the real shocker. Okay. Even Frontier state-of-the-art models, we're talking about massive models like GPT-5-chat, they also underperform on this multi-turn test compared to their own single-turn abilities.
6:11Wait, really? GPT-5-chat? Yes, dropping by five full points. Five points. Just because you said hello a few times before asking your question, that is a massive downgrade. It is. And it proves this multi-turn gap is an industry-wide blind spot. models, everyone is failing. OK, so if everyone is failing, how do researchers actually get the data to teach these models to do better? I mean, why not just use transcripts of real humans talking? Well, collecting real multi-turn data from human interactions is incredibly expensive and slow. Right. Humans are messy. Lots of typers. Exactly. And if you try to speed it up by having AIs simulate users and talk to each other.
6:54Well, yeah, just have two AIs chat. It leads to what researchers call conversational drift. Without strict human guidance, the models just wander off topic. They just ramble. So human data is too slow and AI data just turns into a rambling mess. Right. So the solution detailed in this new analysis is a highly scalable synthetic method called turnwise data. Turnwise data. How does that work? It's incredibly clever. They basically build the data backward. They take a single seed prompt. Like what? Let's say what songs were popular in the 1920s. OK, a specific historical question. Right. And instead of trying to steer a conversation toward that question, they use a strong model to independently generate previous related turns.
7:35Independently. Yes. So it generates a standalone first turn, like what genres of music were big in the 20s, and a separate second turn, like what types of dances were popular. Ah, so because they are independent, they don't drift off topic. Exactly. Then they snack them all together, keeping the original seed prompt as the very last turn. It's almost like the game show Jeopardy. You start with the final destination, the answer, and you synthetically generate the logical questions that would lead up to it. That's a perfect analogy. And placing that seed prompt at the end is crucial. Why? Because it preserves the high quality nature of the original data set, but forces the AI to practice maintaining context to get there.
8:15There's a great example in the data about sports commentators. Oh, let's hear it. The final seed prompt asks the AI to act as a sports commentator for a final game winning play. A fun, complex prompt. Right. But using turnwise data, they generate previous turns about, say, Michael Jordan's last shot and then the 1980 miracle on ice. Okay. So the conversation flows through all this heavy sports history and right into asking the AI to be a commentator. The AI has to navigate all that contextual baggage to do its job. That makes total sense. Okay, so we have the new test and we have the synthetic Jeopardy style data.
8:51What actually happens when a model is trained on it? The results are incredible. They tested this on the Olmo 3-7b Instruct model. Okay. And by adding as few as 10 ,000 of these synthetic multi-turn conversations during post-training. Wait, only 10 ,000? That's tiny in the AI world. I know, but the model's Turnwise Evil self-score jumped by up to 12%. 12 % from just 10 ,000 examples. That's amazing. It is, but there's a delicate balance here. The researchers noted that using supervised fine-tuning, or SFT, actually degraded the model's single-turn abilities. Wait, really? Why would it get worse?
9:26Because SFT forces the model to memorize exact phrasing, which can ruin its general intelligence. Yeah. But using preference tuning, specifically DPO. DPO, right. Yeah. That improved the multi-turn skills without breaking the single-turn smarts. It teaches the concept of focus rather than just rote memorization. Ah, I see. And this profoundly affects what they call turn decay. Normally, an AI's performance decays drastically from turn one to turn eight of a conversation. It just falls off a cliff. Right. But training with this new data significantly flattens that curve of forgetfulness. That is huge because for you listening right now, this is the difference between an AI that gets overwhelmed and gives up after three emails worth of context and one that can truly brainstorm with you all afternoon.
10:13Exactly. It fundamentally changes the utility of the tool. So just to quickly summarize this deep dive, we started by uncovering the hidden gap in AI conversations, right? How it's trained like a flashcard reader. Then we saw how testing AI against itself revealed the distraction flaw. And finally, how this backward engineered synthetic data is flattening the curve of AI forgetfulness. It's a massive leap. And, you know, if an AI's ability to maintain context can be drastically improved just by exposing it to synthetic simulated histories of eight turns, well, what does this mean for the future of digital companions?
10:49Oh, wow. If AI learns to perfectly navigate an afternoon's conversation today, how soon will it be expected to remember a conversational thread you started with it three years ago? Three years. That's intense. It raises an important question about where the line between processing data and true memory really lies. That is such a fascinating, provocative thought to leave you with as you mull this over. Thank you so much for joining us on this Deep Dive. Hopefully your next AI brainstorming session goes a little smoother.
From the publisher
This research addresses the performance gap in large language models between single-turn and multi-turn interactions. The authors introduce TURNWISEEVAL, a new benchmark that isolates conversational ability by comparing model responses in long dialogues against equivalent single-turn prompts. To improve model performance, they also developed TURNWISEDATA, a scalable pipeline that generates synthetic multi-turn training data from existing single-turn instructions. Their experiments demonstrate that even advanced models often struggle with extended context, but incorporating a small amount of this synthetic data during training significantly boosts chat capabilities. Ultimately, the study highlights that multi-turn proficiency is a distinct skill set that requires dedicated evaluation and specialized training data.




