What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT

27 Sep 2025 · 17 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Effective reasoning in large reasoning models (LRMs) is driven by structure and failure management, not by longer chain-of-thought (CoT) or more “review” tokens.

Guests

No guests named; the episode is a solo “Deep Dive” discussion.

Key claims

(1) Across 10 LRMs and datasets (HARP, AIM-25, GPQA Diamond), controlling for question difficulty shows shorter CoT and lower review ratio correlate with higher accuracy. (2) A structural metric, failed step fraction (FSF)—the fraction of reasoning-graph steps in failed exploratory branches—predicts correctness better than length/review. (3) FSF can be used for test-time selection: among 64 candidates, choosing the lowest-FSF trace boosts accuracy ~5–13% on AME 2025. (4) Failed branches “pollute” later attempts: deleting dead-end segments increases continuation accuracy ~8–14%.

Notable examples

Claude 3.7 as an exception where higher review ratio correlated with better math; FSF graphs with “pink” failed branches vs “blue” successful paths; answer entropy declines even when final answers are wrong.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Challenging the Old Wisdom

0:45 to 2:10

Discussion on the traditional belief that more compute leads to better reasoning results.

“It was treated almost as a compute-based magic bullet.”

Evaluating Reasoning Models

2:10 to 3:48

Analysis of a study evaluating different reasoning models using complex problems.

“That's just sheer verbosity measured by the total character count in the chain of thought trace.”

Superficial Metrics Explained

3:48 to 5:04

Defining superficial metrics like length and review ratio and their implications.

“So those really long, verbose chains, the ones filled with all that self-correction and review tokens, they weren't signs of thoroughness usually.”

Discovering Failed Step Fraction

5:04 to 6:40

Introduction of the Failed Step Fraction (FSF) as a new metric for reasoning quality.

“With actual useful structured processing, we need something deeper.”

Understanding FSF's Impact

6:40 to 8:12

Exploring how FSF correlates with accuracy in reasoning tasks across models.

“And I found the method for actually getting this FSF score really slick.”

Interventions to Improve Reasoning

8:12 to 10:44

Discussion of interventions to test the effectiveness of FSF in improving reasoning quality.

“Is it just as important for a simple calculation as it is for, say, a really complex scientific puzzle?”

The Pollution Effect of Failures

10:44 to 12:40

Investigating how past failures influence future reasoning attempts.

“This is the finding that, for me, completely changes how we need to think about context in these models.”

Shifting Focus from Quantity to Quality

12:40 to 14:03

Emphasizing the importance of structural quality over brute-force computation in reasoning.

“It's like having mental scar tissue from a bad idea.”

Understanding Confidence in AI Reasoning

14:03 to 16:08

Explore how AI models assess confidence and the implications of their reasoning flaws.

“Like if a reasoning branch starts accumulating a high FSF score, you actively cut it off before it gets long enough to really pollute the context.”

The Human Element in AI Logic

16:08 to 16:44

Reflect on the parallels between AI reasoning failures and human confidence in logic.

“but in the complex chains of reasoning you encounter everywhere, maybe even in your own work.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. Today we are attempting to hack the brain of the machine. Specifically, we're diving deep into the large reasoning models, the LRM. We want to uncover the genuine secret behind what makes a chain of thought trace, you know, the Cotee, truly successful. Yeah, and for years, the general engineering approach, the sort of common wisdom, was basically just give it more time to think. Right. More compute equals better answers. Exactly. That thinking was really exemplified by early strategies using things like special weight tokens or these complex deliberation systems.

0:35Think S1. It all operated in this pretty simple assumption. If the model thinks longer, spins more tokens, and reviews more, accuracy has to go up. Almost like a brute force method for intelligence. Pretty much. It was treated almost as a compute-based magic bullet. So the sources we're looking at today, they pull back the curtain on this huge systematic evaluation that was specifically designed to challenge that simple idea. Okay. We're looking for the true driver of reasoning quality, moving past just counting tokens and really getting into structural intelligence. Okay, so let's unpack this mission then, because this wasn't some small-scale test, was it?

1:10Not at all. This team did a really rigorous evaluation across 10 different large reasoning models, and they threw some hard stuff at them. Complex math problems using data sets like HARP and AIM-25, and also that high-stakes scientific reasoning from GPQA Diamond. Their goal was to analyze three key metrics. You had the surface level ones everyone used before length and review ratio, and then this brand new structural measure. And right off the bat, the initial results give you that first big aha moment because they pretty much dismantle that old longer is better story. How so? Well, when they started looking at those standard metrics, the old standbys, but crucially controlling for how hard the question actually was, the data just consistently pointed the other way.

1:57Okay, hang on. Before we get to that reversal, let's just clearly define those superficial metrics. Because like you said, for a long time, these were the indicators. Engineers used them to justify, you know, throwing more compute, more tokens at the problem. All right. We can define them pretty easily. First, you've got length. That's just sheer verbosity measured by the total character count in the chain of thought trace. It's basically raw output volume. How much did it type? Okay. Simple count. Simple count. Second, the review ratio. Now, this one required a bit more work. The researchers actually used another LLM, the Mavericks model, to go through and label specific chunks within the CoET.

2:33Label them as what? As review. So if the model was reading back over something, checking its work, maybe restating a premise or explicitly backtracking, that got flagged as review. The key thing is these tokens, they didn't advance the reasoning. They didn't push towards the solution. Ah, okay. So it's like the AI pausing to double check or maybe just rambling while it thinks. Exactly. It could be careful checking or it could be confused rambling. That's the problem with the metric, really. And I think the way they tested the correlation here is really clever because they knew, obviously, a super complex question, like a level six art problem.

3:09It's always going to need a long trace, right? Success or failure. Right. You can't compare apples and oranges. So instead of comparing easy questions to hard questions, they controlled for that difficulty. They looked at multiple attempts, multiple generations for the exact same problem, which isolates the quality of the reasoning path, not just the inherent difficulty of the task itself. And when they did that, the results were almost shocking just in how consistent they were. Across most of the models, most of the data sets, the pattern was clear. For the same problem, shorter reasoning traces correlated with higher accuracy.

3:43Shorter was better. Shorter was better. And similarly, a lower review ratio generally correlated with higher accuracy. So those really long, verbose chains, the ones filled with all that self-correction and review tokens, they weren't signs of thoroughness usually. They were signs of trouble? Often, yeah. Signs of a confused kind of meandering process. But I have to jump in here because there was one notable exception you mentioned. Claude 3.7. On the math problems, it actually showed a positive correlation with review ratio. More review meant better math scores for that specific model. That's right.

4:16So doesn't that suggest that sometimes, maybe for certain tasks or certain model architectures, a disciplined, structured review is valuable? We can't just toss out self-correction entirely, can we? No, and that's a critical point. It absolutely highlights that model-specific styles exist. But it also, I think, underlines the difference between disciplined self-correction, like maybe Claude 3.7 was doing, and just confused self-correction or rambling. For the vast majority of models in this study, that review ratio is just an indicator of inefficiency. They were burning tokens looking back but not getting much benefit.

4:51The Claude 3.7 exception kind of proves the rule that sheer length or just the volume of review tokens, it's a flawed proxy for quality. It mixes up simple wordiness, the model just kind of wandering around. Yeah, wandering is a good word for it. With actual useful structured processing, we need something deeper. Exactly. We need a metric that measures the efficiency of the steps taken, not just how many words it used to get there. So, if just counting characters and review tokens is a distraction, where do we look? We need to stop counting and start mapping, maybe. Mapping the actual thought process.

5:26Precisely. That's the structural shift. The researchers moved past just tokens and characters. They extracted a full reasoning graph for each chain of thought. A graph. Like nodes and edges. Exactly like that. Imagine taking that long, linear text output and breaking it down, segmenting it into connected nodes. Each node represents a distinct logical step, a move forward in the reasoning. Okay, I can picture that. And doing that lets you see the structure of the thinking, the pathways it took, the successful branches that led towards the answer, and crucially, the dead ends, the parts where it went wrong.

6:01And this structural view, this map, that's what let them define the true predictor. This is the key metric, failed step fraction or FSF. FSF. It's conceptually beautiful in its simplicity. It's just the fraction of all the reasoning steps in that graph that belong to failed exploratory branches. Failed branches, meaning? Meaning abandoned attempts. So if the model started down some path, maybe a mathematical derivation, realized the logic wasn't working out and then branched back to try a different approach. Yeah. That whole failed path, all the steps in it, they contribute to the FSF score. So it's measuring the density of failure within the entire thought process, how much of the effort was wasted.

6:39Exactly. How much of the journey was spent going down blind alleys. And I found the method for actually getting this FSF score really slick. They basically made the AI draw a map of its own mind. Tell us about how they used CLOD 3.7 for that. Yeah, it was a very clever sort of metacognitive technique. They used CLOD 3.7 itself, and this is critical with its high-level reasoning switched off. Okay. The model was basically instructed, take this raw account of T-text, this stream of consciousness, and convert it into a structured graph Vs format. Graph Vs, that's for drawing graphs, diagrams. Right.

7:13It's a language for defining graph structures so you can visualize them. So Claude's only job here was to act like a parser, a cartographer. It mapped the sequence of steps and explicitly labeled the successful paths, often shown in blue in the research papers, and the abandoned failed paths, which were usually rendered in pink. So the model wasn't solving the problem again. It was just drawing the org chart of its own previous attempt. That's fantastic. It really is. And once they could calculate FSF this way, its dominance as a predictor was immediate. It was undeniable. It emerged as a far stronger, far more stable predictor of correctness than length or review ratio ever were.

7:52Stronger across the board. Yes. It showed a significant negative correlation. So lower FSF means higher accuracy across all 10 models they tested and across both types of problems, the complex math and the scientific reasoning. It was finally a universal metric for reasoning quality. Which leads us to the really critical application question. When does FSF matter most? Is it just as important for a simple calculation as it is for, say, a really complex scientific puzzle? Well, the data suggests it matters most when quality is absolutely paramount. The link between low FSF and high accuracy was particularly strong, particularly consistent for the harder questions.

8:32Like those high-level heart problems. Exactly. Heart levels four, five, six. The ones where the answer is definitely not trivial. When the reasoning path has to be long and complex, the models that succeed are the ones that somehow maintain an extremely low failure density. They plot a clean path. Well, the unsuccessful ones. They just get bogged down. Multiple dead end branches, lots of pink on their graph, a high FSF score. Okay. So this strongly suggests that avoiding failure, keeping the path clean is way more important than just the sheer volume of effort or thought. But so far, we've mostly talked correlation, right?

9:06Low FSF correlates with success. Right. To prove that FSF isn't just reflecting success, but actually causing it, or at least as a lever you can pull, the research moved into interventions, two key ones. Yes. Moving from correlation to causality. The first intervention was what they called causal intervention one, test time selection. Test time selection. Yeah. They basically took the problem of generating one good answer and flipped it into a selection problem. So for each question on AME 2025 and GPQA Diamond, they generated a big pool of attempts, 64 candidates and identities. Wow. Okay. 64 different ways the AI tried to solve it.

9:42Exactly. Now, normally, you might pick the final answer based on some simple confidence score from the model, or maybe just pick one randomly from the top few. But instead, they used FSF. They used FSF to re-rank all 64 candidates. They selected the trace with the lowest FSF score, the one with the structurally cleanest path as the final answer. Just took the answer from that Sinaiti. This is called pass at one selection. And did it work? It worked remarkably well. FSF yielded the largest performance gains compared to other selection metrics they tried. On that AME 2025 math set, using FSF to pick the best trace boosted accuracy by roughly 5 % to 13 % over just randomly picking one of the 64.

10:22That's a significant jump. It is. It confirms FSF is more than just a measurement. It's a powerful tool for identifying and selecting high-quality reasoning, even when it's hidden within a batch of potentially flawed attempts. Okay, that's strong evidence for its usefulness. But the second intervention, the direct put modification, that one I think really gets under the hood. It tackles that question of whether failed attempts actively poison the well. Do they harm subsequent thinking? This is the finding that, for me, completely changes how we need to think about context in these models. Yeah.

10:56They asked the crucial question, if a model goes down a wrong path, realizes its mistake, and successfully backtracks, is the slate really wiped clean? Or does that memory of failure linger? Exactly. Does the failed branch actively harm the next attempt, even after the model seemingly moved on? So how did they test this pollution effect? It was ingenious. They took kokitis that had already produced an incorrect final answer. They used two different models, DeepSeq R1 and GPT-OS 120b, to get some diversity. They went into these incorrect traces, identified the specific sections corresponding to the field branches, the pink parts on the map, and then they just removed them.

11:31Yeah. Edited them right out of context history. Whoa, okay. So they essentially gave the model amnesia about its mistakes. Precisely. They kept the initial problem set up, kept the parts of the trace that were still on a potentially viable path, but deleted the dead ends. And then they asked the model to continue generating from that cleaned up point. Now, the intuitive expectation, or at least my intuitive expectation, based on how we think about backtracking, is that the model should be fine. It recognized the error, it backtracked. Removing the bad part shouldn't change much if it truly recovered.

12:04You'd think so, but that's not what happened at all. The crucial finding was removing the failed branch substantially increased the accuracy of the continuation. How much? Just deleting the evidence of the past mistake raised the probability of the model generating a correct final answer by a really significant margin between 8 % and 14%. Wow. Okay, so that provides really strong causal evidence that those long-field branches, they actually bias subsequent exploration. Yes. The model doesn't fully unsee its earlier mistakes. So it's not truly resetting its state when it backtracks. It's more like it's trying to build a new path, but it's still navigating around the wreckage of the old one.

12:44That's a great way to put it. It's like having mental scar tissue from a bad idea. The mistake itself, even though abandoned, contaminates the context window. It might divert attention or pollute the probabilities for the next token, making another error more likely. The failure isn't just wasted compute. It's actively harmful. It's like a memory artifact that drags down the rest of the reasoning process. Exactly. And that, I think, is the deepest insight of this whole study. Success in complex reasoning is fundamentally about failure management. Which brings us right to the main conclusion and what this means for the future of building these things.

13:22Effective reasoning isn't about brute force thinking time. It's characterized by this low failure density. The focus has to shift from sheer quantity compute, token count, to structural quality. It really demands a major pivot in how we approach scaling these LRMs instead of just indiscriminately generating longer and longer COPs, which, as we now see, just increases the chance of generating a long, harmful, failed branch. Future development really needs to prioritize what the researchers call structure-aware test time scaling. We need quality-focused strategies. Okay, so what does that look like in practice?

13:56If you're an engineer building the next generation LRM, what do you do differently based on this? Well, it means focusing on techniques that actively limit failure propagation. Maybe highly targeted branch pruning. Like if a reasoning branch starts accumulating a high FSF score, you actively cut it off before it gets long enough to really pollute the context. Intervene earlier. Intervene earlier, yeah. Or it could mean developing better context control mechanisms. Maybe ways to isolate those failed segments automatically or significantly reduce their influence, their attention weight in subsequent steps.

14:30The efficiency gain here isn't just about saving tokens anymore. It's about preventing self-sabotage. Exactly. Preventing the model from tripping over its own past mistakes. That really is a sharp pivot. Shifting from how long the AI thinks to how cleanly it thinks. FSF seems clearly the critical metric now for both efficiency and accuracy. It finds that clean, reliable path through the problem. And there's one final, slightly unsettling finding we should touch on. It speaks to the reliability or maybe the self-awareness of these models. The researchers also tracked the model's internal confidence as it generated the Cothee.

15:06They used answer entropy as a proxy for confidence. Okay, how confident was it feeling at each step? Yeah, and they found a very consistent pattern here too. What was it? The models consistently declined in entropy, meaning they grew more confident as they went along. And this happened regardless of whether the final answer they were heading towards was actually correct or totally wrong. Oh, no. So they become highly confident even when they're fundamentally wrong. Even when their internal reasoning trace, their FSF score, is terrible. That's what the data shows. They get very certain of their conclusion, even if the internal process that led them there was, you know, structurally compromised.

15:45Filled with those pink, failed branches. They're sure they're right, even when the map of their own thoughts shows they got lost multiple times. Precisely. That is genuinely unsettling. Because if even these advanced reasoning models, the ones we're increasingly relying on for complex tasks, can be so highly confident while being demonstrably structurally wrong, well, it makes you think. How should you, the listener, assess confidence, not just in AI, but in the complex chains of reasoning you encounter everywhere, maybe even in your own work. When you're tackling a tough problem, what internal checks, what external metrics do you need?

16:22How do you spot an abandoned or failed branch in your own logic, especially when that feeling of certainty, that low entropy state, feels so convincing? It seems that being highly confident while being deeply wrong is a very human failing. These machines have managed to replicate perfectly. Always look for the pink nodes, I suppose, even in your own head. We'll see you next time for the next deep dive into the sources.

From the publisher

This academic paper investigates what makes a Chain-of-Thought (CoT) trace effective for Large Reasoning Models (LRMs), challenging the prevailing idea that **longer reasoning traces and increased review behaviors automatically lead to better performance**. Through a systematic evaluation across ten LRMs on math and scientific reasoning, the authors demonstrate that **shorter CoTs and lower Review Ratios are often associated with higher accuracy**. To identify a more fundamental predictor, the research introduces a graph view of CoT and defines the **Failed-Step Fraction (FSF)**, which consistently and robustly predicts correctness across models and datasets, outperforming length and review metrics. Finally, test-time selection and direct CoT editing interventions provide causal evidence that **low FSF improves accuracy** by mitigating the bias that failed reasoning branches introduce to subsequent steps.

More from Best AI papers explained

All 475 episodes
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoTBest AI papers explained · 17 min
Listen in VO