In short
Why long “test-time compute” reasoning in LLMs is costly (KV-cache memory and quadratic attention time), and how rolling window reasoner (RWR) prunes redundant middle tokens to keep accuracy while cutting compute/memory.
Guest backgrounds
No guests mentioned; it’s a solo “Deep Dive” discussion.
Key claims
Long reasoning chains contain internal redundancy (repetitive verification, self-correction cycles, exploratory dead ends). Most intermediate tokens don’t matter for the next token. RWR keeps a fixed initial context window (problem “constitution”) plus a recent scratchpad window, discarding middle tokens from the KV cache. It’s inference-time only (no retraining).
Notable examples
AIME/HMMT math, Live CodeBench code generation, GPQA academic QA; models/families include “Qwen3” and “4o Reasoning Plus,” plus DeepSeekLlama8b and Nematron4b. Reported results: ~50% KV-cache budget with near-identical accuracy; ~2x memory savings and ~4x compute reduction.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenges of Complex Reasoning in AI
0:45 to 3:03
Explores the efficiency issues faced by AI models during complex reasoning tasks.
“But the cost side is, well, it's terrifying.”
Understanding the Redundancy Problem
3:03 to 4:33
Discusses the redundancy in reasoning chains and its impact on computation and memory.
“This brings us to the core insight, the thing that unlocks the solution.”
Introducing the Rolling Window Reasoner (RWR)
4:33 to 7:13
Presents the RWR technique to optimize reasoning and manage memory effectively.
“And this leads us directly to the solution that acts like a cheat code against the tyranny of quadratic scaling, a rolling window reasoner, or RWR.”
Performance Benefits of RWR
7:13 to 9:18
Details the efficiency gains and practical implications of using RWR in AI models.
“It's a fundamental change in the economics of running these models.”
Further Implications and Future Research
9:18 to 12:34
Considers the broader implications of RWR and potential areas for enhancing AI efficiency.
“Okay, so RWR is defined by its simplicity, keep the beginning, keep the end.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. Today we are peering directly into the thought process of our most advanced language models. Specifically, what happens when they have to do some truly complex reasoning? Right. You know, think about a model solving a competitive math problem or writing a piece of code that takes, I don't know, multiple internal steps to debug. This whole process, this generation of long chains of thought researchers call it test time compute. It's absolutely crucial for getting high accuracy. But right now, it is just crushing the efficiency of these models. It really is. It's the AI's core conflict, right?
0:38Performance versus practicality. We spent years developing models like the ones used on the A mathematical reasoning benchmark and training them to think longer before giving an answer. And it works. It works great for performance. But the cost side is, well, it's terrifying. Those internal reasoning chains can swell up to tens of thousands of tokens. We've seen averages around 23 ,000 reasoning tokens just for one complex aim problem. 23 ,000 tokens. That's like having the model read a 50-page paper internally just to give you one sentence. That scale sounds impressive, but what is the raw technical challenge that makes those long sequences so ruinously expensive?
1:17The villain here is foundational to the transformer architecture itself. The core mechanism is causal self-attention. That's what dictates how the model processes the sequence of tokens it has generated so far. And the problem is that it scales. It scales quadratically in time complexity. Quadratically. Okay, tell us what that means in plain language, because that's the technical bottleneck that terrifies hardware engineers. Right. Well, think of it like planning a party. If you have 10 guests, the number of unique handshakes you need to worry about is relatively small, maybe 100 interactions.
1:49But if you double the guests to 20, the potential interactions don't just double. They shoot up? They quadruple. If you double the length of your sequence, that's N, the computation explodes fourfold. We call it Ogin on Stala. So longer thought just means exponentially slower generation time. So the speed gets exponentially worse, but we also have a memory problem, right? We do. A big one. Every time the model generates a new token, it has to store the key and value vectors for that token in what's called the KV cache. The KV cache. You can think of the KV cache as the model's short-term memory bank.
2:23It's storing everything it needs to remember about the conversation, or in this case, its own line of reasoning so far. And that scales. It scales linearly in space complexity. It just gets bigger, token by token. That's O-N-N. When you're talking about sequences approaching 23 ,000, or I've seen even up to 64 ,000 tokens, that memory adds up really, really fast. It leads to immense memory demands. I mean, the sources show that when you hit those extreme sequence lengths, you start getting out-of-memory error. Even on the best hardware. Even on high-end, extremely expensive hardware, like 40 gigabyte A100 GPUs, it makes deploying these highly capable, reasoning-optimized models just incredibly cost-prohibitive.
3:02Okay, so we're stuck between this astronomical cost because of quadratic time and massive linear memory, and models that have to think long to be accurate. This brings us to the core insight, the thing that unlocks the solution. If they're generating 23 ,000 tokens, is there any way to prune that history? I mean, do all 23 ,000 tokens really matter equally? And what's fascinating here is that the researchers found the answer is a definitive no. This is the core aha moment. They observed that within these sprawling lengthy reasoning chains, there is immense internal redundancy. Redundancy. Yeah.
3:39Despite the LLM being highly optimized, it doesn't need to reference every single historical token to generate the effective NEXT token. So these highly capable reasoning models, while smart, are actually surprisingly inefficient thinkers. They generate a lot of unnecessary baggage in their thought process. If you dive into the analysis of these reasoning traces, you find these predictable patterns of redundancy. You see repetitive verification steps, cycles of self-correction where the model explores a path and then just discards it, or these exploratory paths that quickly become less relevant as the solution gets closer.
4:12So these intermediate steps, all those middle tokens, they become computational deadweight. Precisely. They are the primary source of the computational and the memory overhead. The goal, therefore, it kind of shifts from how to generate a massive cache to how to surgically and selectively attend only to the most critical information. And eliminate that redundancy. Right. And this leads us directly to the solution that acts like a cheat code against the tyranny of quadratic scaling, a rolling window reasoner, or RWR. It's described as a simple yet effective inference time technique. So if we agree the middle tokens are the problem, what do they propose to manage that crippling KV cache?
4:50Well, RWR introduces a fixed-budget approach. It essentially transforms the model's short-term memory from this, you know, massive endless scroll into a manageable two-part filing system. You can think of it as a two-window mechanism that determines exactly which key and value vectors to keep in the cache. Let's break down those two parts. What has to stay and what gets dropped? Okay, so the first window, you can call it the context window, is always maintained. Technically, it's WUF, the fixed start. And this is absolutely critical because it preserves all the foundational information, the initial problem statement, the instructions, all the constraints.
5:27It's the model's constitution, basically, the core rules. That's a great way to put it. It's the set of rules that must be available for every single step. Got it. The static core of the problem. You can't start solving a physics problem if you forget the initial conditions. So what's the second window? That's the last window, WL. We can think of this as the scratchpad window. It maintains only the most recent reasoning steps and immediate logical dependencies. It's the tokens that capture the model's current line of thinking. Where it's actively trying to figure out the very next step. Exactly.
5:58And everything between that context window and the scratchpad window. All those redundant intermediate middle tokens, the vast majority of those 23 ,000 tokens, are just summarily discarded from the KV cache. They're just deemed temporally distant enough to be unnecessary for the immediate next step. So instead of a cache size that just keeps growing and growing with the sequence length T, we transform it into a fixed manageable budget B, which is the sum of the context window and the scratch pad window. Right. So what does this surgical approach do to that terrifying quadratic time complexity and the runaway memory?
6:33This is where the magic happens. Yeah. By keeping the total budget B fixed and small, the attention time complexity shifts dramatically. Instead of scaling quadratically with the sequence length T, that's OET, the Wissana stretcher board, it shifts to effectively OBT. Okay. And for memory, the space complexity for the KV cache drops from linear with the sequence length, so from OT to just linear with the fixed budget, OB. So for these extremely long thought chains, whereas way, way larger than our fixed budget B, we move from exponential time and massive linear memory down to essentially linear time and a constant memory cost.
7:13That's it. It's a fundamental change in the economics of running these models. Which is huge. It is. And the truly crucial detail here is that this RWR technique is an inference time optimization. It works out of the box. Meaning no retraining. No expensive retraining or fine-tuning of models that were initially trained using that full sequence attention. It's a pure execution time hack. That's a phenomenal technical achievement. Let's talk about the payoff. I mean, what does this actually translate into for the end user, the efficiency gains? The results are consistent and likely impressive across the board.
7:45RWR maintains near-identical accuracy on these challenging tasks. While requiring. Only 50 % of the original KV cash budget. Half the cash for the exact same result. That just sounds like massive savings. It translates directly into tangible deployment benefits, that 50 % cash reduction. That corresponds to 2x memory savings. Okay. Two times the memory? And critically, because you're doing so much less computation for attention, you get a 4x compute reduction overall. Four times the compute reduction. I mean, that is the kind of efficiency gain that fundamentally changes who can afford to run these advanced reasoning models.
8:19Right. That means cheaper API calls, being able to deploy these models on smaller hardware, or serving four times as many users on the same GPU cluster. Exactly. It makes state-of-the-art reasoning far more practical and accessible. Now, what's also important is proving this isn't just a trick for one specific model. How general is this finding? Does it work everywhere? It's highly generalizable. The RWR approach demonstrated consistent performance gains across two major reasoning model families, WEN3 and 5-4 Reasoning Plus. It was also validated on additional architectures like DeepSeekLama8b and Nematron4b.
8:57And across different types of problems. Yes, crucially. It works across diverse domains. Really difficult math reasoning like AIME and HMMT, complex code generation on live code bench, and academic question answering like GPQA. And across all these domains, the performance consistently plateaus at that 50 % cash budget. Which confirms that this redundancy in the middle of a thought process is a general phenomenon of how LLMs reason. It seems to be, yes. Okay, so RWR is defined by its simplicity, keep the beginning, keep the end. But why not try to save the most important middle tokens? Did they compare RWR to more complex methods, maybe something that used attention rates to find the critical steps to preserve?
9:40They absolutely did. they compared RWR to various alternatives, including things like random sampling, strided sampling, and the more sophisticated one. The attention weight scored one. Yes, the attention weight scored middle window, which tries to use the calculated attention weights as a proxy for identifying the most relevant middle tokens to keep. And what was the outcome of that bake-off? They found that RWR, despite its incredible simplicity, it relies purely on, you know, temporal recency for its last window. It performed competitively with and often better than those complex strategies.
10:11That is a massive vindication of the simple approach. It shows that just excluding those intermediate distant tokens is highly effective because they are genuinely just redundant noise. It really is. That's powerful. But let's look closer at that context window, the WF. I know some prior work on long sequences like attention sync used an initial window of just a handful of tokens like four to anchor the generation. RWR needs a much larger first window, right? Why is that difference so critical for reasoning tasks? That is a key finding, and it really validates RWR's design choice for reasoning.
10:45Those previous methods, they used a tiny initial window, mainly just to stabilize the model's ability to generate longer sequences overall. RWR's designers found that for hard reasoning problems, you need a substantially larger first window, often up to 124 tokens. 1 ,024 tokens compared to 4. Why that huge jump? Because for reasoning tasks, you have to preserve the complete problem context, the whole thing. If the model starts forgetting the initial constraints or the full problem statement, its ability to reason logically just collapses immediately. So they tested this theory. They did. They ran what we call ablation studies and found that performance increases significantly as that first window size grows towards 10-ton 24 tokens.
11:28So RWR isn't just solving a memory problem. It's using functional insight. It preserves the critical, persistent context while aggressively pruning the redundant transient steps. So the rolling window reasoner successfully leverages the structural inefficiency in LLM thought chains, those bloated self-corrections and repetitive steps, to get massive efficiency gains. It makes these advanced reasoning models far more practical and accessible, all without any costly fine-tuning. That's the perfect summary. And the final implication here is, I think, pretty profound. We have been so focused on encouraging models to think longer for accuracy.
12:04But the RWR work really proves the long-term path to efficiency isn't just about limiting the length. It's about optimizing the structure of that thought. And if we found such significant, easily discardable redundancy in these intermediate reasoning steps, it raises a really important question for future research. What other structural inefficiencies might exist in complex LLM outputs? You know, in summarization or generation, or even in training data processing. What else could be similarly exploited for efficiency? That's an idea worth pondering the next time you ask an LLM to solve a difficult problem.
From the publisher
This paper studies an inference-time optimization technique designed to reduce the high computational cost of reasoning-optimized large language models (LLMs), which generate long chains of thought. LLMs' self-attention mechanism typically scales quadratically with sequence length, making long reasoning chains prohibitively expensive. RWR addresses this by exploiting the redundancy in intermediate reasoning steps, maintaining only two strategically chosen parts of the key-value (KV) cache: the first window, which holds critical problem context, and the last window, containing the most recent reasoning steps. This simple approach significantly reduces memory and compute requirements, achieving similar accuracy with up to a 50% KV-cache budget reduction, which translates to substantial memory and compute savings across tasks like math reasoning, code generation, and academic question answering, even for models trained with full quadratic attention.




