In short
The episode explains why “long chain of thought” (long-CoT) training is so expensive for LLMs (quadratic compute from transformer attention over growing context, plus linear KV-cache memory growth), and how Delathink reduces scaling by using “Markovian thinking” during reinforcement learning.
Guest backgrounds
No guests are mentioned in the transcript.
Key claims
Standard long-CoT defines model state as prompt + all prior reasoning tokens, making attention cost O(n^2) and KV-cache fill memory; Delathink keeps context fixed by chunking reasoning (e.g., 8K tokens) and resetting attention, while requiring the model to write a compact “carryover” summary (e.g., last 4K tokens).
Notable examples
Cost estimate drops from ~27 H100-months (96K tokens) to ~7 H100-months; Delathink matches/surpasses Long-CoT on AME/HMMT with faster steps (215s vs 248.5s) and ~40% higher token rate; it improves at test time up to ~140K tokens and averages ~42K tokens after 96K training budget; “zero-shot Markovian thinking” appears in models like R1-Distil-1.5B and larger (e.g., GPT-OSS-120B, Qwen-330B).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Cost of Deep Reasoning
0:45 to 4:10
Exploration of the high costs associated with long chains of reasoning in AI models.
“And the economics here are just staggering.”
Challenges with Current AI Training
4:10 to 6:40
Discussion on the limitations of current training methods for AI regarding long reasoning chains.
“And what's crucial is they didn't change the model architecture itself, not the transformer.”
Introducing Delathink and Markovian Thinking
6:40 to 9:20
An overview of Delathink's approach to overcoming the computational bottlenecks in AI.
“Delathink drops that astronomical number down to just seven H100 months for, you know, startups, academic labs, even big research teams.”
The Mechanics of Delathink
9:20 to 12:20
Explanation of how Delathink operates with structured amnesia and efficient note-taking.
“it's an artifact of how we were training it.”
Implications of Markovian Thinking
12:20 to 14:00
Discussion on the implications of successful Markovian thinking in future AI models.
“Okay, so if we pull all this together, we've seen that the quadratic staling wall of Longcott, it isn't necessarily an intrinsic property of deep reasoning itself.”
Revolutionizing AI Thinking
14:00 to 14:13
Discover how improved note-taking can transform AI training methods.
“All because we figured out how to teach them to take better notes along the way.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the deep dive. So imagine your AI gets really smart, right? Smart enough to tackle complex problems, follow, like, a long chain of logic. But here's the catch. Every extra step it takes in that thought process, well, the cost doesn't just double, it quadruples. We're diving into a really fundamental barrier in advanced AI today. It's this runaway expense of deep, multi-step reasoning. Researchers call it long chain of thought or long cot. Yeah, it's one of the most stubborn computational bottlenecks we face. Our sources are pretty clear. The standard way we train this kind of deep reasoning, it just hits a wall.
0:36It makes truly long, sustained thinking, you know, prohibitively expensive. So our mission today is to unpack this really interesting system called Delathink. It proposes a, well, pretty radical solution using something called Markovian thinking to try and break that quadratic scaling barrier. And the economics here are just staggering. I mean, this is why you really need to pay attention. The current standard method, if you want to train a model to think for a decent amount of time, say average around 96 ,000 tokens, the estimate is 27 H100 months. That is a massive amount of high end GPU time.
1:07It's huge. But this DellaThink approach we're looking at, it promises to slash that cost down to just seven H100 months. Exactly. That's almost four times more efficient just by changing the, well, the rules of the game. Right. And that cost difference, it only really clicks once you understand why the standard method is so darn expensive. We probably need to start with the baseline, you know, the long chain of thought reinforcement learning environment, long-caught TRO. OK, so standard practice. We're training the motor step by step, using reinforcement learning to reward it when it gets things right.
1:40What's the fatal flaw in how long-caught TRO handles the context, the information it sees? The fatal flaw fundamentally is how the environment defines the model's state. So when the LLM generates its next token, its next step in reasoning. The standard setup says the state has to be the original prompt plus every single reasoning token is generated up to that point. So the context the model actually reads to decide its next word. It just keeps growing and growing. Exactly. Linearly. With every single token. And the problem is, for basically all the large language models we use today, the ones based on the transformer architecture, well, that self-attention mechanism, it has to compare every new token against all the previous tokens.
2:22Right, the attention mechanism. So that linear growth in context size, it translates directly into this catastrophic quadratic growth in compute cost, O of n squared. Okay, let's make that concrete for everyone listening. What does O-entopny actually mean if you're trying to build or run one of these? It means the math is just punishing, really. If you decide, okay, I need my model to think twice as long, you scale the fault length by two. Your compute cost doesn't double. It goes up by two squared. It quadruples. Wow. Scale the thinking length by 10. The cost jumps by a factor of 100. That's precisely why training for that, you know, average 96 ,000 thinking tokens balloons to that 27 H100 month figure we mentioned.
3:01So it's baked into the math of the attention mechanism itself. But our sources also highlight another problem, right, tied to this growing context, memory. That's right, the memory constraint. It comes from the key value cache, the KV cache. This cache stores activations, these key and value vectors, for all the previous tokens. It lets the model generate the next token faster without recomputing everything. But since the sequence length is growing linearly, well, the KV cache grows linearly too. Okay. And what's the practical impact of that KV cache getting bigger and bigger? Well, the sources gave a pretty chilling example, actually.
3:37Even a relatively smaller reasoning model generating just one single trace of thought, about a million tokens long, that alone fills an entire H100 GPU just with this KV cache. An entire H100 for one thought. Yeah. So for anyone trying to run these in production, that linear memory growth, it just tanks your throughput. You can't easily run requests in parallel. You end up needing specialized hardware complex setups like sequence parallelism just to handle maybe one user's deep thought process long cod is basically choked by both compute and memory so long cut is kind of fighting physics here okay let's pivot then the proposed solution markovian thinking the core idea as i understand it is to decouple the total length of thought from the size of the context the model sees in any given moment basically make sure the policy always looks at a constant limited-sized state.
4:28That's the paradigm shift, exactly. And what's crucial is they didn't change the model architecture itself, not the transformer. They changed the environment, the rules of the reinforcement learning game, the Markov decision process, or MDP. And Delathink is the system that actually implements this idea. It defines a new way the model interacts with its own history. Okay, walk us through the Delathink mechanism, because honestly, it sounds deeply counterintuitive. You're essentially saying, we give the AI structured amnesia. Is that fair? Well, structured amnesia is a pretty good way to put it, yeah.
4:59So reasoning gets broken down into these fixed-sized chunks. Let's say, for example, 8 ,000 tokens per chunk. When the model reaches the end of that chunk, the 8 ,000th token, the environment performs a hard reset. It literally clears the context, wipes the attention cache, everything the model saw in that chunk, gone. Hold on, hold on. So if I'm like halfway through solving a really complex math proof or figuring out some intricate logic puzzle and you just wipe my memory every few thousand steps, how do I possibly continue? I mean, I need that previous work. Ah, but that's where the textual carryover comes in, the Markovian state.
5:35It's not just a brute force reset. When the context is clear, the environment reinitializes the prompt. It gives the model back the original query, plus a short textual summary or a carryover from the very end of the chunk that just finished. Typically, they use the last half of the chunk. So if the chunk was 8K, it might carry over the last 4 ,000 tokens. I see. So the model itself has to learn through the RL process to be its own efficient note taker. It has to figure out how to explicitly write down just the critical stuff, the intermediate results it needs to continue, knowing everything else is about to get deleted.
6:05Precisely. It learns to craft a sufficient textual state to bridge the gap. And because the context size is now fixed by the chunk size, say, that 8 ,000 tokens, it never grows, no matter how long the total reasoning process gets. Okay, the practical consequence there is, well, it's profound, isn't it? The total training cost now scales linearly with the total thinking length, not quadratically. Linear scaling is manageable. Quadratic scaling is, well, potentially bankrupting. And critically, memory usage stays constant, flat, fixed by that chunk size. It solves both the compute and the memory bottleneck simultaneously.
6:40Remember that 27 H100 months for 96K tokens under long code? Yeah. Yeah. Delathink drops that astronomical number down to just seven H100 months for, you know, startups, academic labs, even big research teams. That difference is enormous. It's a difference between maybe running an experiment once versus iterating dozens of times. Right. That's huge. OK, now the proof is in the pudding, as they say, the empirical results. Did this budget friendly Delathink approach compromise performance? Did it actually work as well? Well, according to the research, not at all. They benchmark Delathink against that standard Longco TRL baseline.
7:14They use the R1 Distil 1.5b model. And Delathink actually matched and in many cases surpassed the Longco TRL performance on tough math benchmarks like AME and HMMT. Even though Delathink was operating in those constrained 8K chunks, while the baseline had a bigger 24K context budget. Exactly. Despite the constraint, it performed as well or better. And the efficiency numbers, are they faster too? Yep, faster too. Delathink completed each RL step quicker, something like 215 seconds versus 248.5 for the Longco TRL 24k baseline. And its token generation rate was also significantly higher, about 40 % faster per H100.
7:52Okay, but here's what seems like the really killer advantage they talk about, test time scaling. Ah, yes. This is critical. So Longco TRL models, they're fundamentally capped by their training budget. If you trained it with a 24 ,000 token context limit, it hits a performance plateau pretty quickly if you ask it to think much longer than that during inference, during actual use. Because it never learned how to manage information flow beyond that window. It just learned to look back at everything within that 24K limit. Right. It relies on having everything visible. Dell Think, however, because it was forced to learn that skill of bounded information passing, that efficient note-taking, The research showed it continuing to improve dramatically when allowed to think for longer at test time.
8:34It successfully scaled to solve problems that required it to think for up to 140 ,000 tokens. Whoa, 140K. That's way beyond its, what, 8K chunk size or even the 24K comparison budget. Far, far beyond his training regime's context limit. It learned a generalizable skill. That really is the aha moment, isn't it? It learned a portable, fundamental, skill-efficient summary and transition, not just how to operate within a fixed, large context window. And they even pushed it further, right? Trained Delathink up to a 96K total budget. They did. Yeah. And that version achieved solutions averaging, I think, up to 42 ,000 tokens on really difficult map competition problems.
9:15It's really compelling evidence that this quadratic barrier, maybe it isn't some immutable law of deep reasoning itself, it's an artifact of how we were training it. Which brings us back to that earlier question, why doesn't deleting the context cripple the model? Why does dilithing work so well? It feels like it should break things. Yeah, you'd intuitively think forcing this kind of amnesia would require the model to completely relearn everything, but apparently not. So what's the secret? The key insight, according to the analysis, is that these LLMs, they already have this ability sort of latent within them from their pre-training.
9:48They looked at a range of models from the small R1 to still 1.5b all the way up to giants like GPT-OSS-120b and QUIN-330b. And they found these models exhibit strong zero-shot Markovian thinking. Okay, zero-shot here means they took models that had never seen the Dullifink training process before, subjected them to the chunking and the context resets, and the models just figured it out. They performed almost as well as they did with the full context. Exactly, or at least recovered a very significant portion of their standard full context performance right out of the box. The behavior seems to emerge somewhat naturally, and that's huge for training.
10:26It means when they start the actual Delafink RL training, the policy isn't starting from ground zero, trying to learn a totally alien skill. It's basically refining and perfecting an existing powerful capability. This ensures the RL process starts with good, positive examples of the behavior they want, which makes training much faster and more stable. Huh. It's like finding out your star runner already knows the basics of swimming. You just need to give them some pointers and structured practice in the pool. But they needed to really push this, right? Stress test it. What about tasks where you genuinely need live memory, not just, say, intermediate math results?
11:02Right. They considered that. They tested it on tasks like crossword bench. Solving a big crossword puzzle definitely seems like it requires continuous access to what you've filled in elsewhere. Deleting context in a, say, 14 by 14 puzzle sounds like a recipe for disaster. Yeah, I can't imagine doing one like that. But even there, the model still managed to find a, quote, non-trivial number of valid Markovian solutions. They could still solve parts of it correctly, despite the resets. That's pretty robust then. So why do these models have this, like, hidden talent for Markovian thinking? Where does it come from?
11:37Well, the hypothesis put forward is that it stems from the massive amount of human text they're trained on. Think about how humans do complex, long-form reasoning. When someone writes a 100-page report or a detailed legal argument, they don't constantly reread every single page from the beginning. We learn to write down key intermediate steps, use headings, create sunneries. we explicitly carry forward the necessary information in our writing. Human reasoning, especially complex written reasoning, is often inherently Markovian in practice. Ah, so the LLMs are essentially mirroring the efficient, bounded, note-taking strategies they learn from observing how humans structure complex thoughts in text.
12:17That seems to be the most plausible explanation. They're mimicking our own strategies for managing cognitive load during deep thinking. Okay, so if we pull all this together, we've seen that the quadratic staling wall of Longcott, it isn't necessarily an intrinsic property of deep reasoning itself. By redesigning the environment, by forcing the LLM to learn this bounded, explicit information transfer, Delathink achieves linear scaling for training these long reasoning chains. Exactly. And for you, the lister, whether you're researching, building, or looking to deploy complex reasoning systems, this is a really massive operational insight.
12:52It strongly suggests that for pure reasoning tasks, unlike maybe long context retrieval where you need the whole document, you don't necessarily need the full unlimited history visible at every single step. You might be able to get much longer, more reliable, and crucially, far cheaper reasoning paths just by changing the data flow, the training environment, without even touching the underlying model architecture. And that leads us to our final provocative thought for today. The success. It seems to provide some pretty strong empirical support for the future viability of non-quadratic LLM architectures, specifically for reasoning tasks.
13:29Right. If Markovian thinking works well, if the entire history doesn't need to be instantly accessible via attention at every single moment, then it really opens the door for models with linear time complexity. Architectures like state space models or things like Mamba, these might be perfectly suited, maybe better suited for the next generation of truly deep reasoning LLMs. So the era of attention is all you need for reasoning might, might actually be facing a serious challenge. We could soon see models that can genuinely think efficiently across millions, maybe even billions of tokens. All because we figured out how to teach them to take better notes along the way.
14:05A potential revolution in how AI thinks. Driven by a, well, relatively simple but brilliant change in the training rules. Fascinating stuff. That's a wrap on this deep dive.
From the publisher
Reinforcement learning (RL) methods for training Large Language Models (LLMs) to produce long chains of thought (LongCoT) are constrained by the standard thinking environment, where the state is unbounded, leading to quadratic computational costs as reasoning length increases. This paper propose Markovian Thinking, a paradigm where the reasoning policy conditions only on a constant-size state, effectively decoupling thinking length from context size and yielding linear compute and constant memory benefits. This concept is instantiated with Delethink, an RL environment that organizes reasoning into fixed-size chunks. At the chunk boundaries, the context resets, and the policy must learn to write a concise textual carryover sufficient for seamless continuation of the reasoning process in the next chunk. Models trained with Delethink, such as R1-Distill 1.5B, match or exceed the performance of LongCoT-RL while significantly reducing computational overhead, demonstrating superior test-time scaling capability far beyond their training budget. They emphasize that redesigning the thinking environment is a powerful lever for achieving efficient and scalable reasoning in LLMs.




