In short
Recursive language models (RLMs) to overcome “context rot” and extend reliable reasoning from ~200k tokens to 10M+ tokens by externalizing long prompts into a persistent workspace the model can programmatically query.
Guest backgrounds
No guest names or bios appear in the transcript.
Key claims
Standard transformers fail on dense, high-complexity tasks because effective context windows shrink well before physical limits; context rot causes forgetting/made-up details. RLMs maintain or improve quality on huge inputs by writing code to interact with an external variable (P dollars), selectively peeking at raw text, delegating subtasks to sub-model calls, stitching long outputs, and verifying extracted facts.
Notable examples
S-N-I-A-H (needle-in-haystack), Oolong (dense aggregation), Oolong Pairs (quadratic cross-referencing). GPT-5 F1 on Oolong Pairs: 0.04% vs RLM: 58% F1. BrowseComp+ (6–11M tokens across 1,000 docs): RLM ~91.33% correct vs summary-agent 70.47%. Cost ~$1 average, with high variance; latency can be reduced via async subcalls.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Context Rot in Language Models
0:45 to 2:35
Exploring the limits of transformer models and the concept of context rot.
“say 272 ,000 tokens in a frontier model like GMSI-5, the performance starts to degrade long before you even reach that max capacity.”
The Promise of Recursive Language Models (RLMs)
2:35 to 5:10
Introduction to recursive language models and how they aim to overcome context limitations.
“You mentioned that the effective context window is often much shorter than the physical maximum.”
Challenges Faced by Traditional Language Models
5:10 to 7:50
Discussion on task complexities and how standard models fail in difficult scenarios.
“The RLM just sidesteps that problem entirely.”
Core Innovations of RLMs
7:50 to 11:10
How RLMs use innovative techniques to enhance performance and manage larger contexts.
“The results really do demonstrate superior scaling.”
Performance Evaluation and Cost Efficiency of RLMs
11:10 to 14:00
Analyzing the performance and cost effectiveness of RLMs compared to traditional methods.
“What were some of the most compelling patterns you saw in how the RLMs actually wrote that code?”
Challenges and Solutions of RLM Runtimes
14:00 to 14:46
Explore the challenges of RLM runtimes and how they can be optimized.
“Right now, RLM runtimes can be longer because they use sequential or blocking subcalls.”
Implications of RLM for Complex Applications
14:46 to 15:37
Learn how RLMs can now handle complex applications that were previously infeasible.
“It means that the hard context window limit that has plagued language models for years is now functionally solved for a lot of highly complex applications.”
Transcript
Automatic transcript. May contain errors.0:00Okay, let's unpack this. If you've ever tried to use a modern language model for a truly massive task, I mean, like analyzing an entire software repository or reading a thousand research papers at once, you inevitably hit the wall. Oh, you do. It truly is the foundational bottleneck right now. And it frustrates everyone who tries to push the limits of these systems. No matter how large and powerful these underlying transformer models become, they still have this hard physical limit on how much information they can pay attention to at the same time. And those limits are expanding so fast. We've gone from, what, 4 ,000 tokens to well over 200 ,000 in just a couple of years.
0:37But when you get into the real world scale of tens or hundreds of millions of tokens, the model just can't cope. And what's worse, it's not just the physical limit. say 272 ,000 tokens in a frontier model like GMSI-5, the performance starts to degrade long before you even reach that max capacity. We call that phenomenon context rot. Context rot is such a great term for it. It's that feeling when you ask the model a question about something at the very beginning of a long prompt, and it just starts making stuff up. Or it forgets the crucial details you fed it five minutes ago. The quality just plummets as the task gets complicated, even for the best models.
1:15Right. Context route basically means that an otherwise phenomenal model, when you give it an extremely long, dense prompt, it basically acts like a terrible student just skimming a textbook. Yeah. It's overwhelmed and its ability to reason over all the facts just decays dramatically. So if that's the problem, this fundamental memory crisis for AI here is the big reveal that changes the game. Recursive language models. Yeah. Yeah. RLMs. Yes. This is a brilliant general inference strategy designed to dramatically scale the context size by orders of magnitude. We are moving the goalposts entirely here.
1:50We're talking about handling inputs up to 10 million tokens and beyond, which is what? One to two orders of magnitude beyond what any typical context window can currently manage effectively. And the promise isn't just that they can handle these huge inputs. They actually maintain and in many cases dramatically outperform the quality of base LMs and existing long-context strategies. Across really diverse, complex tasks, too. And all while staying surprisingly cost-competitive. So this isn't just like patching a memory leak. It's building a whole new external storage system for the AI. Exactly. And to really appreciate the innovation of our LMs, I think we need to spend a little more time understanding exactly where and why the standard models fail.
2:33Okay, let's dive into that context crisis then. You mentioned that the effective context window is often much shorter than the physical maximum. What do you mean by that? So the physical limit is just, you know, the number of tokens the model's architecture can fit into its self-attention mechanism. Simple as that. But the effective limit is the amount of context the model can reliably reason over before the performance just dips unacceptably. The problem isn't just length. It's also about the complexity of the problem. Right. So it's not just how long the book is. It's how dense the material is.
3:02Right. Precisely. And researchers have broken down task complexity into three categories that really help explain where these LMs start to fail, catastrophically in some cases. Okay, what are they? So the first category is something called S-N-I-A-H, which stands for single needle in a haystack. This is low constant complexity. Needle in a haystack. I think I can guess this one. Yeah, and think of it as looking for one specific sentence, the needle, in a massive 100-page document full of irrelevant text, which is the haystack. Since you're just looking for one thing, frontier models handle this really well.
3:39That makes sense. It's a simple retrieval task, but I'm guessing the problems get much harder from there. They do. Step two is a benchmark called Oolong, which requires dense aggregation. This is linear complexity. Okay, so what does that look like? Imagine you are trying to summarize all the financial transactions across a year-long ledger. The answer depends on pretty much every single line in the prompt. So the model has to process and connect everything to get the final answer. Exactly. A massive step up in difficulty. And then we get to the absolute killer. The task where traditional LMs just completely break down.
4:11Oolong pairs. Oolong pairs. This is quadratic complexity. Here, you aren't just aggregating all the chunks. You had to cross-reference and aggregate pairs of chunks. Like comparing every clause in a contract against every other clause to find a conflict. That's a perfect example. The number of reasoning steps just explodes quadratically as the input grows. And this is where standard LMs fail. Utterly. The data showing this is what really grabbed me. If you look at the performance of a powerful base model like GPT-5 on that Oolong Pairs benchmark, it got an F1 score of just 0.04%. Think about that number.
4:490.04%. It's basically zero. It's nothing. It made virtually no progress at all. The LLM was given a problem that was mathematically impossible for it to solve with its limited working memory. But the RLM, even when it's operating on inputs that are 10 times longer than the base model's physical context window, it maintains strong performance. The difference is night and day. It really shows that the traditional approach of trying to hold millions of tokens in memory all at once is just doomed to context rot. The RLM just sidesteps that problem entirely. Here's where it gets really interesting.
5:24The core innovation of the RLM design. How do you give the LLM a memory that's orders of magnitude larger than its head? The key idea is inspired by some really clever engineering, specifically out of core algorithms. Okay. This comes from data processing, where a computer with a very small, very fast main memory figures out how to manage massive data sets by cleverly fetching data chunks only when they're needed. So let's put that into an image. If the base LLM is like a student who can only hold, say, three textbooks in their hands at once, an out-of-course system is like a top researcher sitting in a massive library, they can still only hold three books, but they have a perfect index and a way to retrieve any of the other 10 ,000 books instantly.
6:04That's a perfect analogy. And the RLM applies that same logic to the prompt itself. The long prompt, the whole library, is not physically fed into the transformer. That is the fundamental break from how we've always done things. So instead of being ingested, what happens to those millions of tokens? They're initialized as a variable, which the researchers call P dollars for prompt, inside an external computing environment. You can think of it like a command line interface or a Python session running alongside the LLM. Wait, so the LLM itself isn't reading the 10 million tokens. The LLM is acting as a programmer.
6:39It's writing code to interact with that variable, with P dollars. Precisely. The LLM writes code to symbolically interact with P dollars. It can write commands to peek into it, decompose it into smaller chunks, or observe what happens when it runs code against it. And this is where the recursive part of the name comes in, right? It's not just accessing the environment. It's running little mini-analyses inside of it. Yes. Yes. The RLM can programmatically create subtasks and then invoke itself or a smaller, faster sub-language model over specific, manageable snippets of that massive prompt variable tie dollars.
7:15It delegates the hard work. Okay, but if the model is constantly writing and executing code, doesn't that massively increase the delay? Or worse, what if it hallucinates bad code that just breaks the whole process? That seems like a huge risk. That's a valid concern, and we will get to the latency part. But the key advantage is that this design allows the LLM to be strategic. Previous methods tried to decompose the task, but they could never scale the input beyond the context window. The RLM externalizes the context, turning memory management into a coding problem the model solves for itself. Okay, Themy is great, but the proof is in the processing.
7:50Let's talk results. The results really do demonstrate superior scaling. RLMs successfully process contexts up to 10 million tokens and beyond, especially on these enormous tasks like BrowseCom+. And just to frame that for everyone listening, 10 million tokens is, what, the equivalent of a pretty big corporate legal library? Or maybe 20 massive novels? That's the scale we're on now. Exactly. And BrowseCom Plus is the Everest of these context tests. It involves multi-hop question answering, meaning the answer is not in just one place over a thousand documents. A thousand documents. Totaling between 6 and 11 million tokens.
8:26And what did the RLM do with that? The performance gap is huge. RLM using GPT-5 nearly solved all the tasks. It got a 91.33 % correct answer rate. 91%. Now compare that to an existing strategy like a summary agent, which tries to condense the context first. That only managed 70.47%. So why does that summarization approach start to fail when the context gets really dense? It fails because summarization is inherently lossy. When you compress a million tokens down to 10 ,000, you lose all the high-fidelity detail you need for cross-referencing. If the one tiny detail the LLM needs is compressed out of existence, the reasoning just breaks down.
9:04But the RLM avoids this because it can just peek at the raw, uncompressed text whenever it needs to. It selectively views the context, exactly. And let's go back to our catastrophic failure case, Oolong Pairs, the one where base GPT-5 scored 0.04 % F1. The RLM version achieved a 58 % F1 score. Wow. That jump just highlights that this strategy gives the base model an emerging capability to handle these challenges that were just impossible before. It's like giving that student, and not just the library, but an army of interns to instantly cross-reference every single book. What's truly fascinating to me is the cost efficiency.
9:42I would just assume all those recursive calls would make the bill explode. That was the big surprise. RLMs generally maintain comparable or even lower average costs than competing methods. How? Because the RLM is strategic. It doesn't waste time and money processing context it doesn't need. It selectively looks at a tiny fraction of the 10 million tokens instead of processing the entire input lossily. So it's about efficiency of query. Absolutely. For instance, on Browse Comp Plus, the RLM had an average cost of about a dollar. That is way cheaper than the theoretical cost of some infinite context model, and in some cases, up to three times cheaper than the summarization baseline.
10:21But there was a caveat in the cost data, right? The variance? Yes. RLMs show high variance in cost because it all depends on the model's trajectory length, basically how many recursive steps it decide that needs to solve the problem. So if it gets stuck on a hard problem. If it runs into a really difficult query and gets into a long loop of checking, cross-referencing, and rechecking, yeah, the cost can spike for those really tough queries. Let's talk about that trajectory. It's not just that this programming environment is there. The model has to be good at using it. That's right. On those dense tasks, the performance difference between the full RLM and an RLM without the recursive subcalls was huge, up to 59 % better.
11:02So the subcalling is what lets it break down the complexity. That's the key. It's like watching a really efficient human reason through a massive problem using code. What were some of the most compelling patterns you saw in how the RLMs actually wrote that code? The ability to filter input is crucial. So instead of sending the whole 10 million token prompt to the sub-LM, the RLM first uses its own knowledge and some code, like a rejects query, to quickly scan the massive context for keywords. So it's looking for festival or beauty pageant inside the huge document first. Exactly. It finds the right paragraph and only then does it launch a sub-LM call over that specific, small, relevant chunk.
11:41This keeps performance high and costs down because the expensive part is only run on a tiny fraction of the data. That's a fundamentally strategic shift. What about decomposition? We saw some really sophisticated decomposition in chunking. For example, on one task, a model called RLM Quinn3Coder was seen chunking the input by new line and then launching thousands of sub-LMs to classify the content line by line. Wow. It's programmatic reasoning replacing raw token ingestion. And for outputs that are longer than the base model's token limit, how does that work? That's where long output stitching comes in.
12:14Since it's operating in that persistent environment, you can store results from multiple calls, say summaries of 10 different documents, in separate variables. I see. Then it just writes a final bit of code to programmatically stitch those variables together into a final answer that's way bigger than the base model could ever produce. And finally, verification. I love that the model essentially fact checks itself. Yes. RLMs often launch sub-LM calls just to verify information they extracted earlier. That second check is an implicit way to avoid context rot, making sure the details are still accurate before committing to a final answer.
12:52Looking ahead, one interesting finding is that RLM is a model agnostic strategy. It worked with both GPT-5 and QUEN-3 coder. But the models themselves acted very differently. They absolutely did. Gwent 3 Coder, for instance, was way less conservative. It would sometimes use hundreds or even thousands of subcalls for tasks that didn't really need it. It was aggressive. So if GPT-5 is the more conservative, experienced lawyer who only asks for analysis when it's absolutely necessary, Gwent 3 Coder is the eager intern who raises their hand for every single question. Precisely. GPT-5 was more cautious, generally using an order of 10 subcalls.
13:29It highlights that while the framework works universally, the model's internal personality really influences its efficiency and cost. And that disparity points directly to the future of training, doesn't it? It does. The current models are often inefficient decision makers because they weren't trained for this. So a major direction for future work is explicitly training LMs to reason as RLMs viewing those trajectories, the code generation, as a form of reasoning that can be optimized. And let's revisit the runtime issue, that latency friction point. Right. Right now, RLM runtimes can be longer because they use sequential or blocking subcalls.
14:07One call has to finish before the next one starts. But that's fixable. It's largely an implementation detail. We're confident that asynchronous implementations running multiple subcalls at the same time could drastically reduce the time it takes, solving that initial concern. Okay, so the high cost is manageable and the slow runtime is fixable. Was there any context where the RLM strategy just wasn't the best choice? Yes. There's a slight performance drop-off for RLMs in the smallest input regimes compared to just calling the base model directly. So there's a crossover point. Exactly. There's a threshold where it becomes more optimal to switch from a simple direct call to firing up the whole RLM framework.
14:45So what does this all mean? It means that the hard context window limit that has plagued language models for years is now functionally solved for a lot of highly complex applications. We've externalized the problem of memory, turning it into a programmable task that the LLM solves for itself. We've effectively given the LLM of a persistent workspace, a massive filing cabinet, and the strategic tools to manage data that is orders of magnitude larger than its immediate working memory. So if an RLM can now reliably process and reason over 10 million tokens, the equivalent of dozens of books, or an entire software code base, what practical application that was previously impossible due to these limits is now suddenly within reach?
15:26Think about complex legal discovery across a massive historical archive, or building a long horizon agent that has to manage context across years of data. We've just unlocked a whole new scale of problem solving.
From the publisher
This paper introduces Recursive Language Models (RLMs), a novel inference strategy designed to overcome the limitations of context windows and the performance degradation of standard large language models. Unlike traditional approaches that feed long prompts directly into a neural network, an RLM treats the input as an external environment within a Python REPL. This allows the model to use code to programmatically examine, decompose, and filter massive datasets that would otherwise exceed its memory capacity. By recursively calling itself on smaller, manageable snippets of the prompt, the system can handle inputs up to two orders of magnitude larger than standard limits. Experimental results using frontier models like GPT-5 show that RLMs significantly outperform existing methods on complex, information-dense tasks while maintaining comparable costs. Ultimately, this framework provides a scalable way for AI to process millions of tokens without losing the fine-grained reasoning capabilities required for deep research and data aggregation.




