When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token Analysis

18 Sep 2026 · 22 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How LLM agents’ performance scales with huge test-time compute, why “thinking longer” plateaus, and how to allocate tokens to avoid getting trapped in a “sticky basin” of self-generated context.

Guests

None named in the transcript.

Key claims

Standard benchmarks miss continuous progress; the episode uses “ELO per token” (ELO via Bradley-Terry comparisons) to measure skill growth across token budgets. Agents improve early but suffer severe diminishing returns, eventually underperforming an “independent sampling” baseline (memory-wiped restarts). The “sticky basin hypothesis” explains the failure: long context anchors the model to an initial strategy and attention gets overwhelmed. Humans show “superlinear scaling” on AtCoder AHC012, attributed to continual learning/pivoting.

Notable examples

Polyomino packing inflection point at 38M tokens (Kimi K2.7); splitting 100M tokens into three ~33.3M parallel runs yields +264 ELO vs one 100M run.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Test-Time Strategies

1:01 to 2:18

Dive into the implications of giving AI extensive computation time.

“We have a massive, and I mean massive, stack of new data in front of us today.”

The Measurement Problem in AI Intelligence

2:18 to 4:08

Learn about the challenges of measuring AI progress effectively.

“We really have to start with the measurement problem here.”

The ELO Rating System for AI Outputs

4:08 to 6:04

Discover how ELO ratings are applied to measure AI improvement.

“They borrowed the ELO rating system used for human chess players and applied it to continuous AI outputs.”

The Diminishing Returns of AI Performance

6:04 to 8:19

Examine the plateauing of AI performance with increasing tokens.

“And they set them loose on these incredibly difficult open-ended benchmarks like Frontier CS and Flash InfraBenz with that massive 100 million token budget.”

The Sticky Basin Hypothesis Explained

8:19 to 10:06

Understand the concept of the sticky basin and its implications.

“The research proposes a mechanical explanation for this, which they call the sticky basin hypothesis.”

Human Versus AI in Problem Solving

10:06 to 13:34

Compare how humans and AI perform over extended problem-solving sessions.

“But, you know, if AI agents get hopelessly stuck in their own bad ideas the longer they work, it raises a massive question about us.”

Implications for AI Tools and Strategies

13:34 to 14:00

Discuss strategies to enhance AI performance based on research findings.

“So for you, listening to this right now, how do you apply this?”

Testing AI Strategies and Their Outcomes

14:00 to 14:41

Learn about specialized test time strategies tested to improve AI performance.

“And the researchers actually anticipated that instinct.”

Limitations of Current Frameworks

14:44 to 15:12

Understand why advanced AI frameworks still fail under specific conditions.

“The data shows that frameworks like 8Evolve and TTT Discover give the AI a fantastic early lead.”

Scaling Inflection Point Explained

15:18 to 18:15

Discover the concept of the scaling inflection point and its implications for AI.

“Well, the solution lies in a metric the researchers call the scaling inflection point.”
Show all 13 chapters

Practical Applications of the Allocation Rule

18:16 to 20:04

Learn about the allocation rule and how to maximize AI efficiency in tasks.

“By using the inflection point to create three parallel sessions, the AI gained a plus 264 ELO advantage compared to running one single long 100 million token session.”

Engineering AI Interactions

20:05 to 20:57

Explore how to leverage the inflection point for better AI interactions.

“Okay, let's pull all these threads together.”

Risks of Autonomous AI Agents

20:58 to 22:07

Consider the potential risks of deploying fully autonomous AI agents.

“I want to leave you with a final thought today that builds on this data.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, we've all probably noticed this recently, the newer AI models we use every day. they are, well, they're being designed to kind of think a little longer before they actually spit out an answer. Yeah, exactly. The ones we interact with all the time. Right. You see that little indicator spinning, you know, showing you it's a reasoning through the problem. And, I mean, it makes intuitive sense, right? It does. More time spent computing usually equals a higher quality output. Exactly. Yeah. But, you know, it gets you wondering about the extreme limit of that logic. Like, what happens if you take off the training wheels entirely?

0:33Oh, yeah. What if you give an AI an almost infinite amount of time and computing power to solve a deeply complex problem? I mean, does its intelligence just scale upward indefinitely? It's basically the ultimate test of the scaling laws we've been relying on because, you know, if a 10-second pause makes a model noticeably sharper, we naturally assume, well, a 10-day pause might produce a genius-level breakthrough. Right, a total game-changer. Well, welcome to this deep dive into the research. We have a massive, and I mean massive, stack of new data in front of us today. It is a lot to get through, yeah.

1:07It is. And our mission today is to explore how large language model agents actually scale when you give them massive test time budgets. And when I say massive, the research we are looking at today gave these models up to 100 million tokens per session to solve single problems. Which is, it's just a staggering amount of compute. It really is. Just to contextualize a$100 million token budget for you listening right now, that is roughly the equivalent of having the AI read the entire Harry Potter series about 75 times. Oh, wow. 75 times. Yeah, just to build context and iterate on one single task.

1:45That is insane. So we're going to uncover why giving an AI endless time actually leads to a really bizarre plateau. A very unexpected one, too. Totally. We're going to look at data showing how human problem solvers mathematically crush AI over long time frames and explore the underlying reasons why our brains scale differently. Which is a really fun part of the data. Oh, absolutely. And most importantly for you, we're going to reveal a mathematically proven strategy from this research that you can use to squeeze way more intelligence out of the AI tools you use every single day. Okay, let's unpack this.

2:19We really have to start with the measurement problem here. Right, because tracking intelligence over a 100 million token session sounds, well, incredibly difficult. It is, because before we can ask if an AI gets smarter with all that extra time, we need a reliable way to measure continuous improvement, right? And standard benchmarks in the industry just, they completely fail at this. Yeah, those standard tests usually give us a simple pass-fail metric. Like, we hear about benchmarks like SWE Bench all the time. Right, right. So they give the model a software engineering problem, and at the end of the run, well, the code either compiles and works or it doesn't.

2:56Exactly. And a binary grade at the very end of the process tells us absolutely nothing about the journey. Yeah, it doesn't show the work. Exactly. The research we are diving into focuses on highly complex open-ended tasks. We're talking about things like GPU kernel optimization or open-ended algorithm engineering. So definitely not simple math equations with one correct answer. Not at all. Progress in these areas happens in microscopic iterative increments over a very long period. And I imagine even if we assign a raw numerical score to that progress, raw numbers can be, well, incredibly deceptive.

3:33Oh, totally. Like if an AI is optimizing a system and its efficiency score goes from 50 to 60, that might be a relatively easy jump, right? Yeah, it might just be fixing some obvious bottlenecks. Exactly. But pushing that score from 80 to 81, that tiny single point, could require a massive conceptual breakthrough. Right. The raw numbers just do not capture the actual computational difficulty of the improvement. So how do they fix that? Well, the researchers needed a way to track the AI's cognitive progress minute by minute, literally token by token, across wildly different types of tasks without relying on those deceptive raw scores.

4:07So they solved this by introducing a metric called ELO per token. ELO, like in chess. Exactly like in chess. They borrowed the ELO rating system used for human chess players and applied it to continuous AI outputs. Wait, so we're essentially making the AI play a massive, continuous chess tournament against its own past iterations to see exactly when it gains actual skill. Yes. And what's fascinating here is how they mathematically structured that tournament using something called a Bradley-Terry model. Okay, what is that? It's a probability model that takes discrete win-loss matchups and calculates the hidden underlying true skill of the competitors.

4:47Oh, interesting. So at every single budget checkpoint, say, at 1 million tokens and then again at 2 million tokens, the researchers take the absolute best solution the AI has produced so far. Right. And they pit it against solutions generated at different token budgets and even against entirely different AI models working on the exact same problem. Ah, I see. And the Bradley Terry model figures out the actual probability of one solution beating another, which, I guess, smooths out the weird jumps in raw scores. Precisely. If a solution generated at 10 million tokens reliably beats a solution generated at 5 million tokens, its ELO rating goes up.

5:23Makes perfect sense. And because ELO only calculates relative dominance within a specific task, it puts everything on a universal scale. Oh, wow. So they can compare apples and oranges, basically. Exactly. The researchers can now fairly compare how an AI scales on a computer science puzzle against how it scales on, say, a machine learning architecture task. Oh. It gives us a beautifully precise measuring stick for continuous intelligence. That is so smart. Okay, so now we have the measuring stick. And looking at the data, the researchers really didn't hold back. No, they went all out. They tested some of the most powerful agents currently available, right?

5:58Like Claude Opus 4.8, Codex GPT 5.5, Kimi, and Gemini. Yep, the heavy hitters. And they set them loose on these incredibly difficult open-ended benchmarks like Frontier CS and Flash InfraBenz with that massive 100 million token budget. Right. And the assumption, obviously, is that the ELO rating should just climb higher and higher as the AI thinks deeper and deeper. Well, the data reveals a pretty harsh reality check for test time compute. Oh. Yeah. In the early stages of a session, the agents are highly efficient. I mean, they use their initial context, they draft solutions, they revise their work, and their ELO rating shoots upward.

6:36So they do get smarter at first. They do. They convert those early tokens into better performance at a very impressive rate. But I'm guessing the ELO curve doesn't stay on that trajectory. No, it hits a wall, a massive wall. Really? Yeah. As the session drags on into the millions and then tens of millions of tokens, the marginal gains completely collapse. the AI experiences severe diminishing returns. And to quantify just how bad this collapse is, the research establishes a baseline called the independent sampling reference. Okay, let's define that independent sampling reference, because looking at the research, it seems super central to the findings.

7:12It is. Think of it as the ultimate memory wipe. A memory wipe. Yeah. It represents the ELO gain you would get if you forced the AI to start the problem completely from scratch, over and over again, totally independently, without any memory of its previous attempts. You just keep the best random attempt out of the bunch. The math shows that independent sampling scales linearly with the logarithm of compute. So it basically provides a steady, predictable baseline of brute force guessing. Hold on. Wait a second. You are saying that eventually wiping the AI's memory and forcing it to guess from scratch mathematically outperforms an AI that is carefully revising its own work and building context.

7:53Yes, that is exactly what the data shows. But that goes against everything we understand about learning. I mean, if I wipe my own memory while solving a puzzle, I'm just going to make the same mistakes twice. It feels completely counterintuitive, I know. But the data proves that the sophisticated agentic strategy actually becomes a massive liability. A liability. Yeah, because after a certain token threshold, the agents actually fall below that independent sampling baseline. design. The research proposes a mechanical explanation for this, which they call the sticky basin hypothesis. The sticky basin.

8:26Okay, let's explore the mechanics of that. So imagine the mathematical landscape of a complex problem. Early in a session, the AI makes a few assumptions and commits to a specific high-level strategy. It settles into a basin of thought, or what we'd call a local minimum. As it generates more and more tokens, its context window fills up with thousands of lines of code, complex reasoning and micro adjustments all based on that initial strategy. OK, it's like writing a novel. Oh, that's a good analogy. Yeah. Imagine getting 300 pages into a draft and deep down you realize the core plot is fundamentally broken.

9:00Right. Like the protagonist's motivation makes zero sense. Right, right. The rational move is to throw away those 300 pages and start over. But because you have so much history invested in that draft, you just spend weeks endlessly fixing typos, adjusting commas, and rewriting the same broken scenes over and over. You are literally trapped by your own context. That is a perfect way to describe it. The AI is essentially fixing commas on a structurally doomed novel. The more tokens it spends iterating, the stickier that basin becomes. The model's attention mechanism gets completely overwhelmed by the sheer volume of its own historical reasoning.

9:39It gets lost in the weeds. Exactly. It becomes heavily biased toward its past outputs. It can endlessly optimize within that local basin, but it cannot structurally zoom out and realize it needs to climb out to find a better approach. Which explains why the memory wipe eventually wins. I mean, brute force independent sampling might be super inefficient at first, but it guarantees the AI drops into multiple different basins across the problem landscape rather than getting permanently stuck in one bad one. Precisely. The marginal value of adding more context to a stack AI drops to near zero. That is wild.

10:11But, you know, if AI agents get hopelessly stuck in their own bad ideas the longer they work, it raises a massive question about us. Humans. Yeah, humans. Do human problem solvers hit that same sticky basin if we grind on a problem for too long? Well, to answer that, the researchers set up a direct showdown. A showdown. I love it. They pulled historical data from the AtCoder heuristic contests, specifically a contest called AHC012. Okay. This is a platform where the absolute best human programmers in the world compete on the exact same kinds of open-ended, highly complex coding problems we've been discussing.

10:48So this gives us a true apples-to-apples comparison. Humans versus AI, scored by the exact same judge on the same tasks. Exactly. The researchers took the historical performance trajectories of the top 10 and top 50 human experts, and they plotted them against trajectories of AI agents like GPT 5.6 and OPUS 4.8 working on the same AHC01-2-word contest problems. Right. And they tracked both groups over wall clock time. Okay, so what does the data show when you put them head-to-head in a marathon like that? Well, just like we saw on the previous benchmarks, the AI agents explode out of the gate.

11:21Which makes sense. They compute so fast. Right. They process information way faster than humans and improve rapidly. But within hours, or perhaps a day, their progress curve flattens out. They hit the wall. Yep. They hit their sticky basin and stop making meaningful leaps in performance. While the AI is flatlining, what are the humans doing? This is the crazy part. The human cohorts exhibit what the research calls superlinear scaling. Superlinear. Meaning what, exactly? They don't just improve at a steady, linear rate. They actually accelerate. With the AI, the longer the session goes, the slower it gains ELO.

11:57With the human experts, particularly that top 10 group, the longer the contest goes, the faster they convert their time into ELO gains. That is incredible. Their progress curve bends sharply upward, and they eventually blow right past the state-of-the-art AI agents. So time traps the AI, but time makes us fundamentally better at the problem. What is happening in the human brain that isn't happening in the model's architecture? The research points to a critical architectural limitation in current AI. It's the lack of continual learning. Continual learning. Yeah. When a human works on a complex puzzle for three days, we don't just accumulate a static text log of our past actions.

12:35We build a deep structural understanding of the problem space. We learn the underlying physics of the system we are trying to optimize. We update our internal models. We actually learn. Crucially, because we learn, we can recognize when we are in a sticky basin. A human expert can look at a massive block of work, realize the foundational strategy is a total dead end, and completely abandon it for a radically new approach. We throw away the 300-page draft. Exactly. But we keep the structural insights we gain while writing it. And the AI can't do that because its weights are completely frozen during inference.

13:11Exactly. Current LLM agents do not backpropagate or update their core neural pathways during a test time session. All they have is their growing context window, which acts as an anchor. It weighs them down. Yes. The human ability to continually learn, update our mental models, and pivot away from dead ends is exactly why top experts dominate AI over long time horizons. Okay, so we know AI lacks human continual learning, and we know it gets mathematically trapped in a sticky basin of its own context. So for you, listening to this right now, how do you apply this? Right, because we all want better results.

13:48Exactly. We all use these tools, and we want to prevent them from hitting that wall. I mean, the natural instinct is to try and fix the AI's reasoning with better prompts or, like, specialized frameworks. And the researchers actually anticipated that instinct. They tested highly specialized test time strategies designed specifically to force the AI to think better. Like what? Well, they looked at frameworks like AdaVolv, which uses genetic algorithms to constantly mutate candidate programs, and they also tested test time training, or TTT Discover. Wait, TTT Discover actually updates the model's parameters during evaluation, right?

14:22Yeah. I mean, that sounds like it should solve the frozen weight problem. You would think so. TTT Discover runs tiny gradient descent updates during the test to adjust the weights based on the ongoing task. So they threw the most advanced architectural hacks at the problem to see if they could break the AI out of the sticky basin. They did. Did those frameworks stop the AI from hitting the wall? They ultimately failed. Really? Even updating the weights? Even then. The data shows that frameworks like 8Evolve and TTT Discover give the AI a fantastic early lead. They boost efficiency right out of the gate, but they still operate within the gravitational pull of the initial problem formulation.

15:02Eventually, their ELO per token curves bend downward and fall below the independent sampling baseline. Unbelievable. Yeah, they might hit a slightly higher sticky basin, but they still hit it. So if complex frameworks and mutating algorithms don't work, what is the actual solution here? How do we mathematically bypass this limitation? Well, the solution lies in a metric the researchers call the scaling inflection point. Okay, let's define the mechanics of the scaling inflection point. It is the exact mathematical moment in a session where an AI's progress dips below the efficiency of that independent sampling baseline.

15:37The exact moment it gets worse than guessing. Yes. It's the precise token count where the model's attention mechanism becomes too overwhelmed by its own history. At this point, extending the current conversation yields less intelligence than just starting completely over. And the researchers found they could actually pinpoint this exact moment for different models and tasks. They did. Let's look at the specific data they gathered on a special reasoning puzzle called polyomino packing. What's that? The objective is to fit a series of complex geometric shapes into a constrained grid, sort of like Tetris on steroids.

16:12Okay, gotcha. They tested this using the Kimi K2.7 model. And what was the inflection point for Kimi K2.7 on that specific puzzle? The scaling inflection point was exactly 38 million tokens. 38 million. Yep. Before 38 million tokens, the model was navigating its context well. It was iterating on the shapes and gaining ELO fast. Doing great. But around 38 million tokens, the ratio of useful new insights to thousands of past failed permutations just flipped. The attention mechanism couldn't fister the noise anymore, and its efficiency plummeted. Okay, let's apply that practically. Say I have a budget of 100 million tokens I can spend on this spatial puzzle.

16:50My natural instinct is to just open one chat window, drop the prompt in, and let it grind away for all 100 million tokens. Which the research proves is a massive waste of computational resources. Because it hits the wall at 38 million. Exactly. Because we know the inflection point for this task is 38 million, we use what the researchers call the allocation rule. Okay. But wait, if 100 million tokens in one window is bad, my next thought is to be super safe and run 10 completely separate 10 million token sessions. Just wipe the memory constantly so it never gets stuck. Right. But that approach actually swings too far in the opposite direction.

17:27Oh, really? Yeah. 10 million tokens just isn't enough time for the AI to utilize its context effectively and iterate on a deep solution. You are cutting it off way before it reaches its peak efficiency. Ah, I see. So the allocation rule dictates finding the Goldilocks zone. Instead of one long session, or 10 tiny ones, you divide your 100 million tokens into three completely independent, parallel sessions of roughly 33.3 million tokens each. Oh, you set the budget just under the 38 million token inflection point. Exactly. You let the AI run right up to the edge of the sticky basin. It maximizes its early, highly efficient gains, and then you brutally cut it off before the context bloat traps it.

18:09That is brilliant. And you run that process three times in parallel from scratch, and then you simply take the best result of the three. The logic there is flawless, but does the math actually reflect an increase in intelligence? Oh, the performance jump is massive. By using the inflection point to create three parallel sessions, the AI gained a plus 264 ELO advantage compared to running one single long 100 million token session. Wow. Furthermore, the three session split beat the 10 session split by plus 355 ELO. So you gain hundreds of ELO points not by changing the prompt or updating the model, but purely by changing how you allocate the compute relative to the infection point.

18:47Exactly. It completely shifts the paradigm of interacting with AI. It treats the model less like a single human you are endlessly mentoring and more like a statistical resource that must be allocated efficiently to avoid local minima. So what does this all mean for you, the listener? It's a big takeaway. It is. The actionable takeaway here goes against how most of us naturally use these tools. You need to stop arguing endlessly in a single chat window when the AI gets stuck. Stop forcing it. Right. If you're working on a complex coding architecture or analyzing a dense data set and you've been going back and forth for 45 minutes, the AI is not going to suddenly have a massive breakthrough.

19:26It's just not. It is mathematically trapped in a sticky basin of your conversation history. You need to develop a feel for its inflection point. You know, the moment the answers start feeling circular or hyper fixated on minor decent. So we've all seen that happen. Oh, all the time. When that happens, you need to ruthlessly wipe its memory. Start three fresh chats. Feed it the core problem again, perhaps phrase the initial constraint slightly differently, and force it to attack the problem from three parallel angles. The data guarantees you will extract a higher level of intelligence than just beating a dead horse in one long thread.

20:03It really forces us to engineer the AI's environment so that its inherent limitations don't become structural failures. Perfectly said. Okay, let's pull all these threads together. We started by exploring why standard pass-fail tests fail to capture continuous progress. Right. And how the Bradley-Terry model allows us to track ELO per token across massive compute budgets. And then we examined the hard data showing that without true continual learning and the ability to update internal weights at inference time, AI agents inevitably hit a wall. They get trapped by their own historical context. Exactly, while human experts leverage superlinear scaling to accelerate past them.

20:41And finally, we broke down the hack. By identifying a task's scaling inflection point, we can strategically segment our compute budget, wipe the AI's memory just before it gets stuck, and run paralyzed sessions to completely bypass the sticky basin. It is a fundamentally more mathematical approach to prompt engineering. It really is. I want to leave you with a final thought today that builds on this data. And honestly, it's a bit chilling when you consider the current trajectory of the tech industry. Oh, absolutely. Right now, there is an immense push toward deploying fully autonomous agents. The dream being sold is that we will soon give an AI the keys to run a corporation, manage global supply chains, or oversee critical infrastructure for months or even years at a time without any human intervention.

21:28It's the holy grail for a lot of companies right now. But if data we just explored proves that agents inevitably, mathematically, get trapped in a sticky basin of their own initial assumptions without the capacity for true continual learning, what happens when an AI CEO gets stubbornly stuck on a fundamentally flawed business strategy for an entire year? That is a scary thought. If it lacks the human ability to realize its foundation is broken and instead spends 12 months tirelessly optimizing the commas on a doomed corporate plan, what is the real world fallout? It's a structural vulnerability we need to think very carefully about before we hand over the keys to the economy.

22:07Because when you scale a sticky basin up to the enterprise level, the consequences scale right along with it. Something to keep in mind. Well, that is all the time we have for today. Thank you so much for joining us on this deep dive into the research. Keep questioning the data. Keep experimenting with how you allocate your own AI workflows. And above all, stay curious. We'll catch you next time.

From the publisher

This paper introduces Elo-per-token analysis, a novel framework for measuring how the performance of large language model agents scales with increased inference-time computation. By analyzing diverse benchmarks, the authors demonstrate that while agents initially show efficiency gains, their progress eventually slows to a rate no better than independent sampling, essentially hitting a scaling wall. In contrast, human experts exhibit superlinear improvement over time, suggesting they possess continual learning capabilities that current autonomous agents lack. The study identifies a scaling inflection point, which marks the specific budget where extending a single agent session becomes less effective than starting a new one. Utilizing this metric, the researchers developed an allocation rule that optimizes performance by splitting large token budgets across multiple parallel sessions. This strategy significantly boosts results on complex tasks, providing a practical method for managing computational resources in agentic workflows.

More from Best AI papers explained

All 475 episodes
When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token AnalysisBest AI papers explained · 22 min
Listen in VO