In short
Why current LLMs fail on long, multi-step, near-zero-error tasks, and how the “Maker” system using MDAP (Massively Decomposed Agentic Processes) achieves 1,048,575 steps with zero errors.
Guest backgrounds
No guests are identified by name or role in the transcript; it’s a host/interviewer dialogue.
Key claims
Reliability fails due to compounding per-step error and monolithic context/memory limits. Maker fixes this via maximal agentic decomposition (one atomic step per microagent), efficient error correction via dynamic “vote-ahead” using SPRT, and “red flagging” to filter correlated failures.
Notable examples
Towers of Hanoi with 20 disks (exactly 1,048,575 optimal moves). Monolithic SOTA LLMs fail well before 20 disks. Maker uses GPT-4.1 mini microagents as most cost-effective (minimizing cost per token divided by success rate).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI Task Limitations
0:45 to 2:14
Discusses the challenges AI faces in completing long multi-step tasks with low error rates.
“Or maybe a more tangible example for you is the supply chain for something like an iPhone.”
Introduction to Maker System
2:14 to 3:07
Introduction to Maker, a system that successfully executes a million LLM steps with zero errors.
“And what's really critical here, we need to understand, is that this is a totally orthogonal approach to scaling AI.”
MDAP Framework Explained
3:07 to 4:22
Explains the Massively Decomposed Agentic Processes Framework and its components.
“we need to go back to the test bed they used in the sources, which is the Towers of Hanoi challenge.”
Maximal Agentic Decomposition
4:22 to 6:20
Delves into the first key component of MDAP, focusing on task decomposition for error management.
“The approach is conceptually pretty simple, but it's radical in how it's applied.”
Voting Mechanism for Error Correction
6:20 to 8:06
Describes the dynamic voting mechanism that enables efficient error correction in the Maker system.
“And maker's implementation is where the deep math really comes into play.”
Economic Efficiency of Maker
8:06 to 9:00
Discusses the economic implications of the Maker system, highlighting cost efficiency in LLM deployments.
“If your task is 10 times longer, the required voting margin barely increases.”
Empirical Results of Maker System
9:00 to 10:46
Examines the empirical success of the Maker system in solving the Towers of Hanoi task flawlessly.
“And that brings us to the third component, which is this really necessary engineering step, red flagging.”
Future Implications of the MDIP Framework
10:46 to 12:49
Discusses the future of AI reliability and safety through the lens of the MDIP framework.
“The number of undecided steps just decrease exponentially with each voting round.”
Exploring the Limits of Decomposition in AI Tasks
14:01 to 14:38
Discussing the potential of the MDIP framework in managing complex AI tasks.
“It potentially mitigates huge risks associated with a powerful, monolithic superintelligence trying to do unsupervised long-horizon tasks.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we are wrestling with a really fundamental constraint. One that sort of defines the current boundary of artificial intelligence. And that question is, why can't our most advanced AI models reliably complete long multi-step tasks? It really is the silent killer of AI adoption for, you know, for serious real world processes. We hear all this talk about LLMs writing poetry or a bit of code. Sure. But modern societies, they rely on processes with just vast numbers of sequential steps. Yeah. And they require an extremely low, I mean, near zero error rate. Absolutely.
0:37I mean, think about something like the architectural planning for a massive skyscraper. Hundreds of thousands of dependent decisions. Exactly. Or running a major hospital system. Or maybe a more tangible example for you is the supply chain for something like an iPhone. Oh, that's a great one. That relies on contributions from maybe a billion people and just countless logistical microsteps. A single failure point, even a tiny one, in a long chain like that just seizes up the entire process. And that gets us right to the technical problem. It's all about persistent error rates. Right. Even if you have an incredibly competent cutting edge LLM that manages, say, a 99 % success rate per step.
1:18Which sounds fantastic. It sounds great. Only a 1 % error rate. But for long tasks, that is a statistical disaster. After just 100 steps, you're already highly likely to have failed. It's just the mathematics of reliability. You know, current LLMs are benchmarked on these short horizon tasks where they will need to do maybe three or four logical steps. But the moment you apply them to a complex planning problem needing thousands, let alone millions of decisions in a row, failure is basically guaranteed with these monolithic architectures. The problem isn't that the LLM is stupid. The problem is one of compounding unreliability.
1:53We just hit a wall when tasks require high precision across a long sequence. But today that wall has been, well, fundamentally breached. We are diving into a system called Maker, which has successfully solved tasks with over 1 million LLM steps. And with zero errors. With absolutely zero errors. This is just a game changer for reliability and scale. Maker is the first successful implementation of what researchers are now calling the Massively Decomposed Agentec Processes Framework. Yeah. Or MDAP for short. MDAP. And what's really critical here, we need to understand, is that this is a totally orthogonal approach to scaling AI.
2:31Okay, what do you mean by that? We've been pouring, you know, trillions into making these base LLMs bigger and supposedly more intelligent. But MDIP says, stop. Let's focus instead on the system structure, on process decomposition, and on mathematically robust error correction. And that is precisely our mission for you today. To unpack the core components of this maker system and the MDCAPI framework, we're going to look at the three mechanisms that make a million perfect steps possible. Extreme decomposition, some incredibly efficient error correction, and a necessary practical component they call red flagging.
3:04Let's really unpack this technical breakthrough. So to truly appreciate the scale of this solution, we need to go back to the test bed they used in the sources, which is the Towers of Hanoi challenge. Right, the classic recursive puzzle. Exactly. And when we talk about long horizon tasks, we are talking about scaling this up to 20 disks. And a 20-disc task requires exactly 2 to the power of 20 minus 1 optimal steps. Which is? That's 1 ,048 ,575 required actions. It's a perfect measurable proxy for any massive sequential planning problem you'd find in the real world. So what happened when they tasked the best state-of-the-art monolithic LLMs with this problem?
3:45Catastrophic failure. Right. They do okay up to 5 or 6 discs, maybe 7 on a good day. but their success rate just plummets to zero long before they get anywhere near 20. And it's not a gentle decline, is it? No. Their performance doesn't just degrade linearly. It deteriorates dramatically, often exponentially. The LLM just can't handle the memory and the context. Explain that. What's the context burden? Well, think about asking a single agent to plan over a million steps. It has to remember every single move it has ever made, every constraint it has to uphold, it gets clogged. And that leads us directly to the first pillar of MDAP, maximal agentic decomposition.
4:20Or M &S, for sure. MAD. The approach is conceptually pretty simple, but it's radical in how it's applied. You break the task into the smallest possible subtasks. Specifically, where the size M equals exactly one atomic step per agent. So think of it like this. When a single monolithic agent tries to generate, say, the 1 ,000th action. Its context window is just clogged. It's filled with 999 steps of history. And most of that is just irrelevant noise by that point. It gets confused. It loses focus. And, you know, inevitably it violates a constraint. But with MAD, the little microagent responsible for Step 1000 has a tiny, super focused micro role.
5:04Exactly. It only sees the state of the pegs right now. Its entire job is just to confirm one single move. Move disk three from peg A to peg C. That's it. Its context is strictly limited to only the information needed for that one action. This avoids all that confusion from irrelevant built-up context, and it also lets the system use smaller, more efficient LLMs. Now, smart listeners are going to immediately question this. I certainly did when I first read the material. I know exactly what you're going to say. If you break the task into, you know, over a million steps, haven't you just created a million points of failure instead of one long one?
5:35The probability of success still seems like it would decay exponentially. That is the right question to ask. And if you rely on a single call being successful at each of those million points, you are absolutely doomed. But that is where the genius of the framework pivots. The modularity you get from Amadee allows for something that is fundamentally impossible in a single long call. And that is? Robust, efficient error mitigation through voting. This transition from monolithic to modular is what enables the scaling. Error correction isn't just a nice feature. It's absolutely critical for any system that has to persist over a long timeline, whether it's digital or biological.
6:16Right. Think about how mammals fight cancer or repair DNA. Exactly. They're all built around robust error correction to overcome what would otherwise be a linear increase in the probability of failure. And maker's implementation is where the deep math really comes into play. So the second core component is their error correction scheme, which they call first to ahead by voting. This is truly clever. Instead of running a fixed-size poll for every single step, like asking 100 LLMs and picking the majority. Which would be incredibly expensive. Wildly expensive. Instead, this scheme continuously draws independent candidate actions for a subtask until one candidate has been drawn long times more than any other candidate.
6:56So it's not a fixed-size vote. It's dynamic. It's on demand. And it stops as soon as a robust enough consensus is reached. Exactly right. That stopping rule is what makes it so incredibly efficient. It's motivated by something called the Sequential Probability Ratio Test, or SPRT, which is a classical statistical method designed to make highly accurate decisions with the absolute minimum required data. So they only spend the resources they actually need to confirm the right answer. Precisely. And this is where the scaling laws make the entire system economically feasible. The theoretical results, which they borrowed from analysis of the classic gambler's ruin problem.
7:33Which is such a beautiful metaphor for sequential decision making. It really is. It shows this remarkable relationship. Let's break that down because this is really the core insight. To maintain a high probability of overall success, let's say you want a 90 % chance of completing this million-step task perfectly. Okay. The required vote margin, which we call Coramin, it only grows logarithmically with the number of steps. So Coramin is the theta log s. Exactly. Okay, so what does that actually mean in practical terms? It means Maker doesn't have to work exponentially harder as the task gets linearly longer.
8:08Yeah. If your task is 10 times longer, the required voting margin barely increases. Maybe you need four votes ahead instead of three. So if the task gets 1 ,000 times longer. You might only need K-N-L to increase by a factor of three or four, not 1 ,000. And that is why this system can solve a million steps reliably. And the cost implication of that is just immense. Oh, huge. Because coin grows so slowly, the expected cost of the entire maker system scales only, log linearly with the total number of steps. Log linear scaling, that is the holy grail of system design. It is. It means you can manage these astronomically long processes without the cost just exploding into infinity.
8:47This kind of efficiency was, I mean, foundational for the scalability of classical computing. And now they've achieved it for LLM-based agentic processes. But there's a practical side to this. The theory is only as good as the input data. Right. And that brings us to the third component, which is this really necessary engineering step, red flagging. Because the whole voting mechanism, its efficiency depends heavily on the underlying per step success rate, what they call P. Exactly. If P is too low, the system spends way too much money voting. If the LLM is unreliable 50 percent of the time, you'll never reach a consensus sufficiently.
9:24So Maker aggressively tries to increase P by just throwing out outputs that show signs of being unreliable before they even enter the voting pool. And they found two simple, really clear signs of pathological behavior in the LLM microagents. First, overly long response. OK, that's interesting. Yeah. If the LLM takes a simple task, move A to C, and generates three paragraphs of justification, it's a huge sign that the agent is confused and it's overanalyzing the context. It's thinking too hard. And the second sign. Incorrectly formatted responses. If the LLM fails to follow the strict output format, it suggests there was some kind of internal reasoning failure.
10:05And these rules are so crucial because they mitigate what are called correlated errors. They get rid of these specific unreliable samples that would otherwise just drag down the entire voting process. Right. Red flagging is the practical safeguard that makes sure the beautiful theory of logarithmic scaling actually holds up in the real world. Okay, so let's look at the empirical triumph. Maker, using this full stack, the maximal agentic decomposition, the log scaling voting, and the red flagging, it successfully solved the 20-disc towers of Hanoi task. 1 ,048 ,575 optimal steps. Completed with zero errors.
10:39A perfect execution. And what the sources show is that the system behaved exactly as the math predicted it would. The number of undecided steps just decrease exponentially with each voting round. Proving the efficiency of that stopping rule. Absolutely. And critically, the vast majority of the cost, the actual money spent, was in those initial few calls to get that margin K. The cost of resolving the few remaining undecided steps was empirically a rounding error. Now let's talk about the economic insights because this is where that decomposition really, really pays off. Running millions of LLM calls is expensive.
11:16Very. So the key question for the researchers was, which LLM should we actually use for these little microagents? And the decomposition led to a really surprising conclusion. Those state-of-the-art high-cost reasoning models, you don't need them. You absolutely do not need them. Right. Since the task is reduced to a single atomic step, the microagent only needs to be good at following one very simple instruction. The optimal model is the one that minimizes the ratio of C over P. Cost per token divided by success rate. Exactly. You might think the most accurate model always wins, but that's not true if its cost is way too high.
11:51Precisely. Maker found that these highly complex, expensive reasoning models only gave them marginal improvements in the success rate, in P, compared to smaller, faster, much cheaper models. That slight boost in P just didn't offset the massive increase in C. And the result was that a smaller non-reasoning model specifically, they mentioned GPT 4.1 mini, proved to be the most cost-effective choice among the proprietary models they tested. Saving the experiment thousands of dollars compared to models that were technically more powerful, but just completely inefficient for these micro-rolls. This is the practical confirmation of what they call the multi-agent advantage.
12:29It is. They solved a problem that is simply impossible for a monolithic system by creating a framework where the whole is so much greater than the sum of its very, very cheap parts. It just confirms that scaling AI reliability can be achieved through structural engineering and process ingenuity, not just building endlessly bigger and more expensive LLMs. I mean, the whole MDIP framework feels so familiar if you have a software engineering background. Oh, the parallels to microservices architecture are just striking. They're modular, they're scalable, they operate independently. They're explicitly designed for failure.
13:03They expect errors. And natural language is just the communication protocol between them. Modularity enables reliability and efficiency. It's the same principle from classical software systems, now applied to the fundamental substrate of linguistic computation. So what does this all really mean for the future? I mean, the maker system gives us a robust, efficient path to achieving extreme precision and reliability by fundamentally changing how we approach intelligence. Yeah, quite literally smashing intelligence into a million pieces that can then be checked and corrected. And if we connect this back to a societal scale, this has profound implications beyond just efficiency.
13:42It really touches on safety and control. Absolutely. If you break down complex tasks into millions of tiny steps with these limited, clearly defined foci, you strictly limit the LLM's view, its power, and its influence at any given moment. Which allows for a much more effective sandboxing, auditing, and just general control. It potentially mitigates huge risks associated with a powerful, monolithic superintelligence trying to do unsupervised long-horizon tasks. It's a path forward that's defined by control and by transparency. So if this MDIP framework can be extended to handle not just execution, but LLM-based insights.
14:20Which some preliminary results in things like large-digit multiplications suggest is possible. Right. Where the task creation itself is recursively decomposed. What then is the ultimate limit of decomposition for managing and solving these complex, billion-step, real-world societal problems? That, I think, is something truly worth mulling over.
From the publisher
This research paper introduces how we can reliably complete complex, multi-step tasks with zero errors. The core concept is **extreme decomposition** of a task into minimal subtasks handled by focused "microagents," which overcomes the inherent, escalating error rate of monolithic LLMs over long horizons. This modular approach integrates an **efficient error correction** mechanism—specifically, a first-to-ahead-by-$k$ voting scheme—and a process of **red-flagging** unreliable outputs, drastically improving the probability of success. Empirical results on the Towers of Hanoi benchmark demonstrate that MAKER successfully solves a task requiring over one million LLM steps flawlessly, suggesting that MDAPs offer an **orthogonal and scalable path** for AI development beyond merely increasing the size and intelligence of base LLMs. The analysis also provides **cost scaling laws** showing that this framework scales efficiently, with cost increasing only log-linearly with the number of steps, making it an economically viable approach for large-scale applications.




