Towards a Science of Scaling Agent Systems / Google Deepmind

15 Dec 2025 · 16 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

DeepMind research argues there’s no universal “more agents is better” rule; multi-agent LLM performance depends on task structure and coordination topology, with major cost/fragility tradeoffs.

Guests

No specific guests are named in the transcript; it’s a Deep Dive discussion with the host(s) summarizing the study.

Key claims

Benefits are highly task-contingent; adding agents can either boost or catastrophically harm. Agentic tasks require sustained multi-step interaction, iterative information gathering, and strategy refinement. Central finding: centralized coordination can greatly help decomposable tasks, but sequential tasks suffer due to fragmentation and coordination overhead.

Notable examples

Finance Agent benchmark: centralized +80.9%, decentralized +74.5%. PlanCraft (sequential): independent MAS −70% performance. Workbench: ~wash (−11% to +6%). Browse Comp Plus: decentralized +9.2%. Predictive principles: capability saturation (~45% single-agent accuracy threshold), tool coordination tradeoff (worst negative effect; tool-heavy tasks suffer), and topology-dependent error amplification (independent 17.2x vs centralized 4.4x). Costs: 1.6–6.2x token budget; reasoning turns scale superlinearly (~agents^1.724); hybrid drops efficiency to 13.6 successes/1k tokens vs single-agent 67.7. Model biases: Google more cross-architecture efficient; Anthropic stable mainly under centralized; OpenAI best with hybrid. Hybrid also had highest coordination failures (12.4%).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Multi-Agent Systems

0:45 to 1:54

Exploring the concept of multi-agent systems and their inherent assumptions.

“In human teams, specialization and collaboration, they usually win.”

Defining Agentic Tasks

1:54 to 3:11

Identifying the characteristics that distinguish agentic tasks from simpler tasks.

“When we talk about an LLM agent, what really sets it apart from, say, a simple chatbot?”

Architectures for Task Management

3:11 to 4:07

Examining different architectural approaches for multi-agent systems.

“which, for listeners maybe not so deep in the weeds, are basically static knowledge quizzes.”

Context Integration vs Fragmentation

4:07 to 5:07

Discussing the trade-off between context integration and fragmentation in agent systems.

“It combines that hierarchy with lateral communication.”

Benchmark Performance Analysis

5:07 to 7:19

Analyzing various benchmarks to illustrate the performance of multi-agent versus single-agent systems.

“Let's start with the huge success story, the Finance Agent Benchmark.”

Dynamics of Workplace Tasks

7:19 to 7:58

Discussing the impact of multi-agent structures on typical workplace tasks.

“Workbench has realistic office tasks, code execution, heavy tool use, and it showed really marginal effects, ranging from a slight 11 % degradation to a small 6 % game, basically a wash.”

Principles of Team Coordination

7:58 to 10:40

Introducing three major principles that affect the performance of multi-agent systems.

“This all leads directly to, I think, the core innovation of this research, the predictive framework.”

Cost of Coordination in Agent Systems

10:40 to 13:19

Evaluating the costs associated with multi-agent coordination and its implications.

“That centralization contains the error amplification to only 4.4 fold.”

Adapting to LLM Family Differences

13:19 to 14:00

Exploring differences in performance and efficiency across major LLM families.

“It suggests their models are really optimized for depth of reasoning in a single coherent stream.”

Understanding Coordination in Advanced Agent Systems

14:00 to 15:40

Explore the key factors influencing the effectiveness of agent systems in AI.

“What is the core takeaway for anyone building or deploying these advanced agent systems?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. So if you've been following the massive shift in AI over, what, the past two years, the focus has really moved beyond just single-shot prompts. Oh, completely. We're now building these complex LLM-based agents, you know, systems that can reason and plan and actually act based on feedback from their environment. And this whole rush to build agents has been dominated by one really powerful but I'd say unproven assumption. Okay. It's the idea that multi-agent systems or MAS, so a team of LLMs, will just inherently be better than a single agent system or a SAS. Right.

0:38The more is better idea. Informally, yeah. It's the belief that more agents is all you need. Which, I mean, that feels right, doesn't it? In human teams, specialization and collaboration, they usually win. A team of experts should always beat one person trying to juggle everything at once. It's intuitive. I'll give you that. But until now, it's really just been a heuristic, a best guess. It's a hunch. Exactly. So our mission for this deep dive is to move past all that hand waving and establish a principled quantitative scaling science for these systems. We've synthesized research from this huge controlled study, 180 different configurations.

1:14180. Yeah. Across the three big LLM families, OpenAI, Google, and Anthropic. We're here to quantify when adding a teammate is a superpower and when it's, well, an anchor just dragging the whole system down. So we're getting past the hype and into the actual math of team design. I love it. Precisely. And so here's the central finding right up front for you. The benefits of adding more agents are highly, highly task contingent. They are not universal. Performance is entirely dictated by how well your team structure, your coordination topology, matches the problem you're trying to solve. Adding an agent is just as likely to severely harm performance as it is to help.

1:51Okay, let's unpack that. We need some foundational definitions first. When we talk about an LLM agent, what really sets it apart from, say, a simple chatbot? It's all about its operational cycle. It's not just producing a static output from a single prompt. It goes through these iterative loops of observing its environment, reasoning about what it sees, making a multi-step plan, and then, you know, executing an action. And this is the key. It adapts its behavior based on the feedback it gets. So that sustained interaction is the real differentiator. You can't just judge these systems on how they do on a multiple choice test.

2:26You absolutely can't. And that leads to this critical first step. You have to distinguish between agentic and non-agentic tasks. Okay. Agentic tasks, they need three things. First, sustained multi-step interactions. Second, they have to gather information iteratively because they don't have all the facts up front. They're working with partial information. Exactly. And third, they have to be able to refine their strategy when things don't go as planned. So give us some real-world context for that. What kind of tasks are we talking about? Think about dynamic, real-time stuff. Financial trading, maybe, with constant data streams and re-evaluation.

3:03Or web browsing, where you have to click through pages and navigate. Or even complex planning in a game. Right. These are all worlds apart from things like MMLU or Human Evol, which, for listeners maybe not so deep in the weeds, are basically static knowledge quizzes. They are. They rely on single-shot reasoning. or just recalling a fact, the stakes are so much higher in agentic tasks because arrays don't just happen once. They cascade through the entire chain of actions. Got it. Okay. So once we've established a task that's truly agentic, this study tested five main architectures. You have the benchmark, the single agent system, or SAS.

3:39What were the four multi-agent or MAS variants they looked at? So the MAS architectures were all about different ways of coordinating. You had independent agents who just work in parallel and then sort of dump their results at the end. No talking. Then centralized, which has one powerful orchestrator acting like a project manager. Decentralized, which uses more of a peer-to-peer discussion or debate. And then hybrid, which tries to get the best of both. Like a central manager, but also letting the team members chat on the side. You got it. It combines that hierarchy with lateral communication.

4:13And conceptually, the real tension here between the single agent and the team, it comes down to a fundamental trade-off. It's context integration versus fragmentation. That is the heart of the matter. A single-agent system, it maximizes context integration. It has one unified memory, one stream of thought. Everything is done is right there. Total global context. Total global context. But multi-agent systems, by their very nature, they impose information fragmentation. You get parallelism, which is great, but that rich global context has to be compressed, filtered, and sent in these lossy messages between agents.

4:46And that communication has an unavoidable cost. And that's the cost we see reflected in the results, which were, I mean, all over the place. Looking at the four benchmarks, Finance Agent, Browse Comp Plus, PlanCraft, and Workbench, the performance wasn't just variable. It was polarized, massive success, or total meltdown. Absolutely. We have two perfect examples of this. Let's start with the huge success story, the Finance Agent Benchmark. This is a task that involves complex financial reasoning. Like assessing a company's health for investment. Exactly. And what was the performance boost there for the multi-agent approach?

5:23I'm ready for it. Centralized coordination improved performance by a staggering 80.9 % over the single agent. 80%. That's huge. And decentralized was right behind it at 74.5%. It's a massive win because the task is highly decomposable. You can break it apart. You can assign one agent to analyze revenue, another for cost structures, a third for market comparisons, and then have a final agent synthesize it all. The team structure just perfectly matches the problem structure. Okay, an 80 % game is hard to argue with. But does that centralized orchestrator, that final bottleneck, does it become a single point of failure?

5:57Are you just replacing a brittle agent with a brittle coordinator? It's a valid concern, but in this specific task, the benefit just massively outweighed the risk. The pieces fit together cleanly. But then you flip the switch completely when you look at the PlanCraft benchmark. And this one is the opposite. It's strictly sequential, like figuring out the steps to craft an item in the game where every single move depends on the last one. And that is where we saw catastrophic degradation. All the multi-agent setups actively harmed performance. The worst was the independent architecture, a 70 percent drop in performance.

6:32Wait, hold on, a 70 percent drop. So for two years, we've been reading that debate models and collective reasoning are the big breakthrough. And you're saying on a sequential task, that's actively detrimental. Detrimental is exactly the right word. The task demands strict sequential constraint satisfaction. Every action changes the state of the world that the very next action depends on. The moment you introduce coordination overhead and you fragment that reasoning, the system just collapses. The agents are operating on different versions of reality and their errors cascade violently. The cost of just talking to each other degrades the quality of their work.

7:08So we've seen the two extremes, massive success in finance, total meltdown in planning. What about the average everyday workplace task? That would be the Workbench benchmark, and it's the moderate middle ground. Workbench has realistic office tasks, code execution, heavy tool use, and it showed really marginal effects, ranging from a slight 11 % degradation to a small 6 % game, basically a wash. And why is that? It's that balanced trade-off. The problem is complex, which should favor agents, but the heavy orchestration costs we just talked about, they punish them, it cancels out. And what about dynamic, high-entropy environments, like navigating the web in Browse Comp Plus?

7:49That one saw modest but consistent gains. Decentralized, the peer-to-peer model, excelled there with a 9.2 % improvement. Yeah, it suggests that when you're in a big, messy search space, having peers explore different paths at once is more valuable than perfect sequential planning. This all leads directly to, I think, the core innovation of this research, the predictive framework. Because if the results are this variable, you need a science, not just intuition, to know what to build. Exactly. The researchers built a robust model that can predict success or failure based on measurable properties of the task itself.

8:26not just the architecture label you slap on it. And it identified three dominant staling principles that really determine the outcome. Okay, let's hear them. Principle number one, the capability saturation paradox. What does that mean in practical terms? So this is one of the most counterintuitive findings. We observe diminishing or even negative returns once your single agent baseline performance gets above a certain threshold, roughly 45 % accuracy. So if your single agent is already kind of okay? Think of it this way. If your simplest, cheapest single model can already clear 45 % on a task, you should probably just optimize that single model before you even think about paying the coordination tax for a team.

9:05The room for improvement is so small that the overhead just hurts you. So if you've got a really capable top-tier LLM, the simpler design is probably the better economic and performance choice. Don't over-engineer it. 100%. Which leads to the second principle, the tool coordination tradeoff. And this was the strongest negative effect in the entire study. Stronger than anything else. Yes. And it directly counters that intuition that more agents should always help you manage more complexity. This sounds like the digital version of too many cooks in a kitchen with only one spatula. The agents are just fighting over the tools.

9:41That's a great analogy. Tool heavy tasks like the software engineering environments they tested with 16 different tools. They suffered disproportionately from this overhead. When you coordinate, you spend your precious token budget on communication. And that just leaves less capacity for actually orchestrating the complex tool use. The coordination tax is something that heavy tool use just cannot afford. It's an efficiency killer. That is a critical finding for anyone deploying these things in a real software environment. Okay, and the third principle, it's about error handling. Right. This is the topology-dependent error amplification.

10:17The key is that errors made by agents don't just cancel each other out. They cascade and propagate. And it depends on the team structure. Hugely. In that independent architecture, with no verification, errors amplify a shocking 17.2 fold. 17 times. That's a complete nightmare. It is. But compare that to the centralized model, where the orchestrator acts as a validation bottleneck. That centralization contains the error amplification to only 4.4 fold. So the manager slows things down, but it's worth it because they catch the huge mistakes. It pays massive dividends by stopping those catastrophic errors before they pollute the final output.

10:54Let's take a breath there because I think the next set of numbers is going to either terrify your CFO or send them running for the single agent hills. We need to talk about cost and efficiency. The cost of coordination is substantial. To quantify it, multi-agent systems use anywhere from 1.6 to 6.2 times the token budget compared to a single agent to get the same performance. Six times the tokens. The centralized model had a 285 % token overhead, and the super complex hybrid architecture, it soared to 515 % overhead. That's a five-fold jump in your operating expense for a system that might actually perform worse.

11:32That is a massive economic penalty. And this cost, it leads directly to a hard resource ceiling. We found that the number of reasoning turns, so the number of interactions needed to solve the problem, it grows super linearly with the number of agents. Specifically, it scales to the number of agents to the power of 1.724. That super linear growth means the complexity just explodes. If I double my team, I'm way more than doubling my cost. Exactly right. Moving from a two-agent system to a four-agent system doesn't double the cost. It almost triples the interaction overhead. So under any fixed budget, this really constrains your effective team size to just three or four agents.

12:10And beyond that? Beyond that, the per-agent reasoning quality just plummets because so much of the budget is being spent on communication, not on thinking. We can see this really starkly with the token efficiency metric. Success is per 1 ,000 tokens. The simple single-agent system gets 67.7 successes. That's your baseline. Right. But the most complex, overhead-heavy arrangement, the hybrid architecture, that drops to just 13.6 successes per thousand tokens. So five times worse efficiency for maybe worse results. That's a very tough sell for any real-world deployment. A very tough sell. Did this study find any general preferences among the big LLM families, OpenAI, Google, Anthropic?

12:49Do they seem to favor one topology over another? Yeah, there were some subtle but persistent biases. The Google model showed a really robust cross-architecture efficiency. It suggests their internal context management might just be better balanced for these cost-benefit trade-offs. They're more resilient. Interesting. How about Anthropic? Anthropic's models were the most conservative and stable, but really only under centralized coordination. They were also the most sensitive to the coordination overhead. So they need a firm hand. It suggests their models are really optimized for depth of reasoning in a single coherent stream.

13:26So parallelization is particularly taxing for them. They do well when controlled tightly, but they suffer quickly when you fragment their thinking. And OpenAI's models. They showed the strongest synergy with the hybrid architecture, especially on highly structured tasks. It suggests a really effective communication alignment. The models seem particularly good at handling those complex protocols that balance hierarchy and peer-to-peer chat. But these are secondary effects. They are. They might shift the optimal point a little, but the primary driver is still matching the coordination topology to the task structure.

13:59So let's bring it all back home for the listener. What is the core takeaway for anyone building or deploying these advanced agent systems? It's not about just adding more agents, is it? No, absolutely not. Ingentic success isn't about team size. It's about matching the coordination, topology, centralized, decentralized, whatever, to the specific structure of your task. If it's parallel, go centralized. If it's sequential. Stick to the simple single agent. We now have a principled predictive framework that achieves 87 % accuracy in predicting the optimal architecture. That level of predictability is the move beyond heuristics that the field has desperately needed.

14:38That is game-changing certainty. Now, we've talked a lot about the failures and the overhead, but it seems like the most complex architectures, the ones designed to be the best, actually introduced their own unique failure modes. Precisely. We saw this most clearly with the hybrid architecture. While it's powerful in theory, it had the highest rate of what are called coordination failures, 12.4%. And these weren't failures of the LLM itself? No, these weren't bad reasoning. These were failures of the complex protocol. An agent didn't follow the communication rules or it missed a validation step.

15:10It shows that when the complexity of your coordination protocol gets too high, even with very capable underlying models, the system breaks down in new and unexpected ways. Which raises a really important final question for you to consider as you dive deeper into this field. Does the pursuit of optimal coordination, designing this perfectly complex, efficient team, does it inherently create new, brittle dependencies that actually make the system less robust and more fragile than a much simpler design? It's a painful tradeoff between power and fragility that every AI engineer is now going to have to face.

From the publisher

This academic paper by Google Research, Google DeepMind, and the Massachusetts Institute of Technology, systematically evaluates the principles for scaling language model-based agent systems, moving beyond anecdotal evidence that "more agents is all you need." The authors present a controlled evaluation across four diverse agentic benchmarks, testing five canonical architectures—Single-Agent, Independent, Centralized, Decentralized, and Hybrid Multi-Agent Systems—to isolate the effect of coordination structure and model capability. Key findings establish that multi-agent benefits are highly task-contingent, ranging from a significant performance increase (+81%) on parallelizable tasks like financial analysis to substantial degradation (-70%) on sequential planning tasks, primarily due to measurable factors such as the tool-coordination trade-off and architecture-dependent error amplification. Ultimately, they derive a predictive quantitative scaling principle that explains over 51% of performance variance and can predict the optimal architecture for unseen task configurations.

More from Best AI papers explained

All 475 episodes
Towards a Science of Scaling Agent Systems / Google DeepmindBest AI papers explained · 16 min
Listen in VO