Many Agents, Many Problems (The Agents Season, Episode 8)

8 Jun 2026 · 28 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Multi-agent LLM orchestration—when multiple agents help vs when they hurt—using two research papers: “Collaboration Gap” (Microsoft Research/EPFL, 2025) and “Scaling Laws of Multi-Agent Systems” (Google DeepMind/MIT, late 2024).

Guest backgrounds

No named guests; the episode is a solo host discussion.

Key claims

(1) Agents can fail to collaborate even if each agent is strong alone; collaboration requires shared representations and communication skills. (2) Multi-agent gains depend on task structure: tool-heavy work and sequential dependencies often degrade performance; parallelizable work can improve it substantially. (3) Distilled models show sharper collaboration collapse.

Notable examples

Maze task with complementary hidden cells; “relay inference” (strong agent sets representation, weaker follows). Scaling results: tool use “kills” multi-agent value; ~45% solo success threshold; parallel financial analysis improved performance by 80%+, while sequential planning degraded by 70%.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenges of Collaboration

0:45 to 1:30

Exploring the effectiveness of working alone versus in teams, and the limitations of agents.

“working with other people where the challenge of us collaborating with each other is offsetting any of the gains that we might have from having multiple minds working together on a problem.”

Introduction to Multi-Agent Systems

1:30 to 2:54

Defining multi-agent systems and setting up a framework for discussion.

“You are listening to Linear digressions.”

Examining the Collaboration Gap Paper

2:54 to 4:50

Discussion on the research paper regarding AI agent collaboration and its findings.

“And this starts with a very basic question.”

Communication Challenges Among Agents

4:50 to 7:12

Analyzing how agents struggle with shared representations and communication.

“understand how adding more people to a problem, or in this case, more agents to a problem, can actually make things worse than if you had one of them alone.”

Impact of Distilled Models on Collaboration

7:12 to 8:20

Exploring how distilled models perform in collaborative tasks and their limitations.

“They're smaller models that are taken from compressing the larger models, like an LLM.”

Relay Inference as a Strategy

8:20 to 10:10

Introducing the relay inference method to improve agent collaboration effectiveness.

“So I think that's kind of an interesting failure mode.”

Exploring Google DeepMind's Research

10:10 to 11:08

Delving into a study about the scaling laws of multi-agent systems and their effectiveness.

“But let's just think about what that actually means.”

Findings on Multi-Agent Effectiveness

11:08 to 14:00

Discussing findings on when multi-agent systems succeed or fail in problem-solving.

“So let's take ourselves into a world now.”

Understanding Agentic Complexity and Coordination

14:00 to 16:56

Learn about the effects of multiple agents in problem-solving and the coordination costs involved.

“So you have multiple agents that are exploring different paths and it's increasing the chance that at least one of them succeeds.”

Hierarchical Multi-Agent Systems and Their Risks

16:56 to 18:16

Explore how hierarchical structures in multi-agent systems can amplify errors.

“In particular, they looked at parallelizable financial analysis tasks.”
Show all 12 chapters

Task Structuring: Parallel vs. Sequential Planning

18:16 to 21:46

Discover how the structure of tasks affects the performance of single vs. multi-agent systems.

“And here's what we're left with with all of this.”

Best Practices for Multi-Agent Systems

21:46 to 24:46

Understand the best practices for designing effective multi-agent systems and when they are beneficial.

“But that's kind of a narrower use case than the hype suggests.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi and welcome to Linear Digressions. By way of intro today, I want you to think a little bit about how you work. Do you work alone? Or when I say, how do you work, you think of work that you're doing with a team, maybe a group project, maybe it's a group at work, probably a little bit of both. And I think you would also say, if I asked you, what's the most effective way for you to work? well sometimes I can be really effective alone and sometimes I can be really effective with a team but also I can be distracted and not very effective and procrastinate and have all kinds of failure modes when I'm working alone and I can also have all kinds of pathologies when I'm working with other people where the challenge of us collaborating with each other is offsetting any of the gains that we might have from having multiple minds working together on a problem.

1:01Well, as it turns out, agents are not that different from you and me and all the other humans out there that when they're working alone on certain problems, they have limitations that they start to hit. And one of the ways to get around those limitations is to scale your systems out so that you have teams of agents. But just like teams of humans, those start to hit their own types of failure modes. And that's what we're going to be talking about today. Multi-agent orchestration, subtitled Many Agents, Many Problems. You are listening to Linear digressions.

1:41To make this a little bit abstract, let's talk about just what a multi-agent, a demonstration of a multi-agent system might look like. So there's many of these examples out there, but just by way of example, let's say that you have this orchestrator agent that is taking a research task. It's got a bunch of sub-agents that then it spins up. Let's say one of them does a web search. There's a web search agent. You've got a summarization agent. You've got a fact-checking agent. And they're passing results back and forth to each other until the orchestrator synthesizes a final answer. This looks super cool, and it is.

2:21But the question that we're going to dig pretty deeply into today is, does something with this level of complexity and sophistication, does it actually work better than one good agent that's doing the whole thing? And as usual, we'll be grounding this in some of the research that's out there right now, where the researchers are expressly studying this question. Two papers that we're going to use to study this today. The first is the Collaboration Gap paper. This is a group, Davidson et al. from Microsoft Research and EPFL from 2025. And this starts with a very basic question. Before we ask anything about whether multi-agent systems scale, can AI agents actually collaborate with each other?

3:08I'm not talking about a scripted pre-wired workflow, but can an agent coordinate with an other agent that they've never worked with before the way that humans do? In this paper, they set up, it's kind of a clever setup. I actually really like it that the tasks that they gave to the agents, they had two agent teams, like two agents, a pair of agents working together, and they gave them a maze, like a maze that you like trace the way through like as a game as a kid. So each of these agents gets a copy of the maze and half the cells are hidden. Moreover, the two copies are complementary. So when you put them together, you get the full picture of the maze.

3:49So neither agent can solve it all alone. They have to communicate in order to share the information about what the maze actually is so that they can solve it. But there's no prescribed format. So the agents can describe the maze however they want. And in this model, they tried a bunch of different models, 32 different models that they tested, tested them working individually and in pairs. And what they found was that models that solve mazes well solo on their own often fail pretty substantially when they're required to collaborate. And the failure, moreover, is not like a small failure sometimes.

4:26It can be dramatic. It can just fall off a cliff. So the question you may be asking yourself is like, wait, so you're telling me that it's getting worse by working together? Like, how is that even possible? And if you've ever been on, let's say, a bad group project or you've had to work with someone that you don't mesh with where you can't communicate well with that person, I bet you can absolutely understand how adding more people to a problem, or in this case, more agents to a problem, can actually make things worse than if you had one of them alone. But okay, how does it work for agents? The problem that they find as they were digging into some of these failure modes is that the agents in this particular case, they were struggling to establish this shared representation, they weren't communicating in the same way about how to solve the task at hand.

5:18So if you and I were talking about a maze together, what would probably happen pretty quickly, even if we weren't sitting next to each other looking at the same maze, if we only had words to describe it to each other, we would probably end up on something like a grid. Like I'm on the third row from the bottom, I'm in the fifth column over, and there's a wall on my right. And then you would start to build your mental model, and then you would use that same representation like columns and rows maybe, you know, left, right, above, below. And we would start to communicate in a way like that. Obviously, that's not the only way that you could communicate what's in a maze, but that's probably something we would converge on pretty quickly.

6:01The agents struggled with this. So when it failed, what was happening was each of those agents would start to talk about the maze in a different way. So one of them might be using directions like left and right and up and down. One of them might be using coordinates like east and west and north and south or whatever. One of them might start counting from the upper left-hand corner, like over and down. One of them might be starting from the center and then over and up. So they're talking about the same task, but they're not using the same representation. They're not using the same language to talk about it.

6:39And so they're spending all of this conversation time, all of this back and forth, just trying to figure out each other's representation and going in circles, hitting this wall with each other. And they're not actually solving the maze when they do this. And so it doesn't actually work that well. Moreover, and I think this is very interesting, this wasn't something that I was expecting to find in this paper. Some of the models that they looked at were distilled models. And we haven't talked about distilled models a lot on this podcast, but those are basically condensed versions. They're smaller models that are taken from compressing the larger models, like an LLM.

7:20You kind of like slim it down, compress it, distill it. Those are the models that showed some of the most dramatic collapse. So a distilled model in particular, they found distilled models that were pretty good at solving these mazes independently, but they really fell apart when it became this collaborative exercise. And there's an intuition here that I think maybe you can start to build from this, which is that when you're distilling the model, that process is maintaining the ability for the model to produce correct answers. So you're making the model smaller, but you're saying like, preserve your ability to produce correct answers.

7:57But what it's not necessarily optimizing for is producing models that after the distillation process are any good at explaining their reasoning in a way another agent can build on. So they're geniuses at coming up with the correct answer, but they're doing a particularly poor job at communicating and explaining to one another. So I think that's kind of an interesting failure mode. And at least for me, it gives me a little bit of an insight into how we should think about distillation and what that might be doing to the models themselves. Anyway, with all of that on the table, what's the theoretical claim of this paper?

8:39What should we take away from this? Well, it basically says that, or the thesis of this paper is that maybe there's this capability of models to be collaborative with each other, and that might be a separate and distinct capability from just raw task capability or raw intelligence. In other words, just because a model is super, super smart doesn't mean that it's good at collaborating with others. Again, I think we have all worked with people who are really brilliant and terrible at explaining their reasoning or terrible to work with. Not that different for agents. There is one partial fix that I want to mention here because it's going to come up again a little bit later in this episode, which is that in some of these pairings, they purposely put together a very strong agent and a weaker agent as its partner.

9:33And in particular, when they had the stronger agent go first, kind of set forward a shared representation. Hey, we're going to talk about it this way. I'm just going to lay down the rules of communication in this conversation, talk about the maze in this way. And then the weaker agent would just follow that lead. This is a setup that they called relay inference in this paper. And this actually did pretty well. It closed most of the gap that they found between the solo and paired agent systems. But let's just think about what that actually means. This isn't two partners that are collaborating with each other as equals.

10:17there's basically a strong agent that's doing the heavy lifting here and a weaker agent that's just following its lead. So you're not really getting teamwork, you're getting leader and follower. With all of that said, as much as this was a very interesting and fun way of studying, setting up and studying the way that these particular agents collaborated on this maze task, We do know that there are certain systems and they're being built every day. Systems that are multi-agentic inherently, they're building these teams of agents. And the second paper that I want to talk through, the second thing that I want to explore, is research out of Google DeepMind and MIT from late 2024, talking about the scaling laws of multi-agent systems.

11:08So let's take ourselves into a world now. let's say assume that the agents can collaborate. When does having a multi-agent system actually help you do better on the problem that you care about? Because we know that this is sometimes the case. So we're granting that the collaboration problem is solved or at least manageable. We're not in a situation where we've got two models that are good at collaborating. When are they going to do well together? What types of problems characterize successful multi-agentic systems, and where are there places where the communication and collaboration overhead just overwhelms any benefit that you might get from specialization or so forth.

11:53So the group of researchers that wrote this paper ran 180 configurations of different multi-agent systems across four different benchmarks. They used three different model families. So they're really trying to, as much as they can, come up with, they called it about scaling laws of multi-agent systems. They're trying to make it not anchored too much on any particulars of the model that they picked out or the task that they set out. So they're obviously not covering every possible way to solve every possible problem because that's not possible, but they are trying to come up with something that's this generalizing here, this giving them some general scaling laws.

12:36A few things that they found, three patterns that function almost like laws. Law number one, tool use kills multi-agent value. On tasks that require heavy tool use, so lots of web browsing, API calls, retrieval, these are some of the tasks that suffer the most from the multi-agent overhead, from the coordination cost. So every time there's a message between those agents, it's consuming some of the budget that it has for reasoning. And if your agents are spending most of their thinking coordination, so if the agents in cases like this, they end up spending most of their energy coordinating instead of actually doing the task.

13:26These are cases where a single well-equipped agent tends to do very well, almost always wins here. Law number two, we'll call the 45 % threshold. And this is actually a very, I would say, practically useful finding, which is that once you have a single agent that can solve a task about 45 % of the time on its own, adding more agents past that point stops helping and it starts hurting. So below that 45 % threshold, you can get some real boost from parallelization. So you have multiple agents that are exploring different paths and it's increasing the chance that at least one of them succeeds. But above that 45%, the coordination cost starts to exceed the benefit.

14:13This is not a law that you can derive from first principles or anything. It's more of an empirical finding. But it kind of gives you a threshold. It's around 45 % or so of if we started to add more agentic complexity to this system, should we expect it to get better or not? Well, it kind of depends on how well the single agent working on its own is doing. If you'd like, there's an analogy that I think might help, which is imagine that you're working on math homework and there's a problem that's pretty hard. So if you're completely stuck, having a study group is probably going to be helpful because you have different people.

14:53They're trying different approaches. They're brainstorming ideas together. You know, you have different strategies that are in the mix. And then when somebody starts to break through, they can kind of bring along the rest of the group with them. But if you're one of the stronger students in the class and you're already pretty good at your math homework, having five people all in there trying different stuff and second-guessing each other's approaches and trying to figure out who's right, like that can actually slow you down. Agents are kind of the same way. Law number three is that the topology of the task, the structure of the task and of the agentic system, that matters a lot.

15:33And in particular, hierarchies can amplify errors. So in hierarchical multi-agent systems, what I mean by that is you have an orchestrator agent and it is coordinating work, telling workers what to do, and then receiving their results back up. What can happen in that system is the sub-agents, the ones that are actually doing the work and passing it back up to the orchestrator, back up to the boss, those sub-agents can make mistakes. And when those mistaken results get passed back up to the orchestrator, it can corrupt the entire task. And the more layers you have, like sometimes you have sub-agents with sub-agents with sub-agents, the worse that this can get.

16:19Now, you can also have failure modes where you have multi-agent systems that are sort of flat, where everybody's equal and there isn't one that's overseeing the work of all the other ones. These fail too. And they fail definitely. But what they find, and this sort of makes sense, is that all of the multi-agent topologies that they studied, all the different ways of structuring multi-agent systems, have error amplification risks. And those risks simply don't exist with a single agent because there is no way of, with a single agent, having an error from one part of the system propagate to another.

16:56There's one place, though, in particular that was a strong contrast to the overall conclusion here, which is that there's certain types of tasks, very particular types of tasks that actually can see huge gains from multi-agent systems. In particular, they looked at parallelizable financial analysis tasks. So these are things that you can split into independent subtasks that are not dependent on each other. You just run them in parallel. And in those cases, multi-agent topologies improved performance by over 80%. On the other hand, when you had sequential planning tasks where there's dependencies from one step to the next, step A has a dependency on step B, step B depends on step C.

17:46in those sequential planning tasks, having multi-agent topologies degrade performance by 70%. So you're taking the same framework, the same models, this is being done by the same researchers. There's just differently structured tasks they're looking at here. And there's a huge swing in whether the multi-agent systems did better or did worse than a single agent acting alone. So it really comes down to how easy it is to split the task into different pieces and what the dependency is of those pieces, you know, the parts on each other. So here's what we've got at this point. And here's what we're left with with all of this.

18:32The first part of this episode, we said there's this collaboration gap paper that says before you build a multi-agent system, you need to worry about whether your agents can actually communicate effectively. and that current models are generally optimized for getting answers correct, that effective communication is a separate and distinct skill, especially for distilled models, and this isn't a problem that you can prompt engineer your way out of. The scaling laws paper said in the second part of the podcast, even if the agents can communicate with each other, the payoff that you get is very dependent on the task that you give them.

19:16Or even more to the point, it depends on the pairing of the task that you give them and how you structure your agent system. In particular, if you have tasks that have a lot of sequential dependencies, like one step after the next, after the next, if your task is very tool heavy, or if your task is something that a solo agent can do correctly 45 % of the time or better, these are some of the areas where they found the multi-agent systems generally doing worse. Those are regimes where the single agent is going to win. Let's tie a couple of pieces together here, which is remember how the collaboration gap paper, though, it had a bit of a fix, which is don't have two equal pairs.

20:07Have a stronger agent that goes first. It sets the terms of the communication and the weaker agent just follows along. Let's think about that, what that means now in light of the scaling laws. So a relay architecture, that stronger than weaker pairing, that's really similar in some ways to a sequential pipeline where you would have a strong agent that does step one, and then it has a handoff to a weaker agent for step two. And sequential pipelines are exactly one of the cases where the scaling law paper says those can be some of the most fragile and most prone to error amplification, when you have errors in one step that can be handed off to the next step and then just propagate all the way through the system.

20:54So you have this partial solution to the collaboration problem but then it runs into a different structural problem in the second paper that we looked at. So where does this all leave us? I think the overall message here is kind of emerging, even from these two very different takes on multi-agent system, which is that the closer a multi-agent system looks like to a well-equipped single agent, and by that I mean you have a model with good tools working alone, the closer you get to something that looks like that, the better it tends to perform. So multi-agent systems, they do have a real place in the ecosystem.

21:33They're good at genuinely parallelizable tasks. They're better at tasks that are below that 45 % threshold. They're better at tasks where having multiple independent explorers can actually help. But that's kind of a narrower use case than the hype suggests. You kind of have to find the right niche for those. And so what should we take away from this? The first thing I'll say is that both of these are really, really good papers, but they are also at this point for 2024 and 2025. So I'm going to be looking around in the research some more, keeping an eye on this to try to tell if there's updates to this perspective as agenting systems become more sophisticated, as models become better.

22:18But I think for now, it really drives home a general best practice, which is that just because there's a system that says it has a bunch of agents and that they're all orchestrated and you have sub-agents that do all kinds of crazy stuff and there's handoffs and there's sophistication and there's a diagram on a slide and it's got 70 boxes on it and they all have arrows pointing to each other. Like this is the time when your detector should be going off a little bit. A question to ask yourself is like, if you had one really good agent with the right tools, could it do this? And if it could, that in all likelihood, that might be a better approach.

22:59And that just because we threw more agents at this does not necessarily mean that it's doing a better job than a simpler system that has just really good components for each of the pieces. In particular, think about the task that it's trying to solve and the steps that it has to take. If those steps are actually independent, then that's where you could get some pickup from having multiple agents working on them independently where the communication collaboration is less important. But if you have a task that's kind of artificially split into different pieces, then chances are just the communication back and forth between the agents might be negating any gain that you get from the specialization.

23:42So where does this sit in the big season arc that we've been going through for several episodes now about agents as a whole? And what I would say is that, especially thinking about what might it mean to delegate to an AI, delegation of a task to a team of AIs is not obviously better than delegation to one good one. And in many cases, it's worse. And there's an intuition that you might have that having more specialized agents equals more capability, and that just isn't holding up empirically in the research that we're looking at today. But that said, the research also tells you exactly when it does work.

24:19these parallel independent tasks that are below that 45 % threshold. And that's real. So the point isn't that multi-agent is bad. It's that it requires matching the architecture to the task in ways that require some real thinking that are not necessarily automatic. And that the current training doesn't necessarily optimize for giving agents collaborative skills that makes them good at working with each other in unstructured settings. So just throwing more agents at the problem is not going to solve. Just throwing more agents at the problem doesn't mean you're building a better system. It runs into the same core problem as building a team of humans.

25:00So just getting smart people in a room together doesn't necessarily mean that you have something that's effective. You have to design for collaboration. You have to find humans that can talk to each other, that can collaborate, and where their skill sets and the tasks that they're working on together fit into each other in an architectural way. So I would say there is some real remaining engineering and design and thinking through of these types of problems that you and I get to do, at least for a little while longer, because just throwing more agents at a problem is not a one-way ticket to an easy solution.

25:42With that, I want to thank you, as always, for joining us this week. A few quick reminders. If you're not a subscriber to Linear Digressions and you like this, we'd love to have you subscribe. You can do that on Spotify, iTunes, wherever you'd like. Also want to remind you that we have a newsletter. It gets a lot of the content that we cover in each week's episodes kind of in a distilled, fitted in your back pocket kind of format. And it also has a little bit of content that didn't make it into the main episode. At this point, it's not even just stuff that I find in my research. Sometimes it is just stuff that I find in my research.

26:20At this point, though, I often go out and just find additional extra stuff that I think is cool and interesting. Anyway, this week, what we're going to talk about is a really fun paper about, let's say you have a multi-agentic system. And one of the things that you have to do sometimes is you might have agents that disagree with each other, and you need to figure out some way to like reach a resolution. And so what this paper is looking at is should you have the agents, you can have them debate with each other and each one tries to convince the other one. That's one way you could try to resolve the conflict.

26:53Or it horse races that against having them vote. And you know, which one gives better decisions for multi-agent LLMs. I am not going to spoil the punchline for you on that one. I think it's actually really fun. So if you subscribe to the newsletter. I will have the link to that paper about debating versus voting and a few notes on what they actually found in that research. And of course, you can go read it yourself if you have the inclination. Anyway, with that, I want to thank you again for joining us. We will be back next week with even more on agents. We're starting to get near the end of the agent season, which is kind of crazy.

27:33But in the meantime, this has been a really fun one for me. I hope for you too. and we'll talk to you again soon. Thanks.

27:44This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today? If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

From the publisher

Whether you work best solo or thrive in a team, you know collaboration is complicated — and it turns out AI agents face the same tensions. This episode dives into multi-agent systems, exploring how networks of AI agents can overcome the individual limitations of a single model, and what the research says about when collaboration actually helps versus when it just adds noise. Think scaling laws, but for teamwork.

---
Website: https://lineardigressions.com
Apple Podcasts: https://podcasts.apple.com/us/podcast/linear-digressions/id941219323
Spotify: https://open.spotify.com/show/1JdkD0ZoZ52KjwdR0b1WoT
Substack: https://substack.com/@lineardigressions

More from Linear Digressions

All 35 episodes
Many Agents, Many Problems (The Agents Season, Episode 8)Linear Digressions · 28 min
Listen in VO