Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

30 Aug 2025 · 20 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Chain-of-Agents (COA) proposes end-to-end “agent foundation models” (AFMs) that simulate multi-agent collaboration inside one model, avoiding slow multi-agent back-and-forth while retaining planning, reflection, and tool use.

Guest backgrounds

No guests are named; it’s a host-led “Deep Dive” discussion of an OPPO AI agent team paper.

Key claims

AFMs bridge multi-agent systems (MAS) inefficiency and tool-integrated reasoning (TIR) lack of orchestration by dynamically coordinating internal “role-playing” and “tool” agents. Two-stage training: multi-agent distillation with progressive quality filtering, then agentic RL with LLM-as-judge rewards.

Notable examples

GAIA (32B AFM 55.3% vs WebSailor 53.2%, WebDancer 51.5%; near OWL 55.8%, OVGEN 58.3%), Humanity’s Last Exam (18.0% with RL vs 12.6% OWL), math (78.0% avg; AME25 +10.5%), coding (LiveCodeBench/code contests gains). Efficiency: 84.6% lower token consumption than traditional MAS. Generalization: code agent adapts to unseen tools (e.g., Honey Density using new web search + Python). Test-time scaling: GAIA 55.3%→57.3% (best-of-3) and 69.9% (pass-at-3).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Chain of Agents

0:45 to 3:36

Exploring how Chain of Agents revolutionizes AI collaboration for complex tasks.

“So this chain of agents or COA, it shifts how multiple AIs work together.”

Current Limitations in AI Systems

3:36 to 4:50

Discussing the inefficiencies and limitations of current AI collaboration methods.

“So it's not just using tools, it's simulating the team using the tool.”

Introducing Agent Foundation Models

4:50 to 8:20

Overview of Agent Foundation Models and their potential benefits for AI efficiency.

“Role-playing agents, tool agents, what are they?”

Two-Stage Training Process

8:20 to 11:14

Diving into the two-stage training framework for building sophisticated AI models.

“they figure it's too easy, not complex enough for the AFM to learn deep reasoning.”

Performance and Efficiency Gains

11:14 to 14:00

Examining the performance results and efficiency improvements of AFMs in AI tasks.

“Distillation from expert examples, then refinement through rewarded trial and error.”

Efficiency Gains of AFMs

14:00 to 14:40

Learn about how AFMs drastically reduce token consumption and costs.

“Okay, so it seems clear that the SFT stage builds that core chain of agents capability, the planning, reflection, coordination.”

Adapting to New Tasks

14:40 to 15:20

Discover how AFMs can generalize to unseen tools without retraining.

“What does token consumption mean in practice?”

Learning Tool Use

15:20 to 16:10

Uncover the implications of AFMs learning to use tools effectively.

“That feels like the next level of intelligence.”

Challenges with New Tools

16:10 to 16:40

Examine the limitations AFMs face when dealing with unfamiliar tools.

“It means the model didn't just memorize how to use its tools, it learned a deeper, more fundamental understanding of how to use tools in general, based on a description.”

Test Time Scaling Strategies

16:40 to 18:00

Learn about various strategies that improve AFM performance after training.

“But the code agent, because it was already trained with very strict formatting rules for code, seemed more robust and handled the new tools, including their syntax, more consistently.”
Show all 13 chapters

Breakthroughs in AI Problem Solving

18:00 to 19:10

Explore the significant achievements and efficiency of AFMs in AI.

“That's a massive 14.6 percentage point increase over the base AFM score of 55.3%.”

Open Sourcing AI Models

19:10 to 19:50

Discover the importance of open-sourcing AI models and its impact.

“And something we haven't stressed enough for me, maybe the researchers open sourced everything.”

Future of AI Interaction

19:50 to 20:15

Reflect on the transformative potential of advanced AI agents.

“what kinds of problems previously thought impossible for AI might suddenly become tractable?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hey there and welcome to the Deep Dive. We're the place that cuts through the noise to bring you the really fascinating stuff from the sources you give us. And today, wow, you've handed us a really exciting paper. It's called Chain of Agents, End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL. It's from the OPPO AI agent team just published this month. Super fresh. Yeah, really cutting-edge stuff. And what we're going to do today is unpack this groundbreaking approach. It's really changing how large language models, you know, LLMs, tackle the really complicated problems.

0:34Right. Like think of it as giving an LLM the power to be a whole team of experts all working together smoothly, but like inside itself. Exactly. It's sort of a shortcut to getting these super intelligent AI systems that can handle all sorts of tasks faster, smarter and surprisingly efficiently, too. So this chain of agents or COA, it shifts how multiple AIs work together. Fundamentally, yes. Instead of having separate AIs talking back and forth, which can be clunky and slow, COA basically integrates that whole collaborative process into one single model. Huh. So the LLM itself learns to wear different hats, like a strategist, a planner, a critic.

1:11Precisely. And it does it dynamically as it's solving the problem. It's a really different way of thinking about aging collaboration. Okay, let's unpack this then. But maybe before we get into the solution, the COA thing, it helps to understand what problems current AI systems have with these complex tasks. Good point. The paper highlights two main ways things are done now. Multi-agent systems, or MAS, and tool-integrated reasoning, TIR. Right. MAS. That's where you have different AIs collaborating, like for deep research or writing complex code. Yeah. And they can do some remarkable things, no doubt.

1:46But they come with significant downsides. They tend to be computationally inefficient. Meaning they use up a lot of computing power. Exactly. All that communication back and forth between agents is redundant, costly. Plus, they often struggle to generalize. You teach them one task, they're not great at adapting to a new one without a lot of reengineering. Hmm, okay. And they don't easily learn from new data over time. That's another issue. Continuous data-centric learning is tricky with those distributed systems. It's like having a committee trying to learn together slow and cumbersome. So a team of AIs sounds good on paper, but it's slow, hard to manage, hard to update.

2:24Okay. Then what about the other one, TIR, tool integrated reasoning? Right. TIR is where a single LLM uses external tools. Think of it like you using a search engine or a calculator while you work. Okay. That sounds more straightforward. It is in a way. TIR models follow this kind of think action observation loop, which is definitely a step forward for, let's say, independent problem solving by a single AI. Well, it's the catch. The main limitation is, well, it's generally just that single agent. It lacks the sophisticated orchestration. It can't really manage a diverse set of role-playing agents and a wide array of tools all working together in a truly collaborative, multi-step fashion.

3:03Ah, so it misses that complex teamwork aspect that MAS, despite its flaws, could sometimes achieve. Exactly. MAS had that potential for deeper collaboration, even if it was inefficient. TIR is often too limited in its scope. Okay. So this is where chain of agents COA comes in. This sounds like where things get really interesting. This is the core idea. COA aims to bridge that gap. Imagine, like you said, a single LLM that can natively, inherently perform that complex problem solving the way a whole team of agents would. But it's all happening inside its own architecture. Wow. So it's not just using tools, it's simulating the team using the tool.

3:43Precisely. Instead of needing separate agents, separate frameworks, COE lets one model dynamically call upon different internal tool agents and role-playing agents. It simulates the whole multi-agent collaboration end-to-end. And these models they create are called agent foundation models, AFMs. That's the term they use, yes, agent foundation models. It signifies a step towards more capable economist AI. So, okay, let's bring this back to the listener. Why should, you know, why should you care about this AFM stuff? Well, think about it. It could mean less need for complex prompt engineering, you know, figuring out the exact perfect way to ask the AI to do something complex.

4:19Right. Which can be a real headache. Exactly. And less need for designing intricate workflows. Plus, a huge one, significantly lower computational cost. OK. Because you're cutting out all that inefficient agent to agent chat. So less computing power, less money, faster answers. Potentially, yes. It points towards AI that's more efficient, easier to integrate, and maybe capable of tackling problems that were just too cumbersome before. Okay. I'm sold on the why. Now, but how? How does one model actually wear all these hats? Role-playing agents, tool agents, what are they? Let's pull back the curtain a bit.

4:55The paper describes these two core components, right? Role-playing agents and tool agents. Okay. Think of the role-playing agents as the strategists, the coordinators, the internal management team, if you will. Gotcha. Like different jobs within the team. Exactly. There's a thinking agent. It's the orchestrator, guiding the whole reasoning process, deciding what needs to happen next. Okay. Then a plan agent. Its job is to take your complex request and break it down into structured, manageable tasks. Makes sense. Like creating a project plan. Kind of, yeah. And then, crucially, there's a reflection agent.

5:27This one performs self-critique. It asks things like, are we on the right track? Are there inconsistencies here? Could this be done better? Whoa, self-critique built in. That's interesting. Very. And finally, a verification agent validates the reasoning or the final answer, checks the integrity of the whole process. So if I asked it, summarize recent AI breakthroughs and discuss their potential impact on, say, climate change research. Right, the plan agent would likely break that down. Step one, search for recent AI papers. Step two, crawl relevant articles, maybe focus on climate applications.

6:03Step three, synthesize the key breakthroughs. Step four, analyze potential impacts on climate research. Something like that. That sounds remarkably organized, almost human-like. That's the goal, I think. And then you have the tool agents. These are the specialized workers, the ones that actually do things. Like the specialists on the team. Precisely. You've got a search agent to formulate smart queries for web search. a crawl agent to pull specific content from web pages, and a code-generated agent that can write and run code, maybe Python, in a safe sandbox. Useful for calculations, data analysis, whatever's needed.

6:36Okay. And the key here, the real magic, is that the thinking agent dynamically coordinates all of these, both the role players and the tool users. It manages the state of the problem as it evolves. So it's not just a fixed sequence, like do step one, then step two. No, not necessarily. It's adaptive. It might realize midway it needs to go back and search again or maybe run some code to verify something. It's much more dynamic than traditional tool integrated reasoning. Okay, that sounds incredibly powerful, but also incredibly complex to teach an AI to do. Yeah. How on earth do you train an LLM to be the sophisticated internal orchestrator?

7:12Yeah, it's definitely not simple. It involves a pretty clever two-stage training framework. Two stages. Okay, what's the first one? The first stage is called multi-agent distillation. Think of it as a specialized form of supervised fine-tuning, or SFT. SFT. That's where you show the model lots of examples of good input and output, right? Basically, yes. But here, it's agentic SFT. The AFM learns by observing and essentially copying the successful trajectories, the step-by-step actions of existing state-of-the-art multi-agent systems, like watching the best teams work and learning their strategies.

7:49So it's learning from the best examples of how a diverse team would solve a problem, capturing that whole sequence. Exactly. It internalizes those successful patterns of decision-making and agent activation. But how do you ensure it's learning from good examples, not just any random attempts? Ah, good question. They're very meticulous about the data quality. They use something called progressive quality filtering. Progressive quality filtering. Okay, what does that filter out? Several things. First, overly simple tasks. If a problem is solved in fewer than, say, five-agent tool interactions, they figure it's too easy, not complex enough for the AFM to learn deep reasoning.

8:24Right, needs to learn the hard stuff. Exactly. They also filter out dirty data. Things like incorrect final answers. Or maybe runs with redundant steps or inputs. Makes sense. Clean data in, better model out. And here's the really smart part. They specifically prioritize trajectories, where the system showed self-critique and reflection. using that reflection agent we talked about. Oh, interesting. So it learns the value of self-correction. Precisely. And they even upsample examples where the model explicitly identifies and corrects its own errors during the process. Wow. So it's not just learning to get the right answer.

8:58It's learning how to think critically and fix mistakes along the way. That's the idea. Building in robustness. And they focus on problems needing 5 to 20 hops or steps, which is way more complex than typical benchmarks. Okay, that's stage one. Learning complex, high quality, self-critical problem solving from the best examples. Very cool. What's stage two? Stage two is agentic reinforcement learning or RL. Ah, reinforcement learning. Learning by trial and error with rewards. You got it. If SFT is like learning from a textbook or watching experts, RL is like getting out there and actually doing it, learning from the consequences of your actions.

9:34Okay. So how does RL refine the AFM? After the initial SFT, RL fine-tunes the model's internal policy, basically. Its strategy for orchestrating the tools and roles. It uses rewards based on the final outcome. Did it succeed at the task? So it learns to optimize for long-term success, making better strategic choices about which tool to use when. Exactly. Optimizing for task completion and efficient tool use in these dynamic situations. And again, they're smart about applying RL. They use strategic sampling to focus the RL training on the really challenging problems. How do they define challenging here?

10:11They filter out problems that seem too easy based on initial tests. For example, web queries that seem solvable without tools more than 30 % of the time. Or coding tasks that simpler versions of the model already solved easily. They want the RL effort focused, where it adds the most value on the hard stuff. Makes sense. Maximize the learning from the expensive RL process. And how does it get the reward signal? How does it know if it did a good job, especially for complex tasks? Good question. For web tasks, they actually use another LLM as a judge, an LLM as judge, to give a simple correct or incorrect assessment.

10:45Using an AI to judge another AI. Interesting. Yeah, it's becoming more common. For code and math problems, the reward is more direct. It's based on correctness. Did the code pass all the tests? Did the math answer match exactly? and also importantly, was the output format correct? So a clear feedback loop. Did you get it right and did you present it correctly? Precisely. That feedback is crucial for the RL process to hone the model's abilities. Okay, so that's a very sophisticated two-stage training process. Distillation from expert examples, then refinement through rewarded trial and error. The big question is, does all this intricate training actually work?

11:22Does it pay off? The results in the paper suggest a pretty resounding yes. These agent foundation models, AFMs, are setting new state-of-the-art results across a range of really tough benchmarks. Okay, let's hear some examples. What about those web agent tasks, navigating websites, finding info? Right, these are complex. On benchmarks like GAIA, which involves that kind of intricate web navigation and information extraction, the AFMs showed consistent superiority. For instance, their 32 billion parameter AFM hit a 55.3 % success rate on GAIA. And how does that compare? Well, it outperforms other methods like WebSailor at 53.2 % and WebDancer at 51.5%.

12:00And what's really striking to see is it gets very close to much larger systems based on GPT 4.1, like OWL at 55.8%, and the original OVGEN system at 58.3%. Wow. So a 32B open source model is competing neck and neck with potentially much larger proprietary models. That's impressive. It really is. And on another tough benchmark, Humanity's Last Exam, or HLE, which has questions needing expert-level reasoning across many fields. Yeah. The 32B AFM, enhanced with reinforcement learning, scored 18.0%. That's significantly better than the 12.6 % scored by OWL, which again used GPT 4.1. Okay, so it's not just good at web tests.

12:38It shows strong reasoning on very broad, difficult questions, too. Exactly. It points to really robust problem-solving capabilities. And it also showed strong generalization on things like multi-hop question answering MHQA. Those are the ones where you need to connect pieces of information from different places. Right. The AFM achieves state-of-the-art average performance, even on data sets it hadn't seen during training, improving over previous best methods by up to 6.8 % for some model sizes. So it learns the skill, not just the specific data. What about coding and math? Critical areas. Yeah, they tested that thoroughly too.

13:12And AFM shined there as well. For mathematical problem solving, the 32B AFM RL model got an average accuracy of 78.0 % across five different math benchmarks. How does that stack up? That's a solid 3.6 % improvement over the previous best overall. And notably, on one particularly challenging benchmark, AME25, which is designed for, you know, math Olympiad level thinking, it showed a huge 10.5 % absolute improvement. 10 % jump on a really hard math test. Okay, that's significant. And code generation. Similar story. Substantial gains there too. The RO enhanced AFMs boosted accuracy quite a bit over the base models.

13:50They started from like 8.5 % gain for the 7B model and 13.2 % for the 32B model on average across competitive programming benchmarks like LiveCodeBench and code contests. Okay, so it seems clear that the SFT stage builds that core chain of agents capability, the planning, reflection, coordination. Right, the foundation. And then the RL stage really sharpens it, pushing the performance, especially on these tougher, more strategic tasks. That's a great way to summarize it. SFT provides the playbook. RL refines the execution under pressure. Performance is great, but you mentioned efficiency earlier.

14:24That seems like a huge practical advantage. How much more efficient are these AFMs? It's actually quite dramatic. They measured inference cost, basically how much computation it takes to get an answer. Yeah. And the AFM significantly reduced this. Specifically, they reported an 84.6 % reduction in token consumption compared to the traditional multi-agent systems they were learning from. 84.6%. That's massive. What does token consumption mean in practice? Think of tokens as like pieces of words or information the AI has to process. Fewer tokens means less computation, which means lower energy use, lower cost to run, and often faster response times.

14:59So it's not just doing the job of many agents. It's doing it way, way cheaper and faster. That could change how these things are deployed. Absolutely. It makes advanced agent capabilities potentially much more accessible and practical for real-world use. Okay. Cheaper, faster, smarter. What about adaptability? Can an AFM trained on, say, coding tasks suddenly handle a web search task if needed without being retrained? That feels like the next level of intelligence. That's a crucial question about generalization. And the research explored this. Specifically, they called it generalization on unseen agents.

15:34Unseen agents, meaning tools it wasn't explicitly trained to use. Exactly. So they took a code agent model trained only on code and math problems. Then at inference time when it's actually being used, they give it descriptions of totally new tools like a web search tool or even a visual inspector tool just in the prompt. Did it work? Remarkably, yes. The code agent was able to correctly figure out how and when to use these completely unseen tools to solve tasks. They gave an example involving Honey Density, where it first used the new Web Search tool, then used its familiar Python tool for calculation.

16:09It adapted on the fly. That is genuinely incredible. It means the model didn't just memorize how to use its tools, it learned a deeper, more fundamental understanding of how to use tools in general, based on a description. Precisely. It reflects a more robust abstract reasoning capability. It's learning the meta skill of tool use. Now, they did mention some nuances, right? Did it work perfectly every time? Well, they noted the web agent model sometimes struggled a bit with the exact formatting for new tools, like maybe forgetting the specific syntax needed to call it correctly. But the code agent, because it was already trained with very strict formatting rules for code, seemed more robust and handled the new tools, including their syntax, more consistently.

16:50Interesting. So the type of training might influence how well it generalizes to the fine details of new tools. Okay, one last technical bit they mentioned, agentic test time scaling. What is that? Sounds like boosting performance even after training. Yeah, that's exactly what it is. It means applying some relatively simple strategies during inference when you're asking the trained model for an answer can significantly boost its performance further. Like what kind of strategies? One they tested was AFM-BO3. That stands for best of three. The model generates three possible answers or solutions, and then it or another process picks the best one.

17:24Okay, like getting a second and third opinion. Did that help? It did. On the GAIA benchmark, it bumped the score from 55.3 % to 57.3%. And on that HLE reasoning test, it jumped from 18.0 % to 23.0%. So noticeable gains just from trying three times. Okay. Any other strategies? Yeah. An even more effective one seemed to be AFM pass at three. This means the model gets three attempts, and if any of those attempts result in a correct answer, it counts as a success. Ah, so giving it multiple shots at the problem. Right, and this had a huge impact. On GAIA, pass at three, surged the AFM score all the way up to 69.9%.

18:01That's a massive 14.6 percentage point increase over the base AFM score of 55.3%. Wow, almost 70 % on GAIA just by allowing three tries. That's incredible. What does that tell us? It demonstrates that these end-to-end agent models, like AFM, benefit hugely from these kinds of test time scaling strategies, much more so potentially than traditional multi-agent systems. It helps close the gap even further with the performance of top proprietary models. What an incredible deep dive this has been. You brought us a paper that really feels like a glimpse into the future of AI problem solving. It really does.

18:36This chain of agents paradigm and the resulting agent foundation models, it feels like a genuine step change. Yeah, demonstrating that a single, well-trained LLM can effectively embody and orchestrate that whole team, roles, tools, planning, reflection. It's pretty amazing. And the results speak for themselves. State-of-the-art performance across such diverse and challenging areas, web navigation, math, coding, while also being drastically more efficient. That combination is powerful. It really sets a new benchmark for what integrated AI agents can do. Absolutely. That ability to generalize, to learn robustly and to do it efficiently.

19:11It's a remarkable achievement. And something we haven't stressed enough for me, maybe the researchers open sourced everything. The model weights, the code, the training data. That's a huge deal, isn't it? Yeah. It means anyone, any researcher or developer out there can take this work, build on it, experiment with it. It really accelerates progress for the whole field. Yeah, it lowers the barrier to entry for building these super capable agents, really opens up possibilities. Definitely fosters innovation. So as we wrap up, here's something for you, the listener, to think about. A provocative thought, perhaps.

19:43If a single AI model can now effectively act like an entire team of specialized agents coordinating complex tasks internally with this kind of efficiency and adaptability, what kinds of problems previously thought impossible for AI might suddenly become tractable? What's next? And how might this change the way you interact with AI, the kinds of assistance you might expect from it in the future? Lots to ponder there. That's it for this deep dive. Thanks so much for joining us. Been a pleasure. We'll catch you next time for another Journey into Knowledge.

From the publisher

This paper introduces **Chain-of-Agents (CoA)**, a novel method for **Large Language Models (LLMs)** to solve complex problems by simulating **multi-agent collaboration** within a single model. Unlike traditional **Tool-Integrated Reasoning (TIR)** methods, CoA allows for flexible integration of various **role-playing agents and tools** in an end-to-end fashion. The research details a **multi-agent distillation framework** and **agentic reinforcement learning (RL)** to train these **Agent Foundation Models (AFMs)**. Empirical studies showcase AFM's **superior performance and efficiency** across diverse benchmarks, including web navigation, code generation, and mathematical reasoning, ultimately making the entire project **open-source** to foster further development in agent models.

More from Best AI papers explained

All 475 episodes
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RLBest AI papers explained · 20 min
Listen in VO