In short
Challenges “agentic AI” hype by arguing that many multi-agent workflows are inefficient clones. The episode explains the KV cache “pre-fill” cost and proposes OneFlow: a single agent that role-switches internally to keep one continuous memory, avoiding repeated context rereading.
Guest backgrounds
No guest names or external backgrounds are provided; it’s a two-host discussion.
Key claims
(1) Homogenous multi-agent systems (same model per role) pay a computational tax because each agent instance rebuilds KV cache from scratch. (2) OneFlow keeps the shared context by switching roles within one model. (3) A “creative designer” and “critical reviewer” use Monte Carlo Tree Search for an optimized thought workflow before execution.
Notable examples
coding benchmarks (HumanEval, MBPP), travel planner with tool use, and shopping typo-to-recommendation (“crinoline incident” on shopping MMLU).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Multi-Agent Systems
1:39 to 3:16
Learn how multi-agent systems operate and why they can be inefficient.
“When a developer says they're building a multi-agent system right now to, say, build an app, what is actually happening under the hood?”
The Problem with Homogeneous Agents
3:17 to 5:48
Discover the drawbacks of using identical agents in multi-agent systems.
“You put a tie on the second one and say, you're the manager.”
Introducing OneFlow: A New Approach
5:49 to 7:08
Uncover the OneFlow concept that optimizes workflows within a single agent.
“You're recalculating the same shared context over and over again.”
The Designer and Reviewer Models
7:09 to 9:16
Explore how OneFlow utilizes creative and critical models for optimal problem-solving.
“How do you make sure it actually switches hats properly and doesn't just blur it all together?”
Benchmarks and Performance Metrics
9:17 to 13:12
Evaluate how OneFlow performs against traditional multi-agent systems in practical tasks.
“And only once that optimized recipe for thought is finalized, then is it handed to the single agent to execute.”
The Role of Heterogeneous Agents
13:13 to 14:03
Understand the scenarios where multi-agent systems remain relevant, especially with diverse skills.
“If you're a startup burning VC cash on OpenAI credits, a fraction means you might actually have a runway.”
Exploring Multi-Agent Workflows
14:03 to 14:48
Learn about the importance of pairing AI models with different strengths.
“But using GPT-40 to write the CSS code for the PDF is just overkill.”
The Power of a Single Agent
14:49 to 15:34
Discover how a well-organized single agent can outperform a multi-agent setup.
“So OneFlow is a homogenous killer, not a total concept killer.”
Practical Applications for Developers
15:35 to 16:15
Understand how to optimize workflows and reduce costs in AI use.
“We build the Rube Goldberg machine because it looks cool, rather than just optimizing the lever.”
The Future of AI Thinking
16:16 to 17:06
Explore the concept of teaching AI to deliberate and self-critique.
“Force the model to critique itself, but keep it all in one continuous context window.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. Today we are taking a sledgehammer to one of the most popular and honestly most expensive narratives in Silicon Valley right now. It's a narrative that really appeals to our human vanity, which is probably why it sticks so well. It really does. I'm talking about agentic AI. If you've been on Twitter or LinkedIn in the last six months, you have seen the hype cycle in just full swing. The pitch is so seductive. If one AI model is smart, a society of AI models, a CEO, a coder, a reviewer must be a genius. Yeah. Right. It's that classic more is better idea. Right. It's the concept of the multi-agent system.
0:39And it feels intuitive because, you know, it mirrors how human companies work. You don't ask one person to build a skyscraper. You hire a firm with specialists. You want the architect, the engineer, the person pouring the concrete. But, and this is why I love the data we're looking at today, we have to ask a very uncomfortable question. Yeah. What if that complex team of AI agents is actually just a really complicated, really slow way of doing something? a single AI could do all by itself. Yeah. And not just complicated, wasteful. So wasteful. So today we are digging into a concept called OneFlow.
1:13It really challenges this assumption that we need a whole society of agents for every complex task. Yeah. We're going to talk about the hidden costs of these AI teams, specifically a little thing called the KV cache. Which is going to be your new favorite buzzword. It's the key to the castle, really. And we're going to find out why talking to yourself, at least if you're an AI, might actually be the smartest move you can make. Let's dive in. Okay. So first, set the scene for us. When a developer says they're building a multi-agent system right now to, say, build an app, what is actually happening under the hood?
1:48Right. Because the marketing makes it sound like they have these little digital people running around inside the server. Well, typically it starts with decomposition. You have a complex goal, like build a snake game in Python. Okay. If you ask a single chatbot to do that in one go, it might fail or hallucinate or forget the scoring system. So you break it up. Dividing conquer makes sense. Exactly. You create a workflow. You might set up an architect agent to write the plan first. Then that output, that plan, it goes to a coder agent to write the script. Then that script goes to a tester agent to look for bugs.
2:21So it's like a virtual assembly line. Theoretically, yes. They pass messages back and forth. The architect outputs text, the coder reads that text and outputs code, the tester reads the code, and it puts a critique. Okay, now here is where the research throws the first punch. In the vast majority of these systems, let's talk about the homogenous setups, which is a fancy word for... All the same. All the same. Who are these agents? They're clones. Expand on that, because homogenous sounds very scientific, but clones sounds like a horror movie. It's kind of true, though. In a homogenous workflow, the architect is, say, GPT-4O.
2:58The coder is also GPT-4O. And the tester is, well, you guessed it, GPT-4O. So it's not like a construction site where you have a carpenter, an electrician, and a plumber who genuinely have different skills. No, not at all. It's like having three clones of the exact same person standing in a circle. Okay. You put a hard hat on the first one and say, you're the builder. You put a tie on the second one and say, you're the manager. But it is the exact same brain. They're just operating under different instructions, different prompts. That's it. I have to play devil's advocate here, though. As a human, if I put on my editor hat, I do think differently than when I have my writer hat on.
3:37Doesn't separating them force the AI to focus? That was the assumption. That's exactly why the industry drifted this way. Right. We thought separation kept things clean. It keeps the context window manageable. You know, if Agent A messes up, Agent B is a fresh pair of eyes. Compartmentalization. But there is this hidden tax to this approach, a computational tax that almost nobody discusses outside of deep infrastructure engineering. Okay. And it all comes down to how these large language models actually read and remember things. This leads us to the KV cache. I saw this in the notes, and I want to make sure we nail this because it sounds technical, but the financial side is just huge.
4:15What is the KV cache? Okay, so think about how an AI processes text. When you send a message to a chatbot, it doesn't just know what you said previously by magic. Right. It has to process the entire history of the conversation to predict the very next word. So it has to reread the minutes of the meeting every single time I speak. Exactly. Now, mathematically, this process creates something called key value pairs. That's the K and the V. Think of the KV cache as the AI's short-term memory of the conversation so far. It's the mental model of what's been discussed. You got it. It's cached, sitting there, ready to go.
4:50So in a single agent conversation, that cache is preserved. You add a new sentence. The AI just updates the cache with the new info. It's fast. It's cheap. But in a multi-agent system. In a multi-agent system, when the architect passes the plan to the coder, the coder is a fresh instance. It's a new clone. It doesn't have the architect's brain state. It only gets the text file, the transcript of the plan. Oh, I see. So the coder has to read the entire plan from scratch to build its own KV cache. Bingo. We call this the pre-fill phase. And in GPU terms, pre-fill is the heavy lifting. It burns compute.
5:28It costs money. Wait, so reading is expensive. Yeah. I assumed writing the answer was the hard part. Writing or decoding is actually slow, but reading huge amounts of context to set up the brain, that eats up bandwidth and memory like crazy. So if I have five agents and they pass the project around five times. You are paying that rereading tax five times. You're recalculating the same shared context over and over again. Wow. It's like if every time you emailed a colleague, they had to reread every single email in the chain from 2015 just to reply to the latest one. That sounds exhausting and incredibly inefficient.
6:05It is. And the one flow argument is that if the agents are homogenous, if they're the same model, this separation is completely artificial. You're just deleting the memory to rebuild it 30 seconds later. So the alternative is the single agent approach. But wait, if I just ask one AI to build the app, test it and fix it, it usually gets confused. It loses the plot. That's why we split them up in the first place. It does. You're right. And that's the catch. A single agent, unstructured, tends to ramble or get stuck in a loop. So we're stuck. Either we pay the efficiency tax for the team, or we deal with a confused solo agent.
6:43That's where OneFlow comes in. The researchers basically said, what if we keep the structure of the team, but execute it inside a single brain? Talking to yourself. Structured schizophrenia, basically. The agent plays role A, generates the output, keeps it in its memory, the KV cache, and then instantly switches to role B. It remembers the context because it was there five seconds ago. No rereading. Done. Okay, I get the efficiency play. That makes total sense. But how do you keep it from getting confused? How do you make sure it actually switches hats properly and doesn't just blur it all together?
7:14You need a rigorous workflow. You can't just wing it. And OneFlow isn't just a prompt. It's an automated system that actually designs the thought process for the agent. Okay, I want to dig into this designer aspect because the data mentions that OneFlow creates a kind of simulation to figure out the best way to think. This is the coolest part for me. Before the agent even tries to solve your problem, OneFlow spins up two specialized meta models. Think of them as the coaches in the locker room devising the game plan. Who are these characters? First, you have the creative designer. This model's job is to look at the task, say, a complex math problem, and propose a workflow.
7:56The visionary. Yeah, the visionary. It might say, hey, let's try a chain of thought approach here and then let's verify it with a Python script. It's just throwing ideas at the wall. Let's get creative. Right. It even looks at error logs from previous attempts to see where things went wrong. Last time we failed at step three, so let's change that. But then its idea has to get past the boss. The critical reviewer. Exactly. The critical reviewer is the accountant and the strict editor wrapped into one. It looks at the designer's flowchart and just hammers it. It's looking for flaws. Logic flaws for sure, but crucially, it checks for token efficiency.
8:33It's looking at the budget. Ah, so it's asking, is this cost effective? It's asking, do we really need three steps here? Can we do it in two? This step looks redundant. Cut it. I love that. So you have one AI generating wild ideas and another AI just crushing them with logic and budget constraints. It sounds like every writer's room in Hollywood. It is, and they iterate. They use a method called Monte Carlo Tree Search. Whoa, flashback to chess computers like Deep Blue stuff. Same principle. They simulate the future, they rub a proposed workflow on a small set of data, see if it fails or succeeds, and then prune the bad paths.
9:07In the study, they let these two argue for up to 20 rounds. So the workflow that eventually gets used, it's a survivor. It's battle-tested. It's the result of an evolutionary process between the designer and the critic. And only once that optimized recipe for thought is finalized, then is it handed to the single agent to execute. So it's not just act like a team. It's follow this extremely specific optimized sequence of thoughts that we proved works best for this specific problem. Precisely. It turns the chaos of thinking into a predictable, efficient algorithm. Okay, let's get to the scoreboard.
9:43We have the efficient single agent running this one-flow optimized pattern. How does it stack up against the expensive society of agents? Because if it's cheap, it's stupid. I don't care. The scoreboard is pretty impressive. They ran benchmarks across coding, math, and even real-world planning. Let's start with coding. That's the bread and butter of agents right now. On benchmarks like Human Evil and MBPP, the single-agent execution didn't just survive. It matched, or in some cases slightly beat, the sophisticated multi-agent frameworks. The solo developer out-coded the committee. Essentially.
10:15And remember, it did this with a fraction of the compute cost. But the most surprising one for me was travel planner. Oh, I know this benchmark. This one is nasty. Yeah. It's not just text. The AI has to use tools. Right. Like find a flight to Paris under$600, book a hotel that allows dogs, and make sure there's a vegan restaurant nearby. It's messy. It requires database lookups, checking constraints, calendars. Usually developers insist you need a flight agent, a hotel agent, and a scheduler to handle that. Because surely a single agent gets overwhelmed with all those tools. That's just too much context to juggle, isn't it?
10:52That's the belief. But the data show that the single agent following that one flow structure matched the success rate of the multi-agent setup. That is so counterintuitive. Why didn't it get overwhelmed? Because of the context. In the team version, the hotel agent might not fully realize why the flight agent picked a specific arrival time. information gets lost in the handoff, or the summary is too brief. The single agent holds the whole itinerary in that KV cache we talked about. It sees the whole picture continuously. It knows it picked the 6 p.m. flight, so it knows not to book a dinner reservation for 5.30.
11:28Exactly. Intuitively, we think more brains equals better checking. But in AI, one continuous brain means better context retention. There's another example in the data about shopping that I thought was hilarious and insightful. The crinoline incident. Yes. The session-based query recommendation. Walk us through that one. So imagine a user on an e-commerce site. They type in crinoline underskirt. Crinoline. Or the typo and everything. Right. The user meant crinoline, you know, those vintage poofy skirts. The task is complex. You have to correct the typo, understand the user is probably looking for vintage fashion, and then predict what they might want to buy next.
12:06Okay. The standard multi-agent approach. You'd have a spellcheck agent, then an intent detector, then a recommender. And they passed the note down the line, losing context each time. But OneFlow optimized a single agent to do it all in a sequence. It played the analyzer, fixed the typo, inferred the intent, and made the recommendation. And how'd it do? On the shopping MMLU benchmark, it actually outperformed the specialized team. It did better. Yes, because it remembered the typo. Why does remembering the typo matter? Because the type E itself is data. It tells you something about the user. Maybe they're unsure of the terminology.
12:43If you just hand the corrected word crinoline to the next agent, that agent loses the nuance that the user was struggling. It's the difference between reading a summary of a conversation and actually being in the conversation. You nailed it. And again, let's look at the cost. Right. The bottom line. In some of these cases, the inference costs dropped by huge margins. We're not talking 10 % savings. We are talking about doing the same quality of work for a fraction of the price, all because you stopped paying the pre-fill tax. That is going to be music to the ears of any CTO listening. If you're a startup burning VC cash on OpenAI credits, a fraction means you might actually have a runway.
13:22It changes the fundamental unit economics of your product. But I have to push back again. I feel like we are beating up on the multi-agent concept pretty hard. Are we saying multi-agent systems are totally dead? Is there never a reason to use a team? No, and this is a really important nuance. We have to distinguish between homogenous and heterogeneous workflows. Okay, so define heterogeneous for us. Heterogeneous means you are using completely different brains, not just clones with different hats. Give me an example. Let's say you are building a tool that writes a legal contract and then formats it into a pretty PDF.
13:58You might use GPT-4O for the writing because it's brilliant and understands law. Okay. But using GPT-40 to write the CSS code for the PDF is just overkill. It's expensive. It's like hiring a lawyer to paint the office. Exactly. So you might pair GPT-40 with, say, Claude 3.5 Haiku, or a specialized small coding model that's fast and cheap. Okay, so now I have a heavy lifter and a sprinter, a real team with different skills. Right. Now here's the technical kicker. You cannot share the KV cache between them. Because they're different brains. They speak different mathematical languages. Their neural networks are weighted completely differently.
14:34You can't just copy the memory from GPT-4O and paste it into Claude. It would look like noise. So in that case, you have to pay the tax. You have to copy paste the text output from one to the other. And the second one has to read it from scratch. Correct. So for heterogeneous tasks where you genuinely need different skill sets, the multi-agent structure is still valid. The cost is justified. So OneFlow is a homogenous killer, not a total concept killer. Well, hold on. Oh, no. The researchers actually ran a pilot study on this, too. They weren't satisfied with just beating the clones. Okay. They pitted a heterogeneous team, a mix of big and small models, against the optimized OneFlow single agent.
15:16Don't leave me hanging. The single agent held its own. Seriously. It turns out that a single, very smart model, when organized perfectly by OneFlow, is so capable that it often negates the need for the little specialist helper. We are overcomplicating things before we've even maximized the simple solution. It's the classic engineering trap. It is. We build the Rube Goldberg machine because it looks cool, rather than just optimizing the lever. We are trying to build artificial societies before we've even finished building the artificial individual. That's a great way to put it. So let's bring this all the way down to earth.
15:50If I'm a listener, maybe I'm a developer, maybe I'm managing a product team, what do I do with this information tomorrow? You should audit your agents. Look at your workflows. If you are chaining together three instances of GPT-40 just to check each other's work, you are almost certainly burning money for no performance gain. Try talking to yourself first. Try prompt engineering the internal monologue. Use the OneFlow principle. Define the steps. Force the model to critique itself, but keep it all in one continuous context window. It really shifts the perspective from AI as a chatbot to AI as a thinkbot.
16:28That's it. It leads me to a slightly existential thought to wrap this up. We are basically teaching AI to have an internal monologue or a conscience. Well, OneFlow works by having that designer and critic simulate a debate. We are teaching the AI to second guess itself, to look for errors, to optimize its own thoughts before it ever speaks. We're moving from prediction to deliberation. Exactly. That's the frontier, isn't it? The next leap in AI isn't just bigger models. it's models that know how to think before they speak. One flow is just the first step in structuring that silence. Structuring the silence.
17:05I like that. We will leave you with that thought. Go check your API tools, folks. You might be paying for a lot of redundant reading. Thanks for listening. Catch you on the next Deep Dive.
From the publisher
The provided text explores whether multi-agent systems (MAS) can be effectively replaced by a single agent simulating complex workflows through multi-turn conversations. Research indicates that homogeneous workflows, where multiple agents use the same base model, can be replicated by one agent with significant computational efficiency gains via KV cache reuse. The authors introduce OneFlow, an automated algorithm that utilizes dual meta-LLMs and Monte Carlo Tree Search to design streamlined, high-performance workflows specifically for single-agent execution. Experimental results across various benchmarks demonstrate that this single-agent approach matches the accuracy of multi-agent setups while reducing inference costs. However, the study acknowledges that heterogeneous workflows involving different base models still offer unique benefits that a single model cannot yet fully capture. Consequently, these findings establish the single-LLM implementation as a powerful new baseline for future multi-agent research.




