Latent Collaboration in Multi-Agent Systems

29 Nov 2025 · 13 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Collaboration in Multi-Agent Systems (Latent MAS) replaces slow, error-prone text-based communication (“TextMAS”) with instantaneous sharing of continuous latent “thoughts” between agents inside an LLM’s latent space.

Guest backgrounds

No guests are mentioned; it’s a solo host “Deep Dive” episode.

Key claims

TextMAS must decode agent reasoning into tokens and re-encode it, causing major token/cost/speed overhead and compounding errors. Latent MAS uses auto-regressive latent thought generation plus lossless working-memory transfer via the model’s KV cache (layer-wise concatenation). Reported efficiency: 235–471x more efficient reasoning; 70.8–83.7% fewer output tokens; 4x–4.3x faster end-to-end (2.6x–7x vs VLLM-optimized baselines); accuracy gains 2.8% sequential, up to 4.6% hierarchical, max 14.6%.

Notable examples

multi-step math (Amy Math; <50 latent steps vs >20,000 tokens in TextMAS); GPQA Diamond speed/quality improvements.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Challenges of Text-Based Communication in AI

0:46 to 1:48

Discussion on the inefficiencies and limitations of text-based communication in multi-agent systems.

“relies on natural language, just plain old text as the communication medium.”

Understanding Latent MAS

1:49 to 3:52

Explains the concept of latent MAS and its advantages over traditional methods.

“So let's start by framing the problem with text mayas.”

Efficiency Gains of Latent MAS

3:53 to 5:43

Latent MAS offers up to 471 times more efficiency than traditional methods.

“And that flawed text constrains all the downstream agents.”

Mechanics of Latent Communication

5:44 to 7:40

Describes how agents share thoughts directly within latent space without text conversion.

“it's difficult to even conceptualize in normal computing terms.”

Revolutionary Cost and Speed Improvements

7:41 to 9:28

Latent MAS drastically reduces token usage and increases processing speed.

“Agent B is immediately and continuously conditioned on Agent A's complete internal thinking.”

Quality of Problem Solving in Latent MAS

9:29 to 12:11

Latent MAS improves accuracy and mitigates error compounding compared to text-based systems.

“In the world of high-scale inference, time saved is exponentially valuable.”

The Future of AI with Latent MAS

12:12 to 13:13

Explores the potential of latent MAS in reshaping AI architectural designs and future research.

“So Leighton-MAS really is a genuine paradigm shift.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're cutting through the noise around large language models to look at one of the most critical breakthroughs in scaling AI intelligence. collaboration. For a while now, you've heard about multi-agent systems or MAS, where we create these little teams of specialized LLMs, a planner, a critic, a solver to tackle problems far too complex for any single model. And this idea of coordinated system-level intelligence is really the holy grail. It is. We know collaboration boosts performance, especially on really difficult tasks like, say, multi-step math problems or advanced coding.

0:37But there has always been this fundamental speed limit on these systems. And that's how the agents talk to each other. Exactly. The dominant method, what researchers are calling text mares, relies on natural language, just plain old text as the communication medium. Which, you know, on the surface makes perfect sense. These are language models. Language is their native kong. But if the goal is maximum efficiency, maximum fidelity, relying on written communication between two pieces of software sounds, well, cumbersome. Cumbersome is an understatement. The text MAS is painfully slow, it's resource intensive, and it's actually prone to compounding errors.

1:14And that's why we're diving into this groundbreaking work from Princeton, Illinois, and Stanford on a new paradigm, latent MAS. This is an end-to-end framework that enables pure instantaneous collaboration directly inside the model's continuous latent space. So we're moving from them writing complex memos to each other to just sharing pure instantaneous thoughts. That feels like the future. It really is. Our mission is clear then. We need to unpack how Leighton Mayas works, understand why it's so much better theoretically, and then review the truly stunning performance gains the researchers reported.

1:48I mean, in speed, cost, and the final reasoning quality. So let's start by framing the problem with text mayas. If these agents are built on language, why does using natural language become a liability? Why can't they just, you know, pass notes efficiently? Well, the issue boils down to conversion, complexity, and just sheer volume. Yeah. Think about the internal life of an LLM. It processes information as these massive, complex, high dimensional vectors. Right. That's the latent space. That's the continuous latent space where all the rich context and meaning actually lives. You could think of it as a really detailed, like a complex 3D model of an idea.

2:26Okay, a 3D model of an idea. I like that. So in a text Maya system, when Agent A finishes its internal reasoning, it has to take that high-fidelity thought vector, that 3D model, and fully decode it. Into just a flat line of text. Into a linear sequence of discrete, separate tokens, explicit text. It's forcing that rich 3D model to be translated into a long, often very verbose, instruction manual. And then the next agent, Agent B, gets this instruction manual. Precisely. Agent B, say, the critic, then has to take that lengthy text and laboriously re-encode it back into its own high-dimensional latent space so it can even start to process the information.

3:03That whole round-trip decode, past text, re-encode, it just introduces this massive computational overhead. Huge overhead. Every single interaction requires vast token usage, which directly translates to massive costs and delays in getting the task done. That inefficiency is jarring when you think about it. It's about two supercomputers communicating using a slow shared fax machine instead of a direct data bus. That's a great analogy. But the inefficiency isn't just about speed. It's also about quality. One of the biggest findings from analyzing text may us is this problem of error compounding.

3:38Oh, okay. Because the communication is text and text is discrete, it's brittle. So if the first agent, the planner, makes a slight misinterpretation, maybe it phrases a key step incorrectly in its chain of thought output, that error becomes explicit. it. It becomes permanent. It's written in stone. Yes, exactly. And that flawed text constrains all the downstream agents. The critic and the refiner, they're forced to base their reasoning on this brittle, potentially flawed data they just spent a ton of time in tokens re-encoding. It's a huge problem on multi-step tasks. Okay, so let's jump to the solution.

4:12Latent MES proposes the agent stop writing notes and start sharing thoughts. If they aren't talking via text, how do they communicate directly within that continuous latent space? What's the core mechanism here? The shift is architectural, and it really centers on two massive innovations. The first is how each model reasons internally, what they call auto-regressive latent thoughts generation. So instead of generating words, they're generating pure thought vectors. They're staying in that high fidelity internal state. That's the mechanism. Instead of outputting explicit tokens, reasoning unfolds by just auto-regressively appending the hidden representations from the final transformer layer back into the input sequence.

4:54And these hidden representations are the latent thoughts. Those are the latent thoughts. And because they are continuous vectors, each latent step can convey vastly richer and more diverse information than any single discrete token ever could. It's a huge increase in communication bandwidth. That increase in bandwidth suggests an almost revolutionary gain in efficiency right off the bat. If you're transmitting Cure high-dimensional meeting instead of symbols, the data volume has to just plummet. It does. And the researchers back this up with theory. They showed that this latent thought generation can be orders of magnitude more efficient.

5:31In practical terms, for the QIN3 models they tested, it makes the reasoning process anywhere from 235 to 471 times more efficient than generating text. Wait, say that again. Up to 471 times more efficient? 471 times, yes. That number is... it's difficult to even conceptualize in normal computing terms. It completely changes the economics of running these complex systems. It absolutely does. Now, to make this whole loop work where the model feeds its own output back to its input, they had to make sure it was stable. You can't let the model's internal data distribution drift. Right. It would just become noise.

6:08It would. So they introduced a small technical safeguard, a linear alignment operator just to keep things consistent. A necessary consistency check. Yeah. Keeps the system from flying off the rails during this high-speed self-talk. Exactly. But this only addresses how one agent thinks. The crucial step is the MAS part, the multi-agent part. How do you transfer that huge, continuous, latent thought to Agent B and do it losslessly? This is the second innovation, the actual communication channel. Yes. Lossless latent working memory transfer. They realized the perfect shared medium was already built into the LLM architecture, the key value cache or KV cache.

6:46The KD cache, right, which usually acts as a temporary memory bank to speed up attention and decoding. It stores the context of everything the model has seen so far. And that's precisely why it's the perfect medium for shared memory. The KV cache encapsulates the model's entire attention history and, critically, the entire path of the new latent thoughts. It is the purest possible representation of that agent's working memory. So instead of trying to translate the memory into text, they just share the memory unit itself. That's the breakthrough. When Agent A finishes reasoning, its entire layer-wise KV cache, the full working memory, is extracted.

7:21Agent B, the next one in line, then integrates this memory just by layer-wise concatenation. It simply propens Agent A's memory right onto its own. Wait, hold on. So Agent B isn't just getting the conclusion. It's getting the complete internal thought trajectory of Agent A. And it doesn't have to decode or reencode anything. It just accepts the raw memory. That's the key difference. It's truly lossless. Yeah. Agent B is immediately and continuously conditioned on Agent A's complete internal thinking. There's zero time spent on conversion, zero token cost for transferring the reasoning, and zero loss of semantic fidelity.

7:57The agents are tightly integrated at the level of their core data structures. That theoretical complexity reduction is just stunning. But the proof is always in the performance. So here's where we get into the real-world impact across those nine benchmarks. Math, science, coding. What does this actually mean for cost and speed? The empirical results validate the theory completely. So let's look at cost, which is basically token usage. Across the board, latent Mayas slashed output token usage compared to TexMAS by huge margins. We're talking 70.8 % to 83.7%. Over three quarters fewer tokens to solve the exact same problem.

8:32That just fundamentally changes the unit economics of advanced AI. It does. And look at the really reasoning intensive tasks like the Amy Math benchmarks. TexMAS would often have to output more than 20 ,000 tokens for a full verbose chain of thought solution. Just passing texts back and forth. 20 ,000. Layton Mayes achieved comparable accuracy using fewer than 50 Layton steps. That is not an incremental improvement. That's a different order of magnitude entirely. I saw a surprising detail in the data. Layton Mayes actually used 50 % to 60 % fewer tokens than even a single large agent working alone.

9:05How is a team of agents using fewer tokens than a solo agent? That's the efficiency loop working perfectly. In TextMayes, every agent has to output a huge amount of text. In Leighton Mayes, the problem solving is distributed among agents communicating silently in latent space. Only the final agent has a job of decoding the minimal refined answer back into text for you. The quiet consensus. They solve it internally, keep the collaboration silent, and only externalize the final result. Okay, let's talk speed. In the world of high-scale inference, time saved is exponentially valuable. What are the speed-up metrics here?

9:39On average, across all settings, LatentMAS provides a 4x to 4.3x faster end-to-end inference speed compared to traditional TexMAS. And that includes the overhead of that linear alignment operator we mentioned. Four times faster on average is powerful, but what about when the TexMAS baseline is already super optimized, you know, with services like VLLM? Does LatentMAS still beat those? That's where it gets truly transformative. Even when they accelerated the text-may-ass baselines with high-performance systems like VLLM, late-may-ass maintained a massive speed advantage, yielding a speedup of 2.6x up to an astonishing 7x.

10:14Seven times faster. On rigorous benchmarks like GPQA Diamond, that implies the sheer inefficiency of converting thought to text is the primary bottleneck. Removing it is the key. Exactly. It shows that even the most optimized text systems cannot overcome the fundamental computational impedance mismatch created by that decode-encode cycle. So late May-S is significantly cheaper and dramatically faster. But the final crucial question is quality. Is this continuous high-speed thought exchange actually better at solving problems than structured tokenized text? And the answer is a resounding yes. This is maybe the most important finding.

10:51Late Maes consistently outperformed text Maes in accuracy, yielding an average gain of 2.8 % in sequential settings and up to 4.6 % higher in hierarchical setups. And what was the maximum gain? On some of the more challenging benchmarks, the maximum observed gain reached up to 14.6 % higher accuracy. Wow. Okay, so why? Why does receiving a continuous thought vector lead to better quality than receiving clean, discrete text? It goes directly back to that idea of error compounding and expressive capacity. The qualitative analysis they did confirmed that the latent thoughts are semantically consistent with correct text.

11:25But here's the key. Latent thoughts offer greater diversity and expressive capacity. So text is explicit and fixed. But a latent thought maintains nuance and even ambiguity. You've got it. When Agent A makes an error in text maus, the text locks in a brittle mistake that is hard for the critic to undo. But when agent A shares its working memory in latent form, that vector contains richer, more nuanced data. It allows for deeper reinterpretation. Precisely. That inherent nuance means the downstream agents, the critic, the refiner, are far better equipped to reinterpret, refine, and correct upstream reasoning mistakes effectively.

12:04The latent mechanism mitigates error compounding by providing a high-fidelity internal signal instead of a brittle external memo. It allows for a much smoother course correction. So Leighton-MAS really is a genuine paradigm shift. We're moving beyond the structural and economic limits imposed by natural language in these multi-agent systems. We're turning a collection of independent LLMs into tightly integrated, highly expressive components capable of instantaneous coordination. So what does this all mean, really? I mean, this completely changes the architectural design space for high-performance AI.

12:36We're no longer limited by the speed of reading and writing. We are operating at the speed of thought. And crucially, this latent MAS framework is training-free. Existing LLMs can adopt it right now. However, the exciting territory for you to explore is how future research will build on this. Currently, complex post-training paradigms are used to optimize text-based MAS, often refining how agents use language. If those advanced optimization techniques were adapted specifically to refine and optimize latent MAS's latent collaboration protocols, we could unlock far more effective system-level reasoning strategies than we've even seen in this initial research.

13:13It's collaboration truly optimized.

From the publisher

This paper proposes LatentMAS, a novel, training-free framework designed to improve the collaboration efficiency of Large Language Model (LLM)-based multi-agent systems (MAS). Unlike traditional approaches that use explicit natural language, LatentMAS facilitates communication and reasoning entirely within the **continuous latent space** of the models. This is achieved through **auto-regressive latent thought generation** inside each agent and **lossless latent working memory transfer** across agents via shared KV caches. The experimental results demonstrate substantial computational benefits, including **4x to 4.3x faster end-to-end inference** and a significant reduction of **70.8% to 83.7% in token usage** compared to text-based MAS baselines. Furthermore, the system consistently achieves **higher system-level reasoning accuracy**, indicating that collaboration using continuous latent representations offers greater expressive capacity than discrete tokens.

More from Best AI papers explained

All 475 episodes
Latent Collaboration in Multi-Agent SystemsBest AI papers explained · 13 min
Listen in VO