Thought Communication in Multiagent Collaboration

27 Oct 2025 · 17 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Thought communication (ThoughtCom/TCOM) for multiagent collaboration, enabling LLM agents to share latent “intent drivers” directly instead of exchanging language tokens, avoiding language bottlenecks (slow, ambiguous, lossy).

Guest backgrounds

No guests are named in the transcript.

Key claims

Language is a major bottleneck for hive-mind-style coordination; ThoughtCom can provably recover true shared/private latent factors using identifiability theory with sparsity regularization; it preserves cognitive diversity and identifies who-thinks-what structure; it injects recovered thoughts back via prefix adaptation without full model fine-tuning.

Notable examples

Airport scenario (car vs train reasons: luggage vs punctuality, both sharing speed). Results: 19.06% relative improvement over prior multiagent fine-tuning; e.g., 93% accuracy on GSM-8K with a cited model; stable performance with more debate rounds and varying prefix length.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Thought Communication

0:46 to 2:20

Exploration of how thought communication allows LLM agents to collaborate without language.

“We think of language as like the peak of communication.”

Bottlenecks of Language in AI

2:21 to 3:18

Discussion on the inefficiencies and limitations of using language for AI communication.

“We need to understand the mathematical proofs that guarantee they're getting the real internal stuff.”

The Framework of Thought Communication

3:19 to 5:19

Overview of how the Thought Communication framework enables direct sharing of latent thoughts.

“Yeah, the airport one is perfect for illustrating this.”

Theoretical Foundations and Guarantees

5:20 to 8:00

Insight into the mathematical proofs and guarantees that support the Thought Communication system.

“They use a mathematical technique called sparsity regularization.”

Practical Implementation: Connecting Agents

8:01 to 11:33

Details on how the Thought Communication framework is implemented in real systems.

“The actual ThoughtShitcomM framework uses something called a sparsity-regularized autoencoder.”

Results and Performance Metrics

11:34 to 13:20

Results from tests showing the effectiveness of Thought Communication in improving AI collaboration.

“It makes the whole idea seem much more viable.”

Consensus and Accuracy in AI

13:21 to 13:57

Discussion on how Thought Communication achieves both higher accuracy and meaningful consensus.

“Yes, this is super important for understanding why it works better.”

Robustness and Stability of ThoughtCom

14:02 to 15:05

Learn how ThoughtCom maintains accuracy and consensus in varying conditions.

“The system also seemed really robust in the tests.”

Potential Applications Beyond LLMs

15:07 to 15:34

Discover how the core theory can apply to various AI models and data types.

“The researchers suggest the core theory could still hold even if you don't have access to the deep internals.”

Implications for Human-AI Collaboration

15:35 to 16:27

Explore the future of AI teamwork versus human communication methods.

“It could become a kind of universal communication layer operating beneath the surface language or modality.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. Today, we are tackling an idea that sounds pretty much like pure science fiction, honestly. Right. But it's actually becoming a technical reality. We're looking at AI telepathy. Specifically, this radical new paradigm called thought communication or thought TCOM. It lets large language model agents work together without, well, without using language at all. Yeah, exactly. And it's kind of crucial. You know, if you want LLM agents to really achieve that superhuman collective intelligence. Like a hive mind? Sort of, yeah. Like a hundred brains acting like one genius entity.

0:39They just can't be stuck with the same messy, inefficient communication we humans use. And that constraint is language. That's kind of amazing. It is surprising, isn't it? We think of language as like the peak of communication. Right. But the research argues it's actually a major bottleneck, even for these super advanced AIs. Okay, so for the agents themselves, when they're trying to talk using words, what are the big problems? What goes wrong? Well, it really boils down to about five key things. First, natural language is slow, right? It's sequential, one word after another. Okay. It's also ambiguous.

1:12Words can mean different things. It's imprecise. It's indirect. And maybe most importantly, it's incredibly lossy. Lossy. Lossy, meaning information gets lost in translation. Totally. Think about it. When an LLM spits out a sentence, that sentence is just a tiny, maybe even distorted echo of this vast internal reasoning process that produced it. Oh, okay. So when you have, say, two LLM agents in a multi-agent system, an MAS, and they're just swapping tokens or maybe beddings, they're basically just trading these faint echoes, not the real substance. Precisely. And the researchers found that most of the time when these AI teams fail to collaborate effectively, it's because the messages are too vague or the agents aren't quite aligned on what they mean or want.

1:56And both those problems, they come straight back to the indirect lossy nature of language. It's like trying to build a complex machine using only vague instructions passed through a noisy phone line. That's a great analogy. Exactly. It's just not efficient or reliable enough for high stakes collaboration. OK, so our mission for this deep dive then is to figure out how these researchers managed to just bypass that noisy phone line entirely. Right. How did they pull out the actual pure thought from the AI? Yeah. We need to understand the mathematical proofs that guarantee they're getting the real internal stuff.

2:31That's your intent. And obviously look at the results. Did this mind to mind thing actually work in practice? Absolutely. So the big shift, the paradigm shift, is enabling agents to interact directly, mind to mind, sharing their raw latent thoughts. Latent thoughts. The hidden stuff. Exactly. Let's call them their internal intent drivers maybe instead of just sharing the surface level words or tokens. Okay. And this allows for like instant direct transfer of purpose and understanding, no loss in translation. Internal intent drivers. I like that. It makes sense. We're not talking about the final sentence they type out, but the raw map of goals, beliefs, reasoning, the stuff deep inside the model.

3:11That's it, exactly. The sentence is just the linguistic wrapping paper, you know? Wow, yeah. If we can recover those raw drivers, those thoughts, we just skip the whole messy language filter. Okay, so give us an example. Sources had a good one, right? The airport scenario. Yeah, the airport one is perfect for illustrating this. Imagine a user asks two different LLM agents, hey, what's the best way to get to the airport? Standard question. One agent says take a car, the other says take the train. Right. On the surface, two simple answers. But why did they choose those? Their underlying reasons, their intent drivers are way more complex.

3:48Okay, break it down. What's going on under the hood? So agent one, the car agent, might be heavily weighing a private factor like I need to carry lots of luggage. That's unique to its perspective, maybe. But it's also thinking about a shared factor, something both agents care about, like speed, getting there fast. Right. And agent two, the train agent. Agent two might have its own private thought. Maybe schedule punctuality is critical. Trains are usually on time. Makes sense. But it also cares about that shared factor, speed. So ThoughtCom's goal isn't just to share car or train. it's to instantly share the importance and the structure of these deeper factors, luggage, speed, punctuality.

4:29Wow. Okay. That is a huge difference. You're not just getting the answer, you're getting the why, the context, the reasoning, like instantly. Exactly. But here's where I get a bit stuck. If these thoughts are latent, hidden deep inside, how can we be absolutely sure that what SOCOM pulls out is actually the true useful reason? Yeah, that's the critical question. How do they avoid just extracting, I don't know, random internal noise or some irrelevant model state? Right. And that is honestly the most brilliant part of this work, the theoretical foundation. The researchers had to build and prove a really solid identifiability theory.

5:07Identifiability, meaning proof you could find the real thing. Exactly. It means they have mathematical proof, like a guarantee that the method they use can recover the true latent thoughts, separating them from noise and everything else. Okay, so it's not just a fancy algorithm. There's real math backing it up. How does that work? Is it like a filter? Kind of, yeah. They use a mathematical technique called sparsity regularization. You can think of it like a cleaning filter or a constraint that encourages the system to find the simplest, most distinct underlying factors. Okay, sparsity forces it to be clean and clear.

5:39Precisely. And this whole framework provides three essential guarantees. Things you absolutely need for this AI telepathy to work reliably. All right, lay them out for us. What are these guarantees? Okay, guarantee number one, shared understanding. The theory proves that common concepts like speed in our airport example get cleanly pulled out and separated from everything else. Disentangled. Yeah, disentangled is the technical term. Separated from private thoughts or irrelevant stuff. This is vital for creating what they call a faithful common basis, making sure when agents talk about speed, they actually mean the same thing.

6:14Okay, so that handles agreement, making sure they're on the same page for shared ideas. But good teamwork isn't just about agreeing, right? You need individual insights, too. What about those unique private thoughts? Excellent point. And that's the second guarantee, preserving cognitive diversity. The math also ensures that those unique agent-specific factors like carrying luggage for the car agent or schedule punctuality for the train agent are also identified and kept separate. Ah, so you don't just average everything out. Exactly. The source material emphasizes this. Just forcing everyone to agree creating homogeneity is actually bad for complex problem solving.

6:54You need those diverse perspectives. You don't want to throw away a rare but critical insight just because only one agent thought of it. That makes sense. You preserve the potential for innovation. But isn't there a risk there? If agents start sharing all their private niche thoughts, couldn't that just flood the conversation with noise? Distractions. Another great question. And that brings us to the third guarantee. Identifying the structure. It's not just about what the thoughts are, speed, luggage, etc. It's also about knowing the structure, who holds which thoughts, and how strongly. Oh, interesting.

7:27Like an organizational chart of thinking. Kind of. The framework can identify this pattern, this structure. It tells Agent A, okay, Agent B is really focused on this luggage aspect, while Agent C cares more about cost. Knowing the structure is fundamental for smart, adaptive coordination. Agent A can then weigh Agent B's input on luggage more heavily, for example. Okay, wow. So the theory sounds really solid. You can reliably pull out shared thoughts, private thoughts, and the structure of who thinks what. But how do you actually build this thing? How do you hook this telepathy machine into a working LLM?

8:00Right. Moving from theory to practice. The actual ThoughtShitcomM framework uses something called a sparsity-regularized autoencoder. Okay, autoencoder. That's a type of neural network. Exactly. You feed this autoencoder the combined outputs, the responses, or internal stays from all the agents working together. So you feed it the wrapping paper from our earlier analogy. You got it. And the autoencoder acts like a translator. It learns to map that messy world of language output back to the clean, structured, latent world of the underlying thoughts, the intent drivers. The sparsity regularization is key here, making sure the translation sticks to the guarantees from the theory.

8:37So the autoencoder bridges the gap between the words and the real meaning. Perfectly put. Yeah. And once it recovers these latent thoughts, it doesn't just, like, broadcast everything to everyone that would be inefficient, maybe noisy, like you said. Right. Instead, they use a really clever mechanism based on agent agreement. Each recovered thought gets a score based on how many agents' internal states seem to depend on it. Ah. So it figures out how widely shared each thought is. Yep. Is it a consensus thought everyone shares? Is it specific to just two agents? Or is it a completely private thought from one agent?

9:09This allows for customized thought sharing. Personalized telepathy. Basically, yeah. It makes sure agents only receive the thoughts most relevant to improving the collaboration could be shared, could be agent-specific, could be a crucial private insight while filtering out the distracting noise. It's truly adaptive. Okay, that's smart. Yeah. So you extract the thoughts, filter them by relevance and agreement. Then how do you get them back into the agents to influence their next step? How does Agent A use Agent B's thought? They use a technique called latent injection via prefix adaptation. It sounds complicated, but the idea is pretty neat.

9:46Go on. They take the selected relevant thoughts, which are basically just vectors of numbers at this point, and they use a small learned adapter network to convert them into a very short prefix vector. A prefix, like adding something at the beginning. Exactly. This short thought prefix vector gets prepended, stuck right onto the front of the normal token embeddings that the LLM would usually process for its next turn. Huh. So you're not rewriting the agent's brain or anything. No fine-tuning of the whole model needed. You're just giving it this little, like, precognitive nudge, a hint derived from the group's collective thought space to guide its next generation.

10:23That's a perfect way to describe it. A subtle nudge based on distilled collective insight. And this approach, using prefixes, it leads to a huge efficiency win, right? That seemed like a big deal in the paper. Absolutely crucial for making this practical. The computational cost, the overhead of running Thotikom depends only on the LLM's embedding dimension. Okay, the size of those vectors representing words or concepts. Right. And crucially, that embedding dimension often stays the same even when the model gets way bigger. So think about scaling up. You go from a, say, 70 billion parameter model to a monster 405 billion parameter model.

11:02Yeah. The embedding dimension might not change much, if at all, which means the cost of running Thought TECOM stays roughly the same. Whoa. So the communication cost doesn't explode as the models get bigger, unlike, say, trying to fine-tune the whole giant model for collaboration. Exactly. Traditional fine-tuning scales with the model size, which gets incredibly expensive. Thoughtteacom gives you this efficient, high-speed communication channel, like a dedicated Neuralink, whose cost is basically fixed regardless of how massive the base LLM is. That's a massive economic advantage for building large-scale multi-agent systems.

11:37It makes the whole idea seem much more viable. Totally. It makes the framework model agnostic in that sense and way cheaper than constantly retraining huge models. Okay. Theory is solid. Implementation is clever and efficient. Let's talk results. They ran tests, right? Did it actually work? I know they had synthetic tests. Yeah, the synthetic tests were great. They confirmed the theory worked, like showing they could cleanly separate shared and private thoughts, got high scores on metrics confirmed, and they identified the true structure. All the theoretical boxes ticked. But the real test is complex real-world problems.

12:10What happened when they threw thought to you, Comet, stuff like hard math problems? That's where it got really exciting. They tested it on tough benchmarks like math and GSM-8K standard tests for mathematical reasoning. They compared ThoughtCom against, you know, just having one agent try, single answer baseline, and also against the best previous multi-agent method, which involved fine-tuning. Okay, the state-of-the-art linguistic approach, what was the bottom line? ThoughtCom consistently just, well, it blew the baselines out of the water. Oh, really? By how much? On average, across the tasks, it delivered a 19.06 % relative improvement over that state-of-the-art multi-agent fine-tuning baseline.

12:5019%. That's huge in AI benchmarks. It really is. And to give you a concrete example, using one specific model, the Quinn 3-per-21-titty B ThoughtCom hit 93 % accuracy on that hard math benchmark. The best the linguistic baseline could do was significantly lower. ThoughtCom gave an absolute gain of 17.2 % in accuracy. Okay, that's not just tweaking parameters. That's a fundamental jump in capability driven purely by better communication. Exactly. And there was something interesting in the results about agreement, right? Consensus versus accuracy. Yes, this is super important for understanding why it works better.

13:25Often, with older multi-agent systems, you could get the agents to agree more increased consensus. Okay. But sometimes they'd all just confidently agree on the wrong answer. It's a known failure mode. Right. Groupthink, but for AI. Pretty much. But ThoughtTCOM, because it's sharing the actual underlying reasoning, achieved both higher accuracy and higher consensus. The agreement was meaningful. It led to better outcomes. So the shared thoughts actually helped them align on the correct path, not just any path. Precisely. It shows superior alignment that translates directly to better task performance.

14:02The system also seemed really robust in the tests. Like, it wasn't fragile. Yeah, very robust. They tried increasing the number of debate rounds, going from two rounds up to six. Okay, letting them talk more. With the baseline methods, performance often got worse, with more rounds, more noise, more redundancy, agents getting confused. Makes sense. But, ThoughtCom, its accuracy and consensus stayed stable, rock solid. It didn't degrade because it wasn't just repeating noisy messages. It was refining based on pure shared signal. That's impressive. Immune to the too much chatter problem. And they also played with the length of that prefix vector, the injected thought.

14:39They varied it 16-fold from length 1 to 16. Performance stayed stable again. It delivered the gains without needing super precise, painful tuning of that parameter, which is often a headache in complex systems. Okay, so the overall picture is pretty clear then. Moving away from language, moving to these provably identifiable latent thoughts, It just makes AI collaboration stronger, way more efficient, and more accurate. That's the core takeaway. It's a fundamental shift. You mentioned it's potentially broader than just LLMs communicating via internal states. Right. The researchers suggest the core theory could still hold even if you don't have access to the deep internals.

15:17You might be able to use, say, sophisticated embeddings of the text outputs themselves as a proxy for the internal states. Ah, so you could potentially apply this to coordinate closed source models like GPZ-4 or Claude, where you only see the text, or even models dealing with images or other data. That's the potential, yeah. It could become a kind of universal communication layer operating beneath the surface language or modality. Which really brings us to our final provocative thought for you, the listener, to chew on. If these LLM agents can achieve this level of almost superhuman collaboration, bypassing language, directly sharing complex, nuanced, even private thoughts, all backed by mathematical guarantees.

15:59Yeah. What does that really imply for the future of human AI collaboration? When we humans are still stuck with our relatively primitive, lossy biological ways of communicating. Will AI teams fundamentally outperform human teams or even human AI teams just because their communication is so much better? Exactly. Are we heading towards a future where AI collaboration is just inherently fundamentally superior because they ditch the limitations of language that we can't escape? It definitely raises some fascinating, maybe slightly unnerving questions about the future of intelligence, both artificial and collective.

16:34Indeed. Something to think about. We'll leave that with you. Thanks for joining us for the Deep Dive.

From the publisher

The academic paper proposes "thought communication," a new paradigm for multi-agent collaboration that allows large language models (LLMs) to exchange latent thoughts directly, akin to telepathy, instead of relying on lossy natural language. The authors formalize this process using a latent variable model where agent states are generated from underlying thoughts, proving that both shared and private thoughts can be mathematically identified. Guided by this theory, the proposed THOUGHTCOMM framework uses a sparsity-regularized autoencoder to extract these latent thoughts and their structural dependencies, allowing agents to efficiently receive personalized, relevant cognitive information. Experimental results on math reasoning benchmarks confirm that this direct, mind-to-mind communication significantly enhances collaborative accuracy and consensus compared to existing language-based multi-agent systems. The work suggests that leveraging these hidden internal representations is critical for achieving superhuman collective intelligence in machines.

More from Best AI papers explained

All 475 episodes
Thought Communication in Multiagent CollaborationBest AI papers explained · 17 min
Listen in VO