LatentMAS: Breaking the Text Bottleneck for Multi-Agent Intelligence
Latent Collaboration in Multi-Agent Systems
LatentMAS is an end-to-end, training-free framework for multi-agent systems (MAS) that enables Large Language Models (LLMs) to collaborate directly within the continuous latent space. By replacing text-based communication with latent thoughts and shared KV-cache working memory, it achieves state-of-the-art performance across 9 benchmarks including math, science, and coding.
TL;DR
Multi-agent systems (MAS) have long been constrained by the "textual bottleneck"—the slow, lossy process of agents talking to each other through natural language. LatentMAS changes the game by enabling agents to "think" and "collaborate" entirely in the continuous latent space. The result? A framework that is 4x faster, uses 80% fewer tokens, and is actually more accurate than traditional text-based systems—all without any retraining.
The "Lingua Franca" Problem in Agentic AI
In the current LLM landscape, agents communicate like humans: they write out thoughts, pass them as text, and the next agent re-encodes that text. This is fundamentally inefficient.
- Discretization Loss: Forcing a high-dimensional internal representation into discrete tokens loses semantic nuance.
- Computational Waste: Decoding and re-encoding steps consume significant FLOPs and time.
- Inflexible Reasoning: Natural language is linear, while latent thoughts can be O(dh/log|V|) more expressive.
LatentMAS asks a radical question: Can agents collaborate directly through their internal "brains" (hidden states) instead of their "mouths" (text outputs)?
Methodology: High-Fidelity Latent Collaboration
The technical core of LatentMAS rests on two pillars: Latent Thought Generation and Latent Working Memory.
1. Auto-regressive Latent Reasoning
Instead of generating tokens, each agent generates a sequence of last-layer hidden embeddings. To prevent the model from getting "confused" by these high-level embeddings when fed back as input, the authors use a clever Alignment Operator (). Using a pseudo-inverse mapping, realigns the output hidden states back into the input embedding space.
Figure: The LatentMAS Pipeline showing internal thought generation and cross-agent memory transfer.
2. Lossless Memory Transfer
In text-based MAS, a "Critic" agent reads a "Planner's" text. In LatentMAS, the Critic inherits the Planner's layer-wise KV-cache. This "Latent Working Memory" ensures that the successor agent's computation is conditioned on the exact internal state of the predecessor. Theorem 3.3 in the paper proves this is mathematically equivalent to re-processing the entire sequence, but without the re-computation overhead.
Empirical Superiority: Faster, Cheaper, Better
The results across 9 benchmarks (GSM8K, HumanEval+, GPQA, etc.) are striking.
- Inference Speed: LatentMAS consistently delivers a 4.3x speedup because it skips the expensive token-by-token decoding process in intermediate steps.
- Efficiency: Total token usage drops by over 70%. Intermediate agents generate zero text; only the final agent decodes the actual answer.
- Accuracy: In complex tasks like AIME25 and MBPP+, LatentMAS sees up to 14.6% improvement.
Figure: Performance comparison across accuracy, speed, and token usage.
Visualizing Semantic Meaning
One might worry that "latent thoughts" are just noise. The authors used t-SNE to compare LatentMAS embeddings with traditional text-generated embeddings.
Figure: Latent thoughts share the same semantic region as text but offer higher density and diversity.
The visualization confirms that latent thoughts stay within the "meaningful" region of the embedding space while actually offering richer representational diversity than constrained tokens.
Critical Insight: The End of Human-Centric Agents?
LatentMAS suggests a future where "System 2" reasoning happens in a dark, high-dimensional space where humans can't read every step. While this raises challenges for interpretability, the authors address this with a "Debug Mode" that can probe latent thoughts into text.
The Takeaway: If you want efficient multi-agent systems, stop making your agents talk to each other in English. Let them communicate in the language they were born with: Vectors.
Conclusion
LatentMAS provides a scalable, training-free paradigm that effectively separates reasoning from communication. By treating multi-agent collaboration as a distributed latent process rather than a dialogue, it clears the path for highly efficient, system-level intelligence.
