arXiv:2605. 22863v2 Announce Type: replace Abstract: LLM agents today communicate via text, which incurs considerable latency and information loss due to the need to autoregressively decode the sharer model's state and encode at the receiver model.
By Maximillian Rossi, Prajwal Raghunath, Eugene Wu
arXiv:2609.37017v1 Announce Type: new
Abstract: LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natu...
By Shinan Zhang, Tao Zhang, Qihui Zhu, Mengjie Zhang, Dong Jin, Yunpeng Hou, Shuangwu Chen, Xiaobin Tan, Quan Zheng, Jian Yang
The study investigates how different message formats affect the fidelity of information as it passes through multiple LLM agent relays. Using a controlled testbed, the authors encode twelve atomic facts in five formats (free natural language, precision‑instructed NL, JSON, triples, key‑value) across six hops and evaluate recall against programmatic ground truth. Results show that strong relays maintain near‑lossless recall for all formats, while weaker relays exhibit significant format‑dependent recall loss, and that any injected error is faithfully propagated across all formats without causing collateral damage.
By Sicheng Zeng
WhiteMatter introduces a novel architecture for Transformers that connects every attention layer to representations from all layers of each past token, allowing connection weights to vary across consumer layers and adapt to the source token. The design uses a router to mix the $L$ layer states of each token into $k$ KV channels, which are cached for subsequent tokens; each consumer layer attends to one channel. Experiments show that WhiteMatter outperforms a vanilla Transformer with 50% more layers and maintains most of this advantage even when the KV-cache is compressed by 50%.
By Wenbo Zhang, Xiang Ren
arXiv:2609.32259v2 Announce Type: replace
Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires ea...
By Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy
The paper introduces XKV, a latent protocol that enables efficient communication between heterogeneous language models by translating a sharer's key‑value cache into a receiver's context. XKV overcomes limitations of prior methods by jointly pooling both caches, reconciling differing layer depths, and allowing each receiver position to retrieve its own residual in native KV geometry. Across 45 dataset‑model pairings, XKV outperforms previous protocols and text communication while using fewer parameters and achieving faster translation times.
By Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang