arXiv Machine Learning By Yiping Li, Zhiyu An, Wan Du

When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration

Read the original on arXiv Machine Learning →

arXiv:2604. 13349v2 Announce Type: replace Abstract: Communication in Large Language Model (LLM)-based multi-agent systems is moving beyond discrete tokens to preserve richer context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

Faithful, Not Corrective: Model Capability Governs Message-Format Effects in Multi-Hop Agent Relays

The study investigates how different message formats affect the fidelity of information as it passes through multiple LLM agent relays. Using a controlled testbed, the authors encode twelve atomic facts in five formats (free natural language, precision‑instructed NL, JSON, triples, key‑value) across six hops and evaluate recall against programmatic ground truth. Results show that strong relays maintain near‑lossless recall for all formats, while weaker relays exhibit significant format‑dependent recall loss, and that any injected error is faithfully propagated across all formats without causing collateral damage.

By Sicheng Zeng
arXiv Machine Learning
Aug 20

WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing

WhiteMatter introduces a novel architecture for Transformers that connects every attention layer to representations from all layers of each past token, allowing connection weights to vary across consumer layers and adapt to the source token. The design uses a router to mix the $L$ layer states of each token into $k$ KV channels, which are cached for subsequent tokens; each consumer layer attends to one channel. Experiments show that WhiteMatter outperforms a vanilla Transformer with 50% more layers and maintains most of this advantage even when the KV-cache is compressed by 50%.

By Wenbo Zhang, Xiang Ren
arXiv AI
Aug 24

Dual-Cache Latent Space Communication between Heterogeneous Language Models

The paper introduces XKV, a latent protocol that enables efficient communication between heterogeneous language models by translating a sharer's key‑value cache into a receiver's context. XKV overcomes limitations of prior methods by jointly pooling both caches, reconciling differing layer depths, and allowing each receiver position to retrieve its own residual in native KV geometry. Across 45 dataset‑model pairings, XKV outperforms previous protocols and text communication while using fewer parameters and achieving faster translation times.

By Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang