arXiv AI

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

arXiv:2608. 11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns.

arXiv AI
Aug 24

Dual-Cache Latent Space Communication between Heterogeneous Language Models

The paper introduces XKV, a latent protocol that enables efficient communication between heterogeneous language models by translating a sharer's key‑value cache into a receiver's context. XKV overcomes limitations of prior methods by jointly pooling both caches, reconciling differing layer depths, and allowing each receiver position to retrieve its own residual in native KV geometry. Across 45 dataset‑model pairings, XKV outperforms previous protocols and text communication while using fewer parameters and achieving faster translation times.

By Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang
arXiv AI
Jun 2

Latent Collaboration in Multi-Agent Systems

arXiv:2511. 20639v3 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) extend large language models (LLMs) from independent single-model reasoning to coordinative system-level intelligence.

By Jiaru Zou, Ruizhong Qiu, Gaotang Li, Xiyuan Yang, Katherine Tieu, Pan Lu, Ke Shen, Hanghang Tong, Yejin Choi, Jingrui He, James Zou, Mengdi Wang, Ling Yang
arXiv AI
6d ago

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.

By Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)