arXiv:2607. 26773v1 Announce Type: new Abstract: Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information.
By Huixiang Zhang, Mahzabeen Emu
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent.
arXiv:2604. 13349v2 Announce Type: replace Abstract: Communication in Large Language Model (LLM)-based multi-agent systems is moving beyond discrete tokens to preserve richer context.
By Yiping Li, Zhiyu An, Wan Du
arXiv:2607. 09678v1 Announce Type: new Abstract: When LLM agents hand off information to one another, does the message format matter?
By Zayx Shawn
The study investigates how different message formats affect the fidelity of information as it passes through multiple LLM agent relays. Using a controlled testbed, the authors encode twelve atomic facts in five formats (free natural language, precision‑instructed NL, JSON, triples, key‑value) across six hops and evaluate recall against programmatic ground truth. Results show that strong relays maintain near‑lossless recall for all formats, while weaker relays exhibit significant format‑dependent recall loss, and that any injected error is faithfully propagated across all formats without causing collateral damage.
By Sicheng Zeng
arXiv:2609.24635v1 Announce Type: new
Abstract: When a language model reads an operation such as "Swap the contents of Box F and Box B", its forward pass writes keys and values for those tokens into...
By Lingfeng Wu, Behzad Shomali
arXiv:2609.06872v1 Announce Type: new
Abstract: When a user asks an assistant to forget a record, the test is whether the memory now matches the state it would hold if the record had never been store...
By Vishwajith Ramesh
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
By Alex Kwon
SCIT (Suffix Cache Interchange Test) is a causal protocol designed to identify which transformer components carry counterfactual computations in latent chain-of-thought models. By constructing exact source‑recipient counterfactuals and applying sufficiency tests, K/V splits, hidden‑state controls, and semantic source controls, SCIT demonstrates that counterfactual arithmetic primarily transfers through value‑cache suffix trajectories rather than hidden states or keys. The method reveals carrier‑regime shifts across different GPT‑2 checkpoints, providing a cache‑level diagnostic and a competence‑gated carrier map for arithmetic mechanisms.
By Yi Ding, Lijun Huang, Menglin Yang
arXiv:2608. 16190v1 Announce Type: cross Abstract: Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost.
By Anik Jha
arXiv:2608. 20280v1 Announce Type: cross Abstract: Semantic caches reuse an LLM response when the incoming query embedding lies near a cached query, but proposed eviction policies have rarely been compared under one protocol.
By Yash Kulkarni, Shubham Harkare, Arvind Suresh Yogesh Babu
arXiv:2609.24146v1 Announce Type: new
Abstract: Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one anot...
By Cong Li, Cheng Chen, Thomas Fung, Alex Rossi, Yi Li