arXiv AI

When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

arXiv AI
1d ago

Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

The paper introduces Independent–Communicate–Revise (ICR), a framework that isolates communication effects in large language model multi‑agent systems by fixing initial reasoning and measuring how messages influence answer revision. ICR evaluates correction, preservation, and selectivity across four reasoning benchmarks, revealing that similar overall accuracy can mask divergent revision behaviors. The study shows that richer messages can both improve and harm outcomes, and that receiver policies can shift preservation and correction dynamics differently across tasks.

By Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan
Hugging Face Trending Papers
Jul 29

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent.

arXiv AI
Sep 11

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

The paper introduces NetArtifactBench, a benchmark designed to evaluate whether AI agents can detect and repair inconsistencies in network experiment records while preserving supported claims. It tests 23 agent configurations on 52 instances with injected inconsistencies, finding an average pass rate of 65.3 % but no runtime exceeding 30 % for complex repairs that require recovering implicit relations and propagating changes across artifacts. The results highlight a clear distinction between local corrections and full record-level repair, leading the authors to argue that artifact integrity should be a primary design and evaluation criterion for AI agents in network systems.

By Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu
arXiv AI
Sep 2

EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems

EDGE is a framework that attributes multiple related errors in multi-agent large language model systems by constructing an error dependency graph from observed error events. It validates a reliable causal subset through counterfactual rollout and uses this inference graph to guide a two-stage LLM-as-judge detector for error attribution. Experiments on TRAIL and MAST demonstrate that EDGE improves category-level multi-error attribution across most models and settings, and that the graph aids explanation and repair analysis.

By Jun Hou, Priya Pitre, Yi Fang, Xuan Wang