arXiv AI By Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

Read the original on arXiv AI →

arXiv:2608. 14375v1 Announce Type: new Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

The paper introduces Independent–Communicate–Revise (ICR), a framework that isolates communication effects in large language model multi‑agent systems by fixing initial reasoning and measuring how messages influence answer revision. ICR evaluates correction, preservation, and selectivity across four reasoning benchmarks, revealing that similar overall accuracy can mask divergent revision behaviors. The study shows that richer messages can both improve and harm outcomes, and that receiver policies can shift preservation and correction dynamics differently across tasks.

By Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata