arXiv:2609.36855v1 Announce Type: new
Abstract: Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful inf...
By Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen, Chenxiao Fan, Jiakai Wang, Dong Wang, Yang Liu, Wenjie Wang, Xiangnan He
arXiv:2608. 14375v1 Announce Type: new Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer.
By Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur
arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.
By Yefan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran, Luyang Kong
arXiv:2606. 01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers.
By Jiaming Qu, Lucheng fu, Yibo Hu
The paper investigates how agentic systems decide between acting and abstaining, focusing on the fidelity of their reasoning explanations. Using Qwen3‑8B in a multi‑party conversation setting, the authors compare direct decision policies, reasoning policies, supervised fine‑tuning, and reinforcement learning, finding a trade‑off: strong direct policies yield higher performance but no traceable reasoning, while reasoning policies provide an audit trail at the cost of lower recall. The study also uncovers that exposing reasoning can alter the agent’s policy and that common faithfulness metrics may overstate the alignment between reasoning and decisions.
By Shreya Mendi, Brinnae Bent
The paper investigates intrinsic self‑correction, where a language model revises its own answer without new evidence. Across 29 open‑weight LLMs on BoolQ, GSM8K, and Corr2Cause, the study tracks how revisions change correctness, revealing that while some models improve significantly, others lose a notable fraction of correct answers. The authors compare three runtime strategies—keeping the initial answer, always accepting the revision, and selectively gating revisions—and find that the best approach depends on the model and task, suggesting that self‑correction should be treated as a revision policy rather than a uniformly beneficial second pass.
By Tianzhu Zhang
arXiv:2606. 29026v1 Announce Type: new Abstract: Multi-agent AI systems can improve answer selection by allowing different language models to exchange reasoning traces, revise initial predictions, and support a final decision.
By Shahnewaz Karim Sakib, Anindya Bijoy Das
The paper introduces Revision‑Aware Independent Agent Graphs (RIAG) to address dynamic task routing, where an event stream continually revises task bindings and a system must select the correct document version at query time. By repurposing six benchmarks into over 31,000 dynamic episodes, the authors demonstrate that RIAG balances recomputation and reuse, achieving 54.24 % joint routing‑and‑answer accuracy with only 0.62 calls per query—substantially better than the strongest baseline. The study highlights the trade‑off between stale conclusions and wasted work in dynamic reasoning settings.
By Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang
arXiv:2609.21423v1 Announce Type: new
Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...
By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent.
arXiv:2606. 10747v1 Announce Type: new Abstract: As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisions together: discussing, negotiating, and acting on shared tasks.
By Filippo Tonini, Federico Torrielli, Anton Danholt Lautrup, Peter Schneider-Kamp, Mustafa Mert \c{C}elikok, Lukas Galke Poech
arXiv:2608. 12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging.
By Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang