arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.
By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv:2510. 20963v2 Announce Type: replace Abstract: Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasoning and provide effective supervision to superhuman LLMs.
By Yongqiang Chen, Gang Niu, James Cheng, Bo Han, Masashi Sugiyama
arXiv:2604. 02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion.
By Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah, Krishna Agaram, Ritwik Garg, Sagar Jha, Nimet Beyza Bozdag, Dilek Hakkani-Tur
arXiv:2606. 13591v1 Announce Type: new Abstract: Confidence is used for reliability, oversight, and a range of downstream decision tasks in Natural Language Processing (NLP), yet no existing method produces or evaluates a confidence for the output of a multiagent system.
By Ali Elahi, Barbara Di Eugenio
arXiv:2607. 26212v1 Announce Type: cross Abstract: Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems.
By Quim Motger, Marc Oriol, Jordi Marco, Xavier Franch
arXiv:2608. 01463v2 Announce Type: replace Abstract: Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims.
By Weijun Gao, Xiang Ding, Haoyang Liu, Tiancheng Xing
Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signals, token-level log-probabilities from decoding, can predict reasoning quality as assessed by LLM-as-judge evaluation.
Meta-Moderator is a learnable framework that treats moderation as a meta‑cognitive process, monitoring debate utility, controlling deliberation, and adjudicating final answers. It is trained independently of the debaters through outcome‑driven policy optimization, allowing dynamic regulation of debate rather than relying on fixed budgets or untrained judges. Across five benchmarks, Meta‑Moderator outperforms common decision layers, transfers across tasks and system configurations, and selectively allocates debate to reduce mis‑aggregation after informative hypotheses appear.
By Wentao Hu, Zhuoyue Wan, Jinhao Shen, Chen Jason Zhang, Xiaoyong Wei, Qing Li
arXiv:2601. 05746v2 Announce Type: replace Abstract: Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making and complex problem-solving.
By Zhenghao Li, Zhi Zheng, Wei Chen, Jielun Zhao, Yong Chen, Tong Xu, Enhong Chen
arXiv:2606. 13197v1 Announce Type: new Abstract: Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents.
By Fuqiang Niu, Bowen Zhang
arXiv:2606. 29425v1 Announce Type: new Abstract: Existing multi-agent debate frameworks suffer from two critical limitations: they rely on static architectures where agent roles and coordination patterns are fixed at design time, and they require instantiating multiple model copies, incurring substantial computational overhead.
By Dayong Liang, Kaisong Gong, Yi Cai, Changmeng Zheng, Xiao-Yong Wei
arXiv:2608. 19701v1 Announce Type: new Abstract: Long-term multi-agent systems continuously accumulate the memories produced by different agents.
By Chenchen Lin, Wenhao Yuan, Xuehe Wang, Edith Cheuk Han Ngai