arXiv:2606. 13591v1 Announce Type: new Abstract: Confidence is used for reliability, oversight, and a range of downstream decision tasks in Natural Language Processing (NLP), yet no existing method produces or evaluates a confidence for the output of a multiagent system.
By Ali Elahi, Barbara Di Eugenio
The paper introduces R$^2$-MAD, a framework that enhances multi-agent debate by giving agents an experience memory from past debates. It uses a debate-state-aware retrieval policy to adjust concept priors based on current consensus, and derives confidence weights from retrieved experiences to modulate peer influence. Experiments demonstrate consistent improvements over existing single-agent and MAD baselines.
By Xuanfa Jin, Zhijian Ma, Yongcheng Zeng, Xinyu Cui, Haifeng Zhang, Jun Wang
Meta-Moderator is a learnable framework that treats moderation as a meta‑cognitive process, monitoring debate utility, controlling deliberation, and adjudicating final answers. It is trained independently of the debaters through outcome‑driven policy optimization, allowing dynamic regulation of debate rather than relying on fixed budgets or untrained judges. Across five benchmarks, Meta‑Moderator outperforms common decision layers, transfers across tasks and system configurations, and selectively allocates debate to reduce mis‑aggregation after informative hypotheses appear.
By Wentao Hu, Zhuoyue Wan, Jinhao Shen, Chen Jason Zhang, Xiaoyong Wei, Qing Li
The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines.
whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."
By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.
By Ali Elahi, Michael J. Curry, Barbara Di Eugenio
arXiv:2601. 05746v2 Announce Type: replace Abstract: Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making and complex problem-solving.
By Zhenghao Li, Zhi Zheng, Wei Chen, Jielun Zhao, Yong Chen, Tong Xu, Enhong Chen