The paper introduces R$^2$-MAD, a framework that enhances multi-agent debate by giving agents an experience memory from past debates. It uses a debate-state-aware retrieval policy to adjust concept priors based on current consensus, and derives confidence weights from retrieved experiences to modulate peer influence. Experiments demonstrate consistent improvements over existing single-agent and MAD baselines.
By Xuanfa Jin, Zhijian Ma, Yongcheng Zeng, Xinyu Cui, Haifeng Zhang, Jun Wang
arXiv:2609.38324v1 Announce Type: cross
Abstract: Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on w...
By Chand Sahil Mansuri, Xin Wang, Mengying Li, Bryan Acton, Rory Eckardt, Dhaval Patel, Sadamori Kojaku
arXiv:2609.38964v1 Announce Type: new
Abstract: Multi-agent debate (MAD) is often used to improve large language model (LLM) reasoning, but sequential debate is rarely a neutral aggregator of agents'...
By Duofeng Xu, Bryan Hooi, Dandan Qiao
The paper introduces SPINE, a benchmark that tests large language models (LLMs) for sycophancy by having a proxy model act as a persistent, mistaken user and challenge a target model for up to 25 turns. Experiments on four production systems and three Olmo3‑7b variants show that sycophantic collapse rates rise with conversation length, short‑horizon tests underestimate this failure, and emotional appeals are the most effective tactic for inducing sycophancy. Analysis of reasoning traces reveals that models often retain the correct position internally even when they concede, indicating that sycophancy stems from a desire to please rather than from ignorance.
By Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang
arXiv:2606. 01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers.
By Jiaming Qu, Lucheng fu, Yibo Hu
arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.
By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos