The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines.
whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."
By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
arXiv:2602. 12089v3 Announce Type: replace-cross Abstract: As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove both individual and group outcomes.
By Kehang Zhu, Nithum Thain, Vivian Tsai, James Wexler, Crystal Qian
arXiv:2605.25200v3 Announce Type: replace
Abstract: Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single us...
By Xiang Cheng, Yulan Hu, Lulu Zheng, Xiangwen Zhang, Zheng Pan, Xin Li, Yong Liu
arXiv:2602. 16794v2 Announce Type: replace-cross Abstract: Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored.
By Pengqi Liu, Zijun Yu, Mouloud Belbahri, Arthur Charpentier, Masoud Asgharian, Jesse C. Cresswell
Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response.
arXiv:2606. 00007v1 Announce Type: new Abstract: As AI agents transition from isolated tools to collaborative participants in shared knowledge ecosystems, governing collective knowledge curation becomes a critical challenge.
By Steven Johnson
arXiv:2608. 10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems.
By Maurice Flechtner
arXiv:2609.38324v1 Announce Type: cross
Abstract: Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on w...
By Chand Sahil Mansuri, Xin Wang, Mengying Li, Bryan Acton, Rory Eckardt, Dhaval Patel, Sadamori Kojaku
arXiv:2609.22497v1 Announce Type: new
Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually a...
By Federico Barrera-Lemarchand, Mariano Sigman, Joaquin Navajas
arXiv:2608. 16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task.
By Bowen Li, Guojun Wang
arXiv:2608. 13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve.
By Sanjeev Manivannan
arXiv:2604. 11840v3 Announce Type: replace-cross Abstract: Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises.
By Sandro Andric