arXiv:2607. 18269v1 Announce Type: new Abstract: The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters.
By Igor Douven
The paper introduces a behavior‑aware framework to build diverse crowds of large language models (LLMs) for future prediction. By analyzing reasoning traces on independent tasks, clustering models by behavioral similarity, and selecting representative medoids, the authors demonstrate that a small, well‑chosen crowd can outperform a larger, conventional voting ensemble. Experiments with 25 LLMs across multiple benchmarks show significant reductions in model calls and inference cost while improving prediction accuracy.
By Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen
The paper examines how Preference Inference (PI) models used in large-scale participatory democracy platforms can alter the perceived consensus and minority support by predicting missing votes. It introduces a collective‑centric evaluation framework that assesses whether inferred votes maintain key properties of the overall preference landscape, rather than focusing solely on individual prediction accuracy. Using the largest multilingual dataset to date—four consultations with over 90,000 participants, 1 million votes, and 22 languages—the study finds that models with similar predictive accuracy can differ markedly in how well they preserve the collective structure, underscoring that accuracy alone is insufficient for evaluating PI in democratic contexts.
By Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner, Nazanin Shafiabadi, Laur\`ene Cave, David Mas, Jean-Philippe Cointet, Benjamin Piwowarski, Fran\c{c}ois Yvon
arXiv:2602. 13792v2 Announce Type: replace Abstract: Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these systems remain isolated and cannot readily share their capabilities.
By Siyang Li, Chenhao Liu, Dongrui Wu, Zhigang Zeng, Lieyun Ding
arXiv:2608.30311v1 Announce Type: cross
Abstract: Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-makin...
By Zhuoran Lu, Weilong Wang, Yangyang Yu, Xinru Wang, Zhuoyan Li, Zhiwei Liu, Sophia Ananiadou
arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.
By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv:2609.39483v1 Announce Type: new
Abstract: Ownership establishes rights over the use, control, and transfer of objects. Understanding these relations is essential for AI systems to interact appr...
By Xizhi Xiao, Yue Wu, Shan Xu, Jia Liu
The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.
By Ionel Eduard Stan, Paolo Napoletano
arXiv:2607. 03695v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in interacting populations, raising the question of what such populations come to believe collectively.
By Kaixuan Liu, Guojun Xiong, Weinan Zhang, Shengpu Tang
arXiv:2608. 14629v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI).
By Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu
arXiv:2609.39645v1 Announce Type: cross
Abstract: Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion coupl...
By Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang
The paper investigates how a minority of biased agents in a multi‑agent system of large language models (LLMs) can amplify bias through textual interactions. Even a small percentage of persistently extreme agents causes significant opinion shifts among the non‑biased agents, with the effect occurring faster in the Llama 3.2 model than in a classical Friedkin‑Johnsen model. Semantic analysis shows that rhetorical consistency rises with biased exposure and that non‑biased agents adopt the biased vocabulary even when their numerical opinions change only modestly.
By Omran Berjawi, Giuseppe Fenza, Rida Khatoun