The Wisdom of Artificial Deliberative Crowds
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 18269v1 Announce Type: new Abstract: The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters.
The paper introduces a behavior‑aware framework to build diverse crowds of large language models (LLMs) for future prediction. By analyzing reasoning traces on independent tasks, clustering models by behavioral similarity, and selecting representative medoids, the authors demonstrate that a small, well‑chosen crowd can outperform a larger, conventional voting ensemble. Experiments with 25 LLMs across multiple benchmarks show significant reductions in model calls and inference cost while improving prediction accuracy.
The paper examines how Preference Inference (PI) models used in large-scale participatory democracy platforms can alter the perceived consensus and minority support by predicting missing votes. It introduces a collective‑centric evaluation framework that assesses whether inferred votes maintain key properties of the overall preference landscape, rather than focusing solely on individual prediction accuracy. Using the largest multilingual dataset to date—four consultations with over 90,000 participants, 1 million votes, and 22 languages—the study finds that models with similar predictive accuracy can differ markedly in how well they preserve the collective structure, underscoring that accuracy alone is insufficient for evaluating PI in democratic contexts.
arXiv:2602. 13792v2 Announce Type: replace Abstract: Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these systems remain isolated and cannot readily share their capabilities.
arXiv:2608.30311v1 Announce Type: cross Abstract: Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-makin...
arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.