arXiv AI

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

arXiv:2607. 18269v1 Announce Type: new Abstract: The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters.

Hugging Face Trending Papers
Jun 29

Diversity is the Strength of the AI Crowd

Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy?

arXiv AI
Jun 30

Diversity is the Strength of the AI Crowd

arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.

By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day
arXiv AI
Aug 26

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

The paper introduces a behavior‑aware framework to build diverse crowds of large language models (LLMs) for future prediction. By analyzing reasoning traces on independent tasks, clustering models by behavioral similarity, and selecting representative medoids, the authors demonstrate that a small, well‑chosen crowd can outperform a larger, conventional voting ensemble. Experiments with 25 LLMs across multiple benchmarks show significant reductions in model calls and inference cost while improving prediction accuracy.

By Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen
arXiv Computation and Language
Aug 21

The Asymmetric Harms of LLM Compression

arXiv:2608. 19670v1 Announce Type: new Abstract: Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts.

By Yuan Wu, Mairui Li, Lesia Semenova, Chudi Zhong
arXiv AI
Sep 15

The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems

The paper introduces the Universe of Universes (UoU) framework, treating the ecosystem of major large language models as a structured retrieval corpus and proposing a compositional Automated Reasoning and Machine Learning architecture for cross-model retrieval‑augmented generation. It formally defines the Benefit Yield Function (BYF), measuring marginal performance gain per added model, and identifies an implosion threshold θ* where BYF becomes zero and ensemble performance degrades. The work highlights gaps in current LLM ensemble research, such as lack of performance analysis across full model universes, and connects these findings to implications for DoD AI acquisition policy and testing of AI‑enabled systems.

By Danielle Franklin, Vasu Raj Jain
Hugging Face Trending Papers
Aug 20

The Asymmetric Harms of LLM Compression

Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts. In this work, we systematically evaluate 3 LLMs across 11 compression methods to investigate the effects of compression on knowledge retention, model confidence, and social bias.