arXiv AI

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

The paper introduces a behavior‑aware framework to build diverse crowds of large language models (LLMs) for future prediction. By analyzing reasoning traces on independent tasks, clustering models by behavioral similarity, and selecting representative medoids, the authors demonstrate that a small, well‑chosen crowd can outperform a larger, conventional voting ensemble. Experiments with 25 LLMs across multiple benchmarks show significant reductions in model calls and inference cost while improving prediction accuracy.

arXiv AI
Jun 17

Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models

arXiv:2504. 03991v2 Announce Type: replace-cross Abstract: Understanding how humans collaborate and communicate in teams is essential for improving human-agent teaming and AI-assisted decision-making.

By Siddharth Srikanth, Varun Bhatt, Boshen Zhang, Werner Hager, Charles Michael Lewis, Katia P. Sycara, Aaquib Tabrez, Stefanos Nikolaidis
arXiv AI
4d ago

Population Fidelity: Evaluating Population Representativeness in LLMs

The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.

By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva
arXiv AI
Jun 30

Diversity is the Strength of the AI Crowd

arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.

By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day
Hugging Face Trending Papers
Jun 29

Diversity is the Strength of the AI Crowd

Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy?

arXiv AI
Jun 9

Scaling Participation in Modular AI Systems

arXiv:2606. 07812v1 Announce Type: new Abstract: Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness.

By Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer, Yejin Choi, Yulia Tsvetkov
arXiv Machine Learning
Aug 11

On the Effect of Sampling Diversity in Scaling LLM Inference

arXiv:2502. 11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it.

By Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng