arXiv Machine Learning

Epistemic diversity across language models mitigates knowledge collapse

arXiv:2512. 15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems.

arXiv Machine Learning
Aug 11

On the Effect of Sampling Diversity in Scaling LLM Inference

arXiv:2502. 11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it.

By Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng
arXiv Machine Learning
Sep 4

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

The paper introduces a new measure of generative‑process diversity for language models, using Normalised Compression Distance on raw outputs after controlling for permutation effects. Across 38 models, this metric uncovers population structure that semantic similarity misses and predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability. The authors argue that higher generative‑process diversity reduces correlated failures in multi‑model systems, offering a practical tool for safety‑relevant applications.

By Ross Tieman, Evan Markou
arXiv AI
Sep 15

The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems

The paper introduces the Universe of Universes (UoU) framework, treating the ecosystem of major large language models as a structured retrieval corpus and proposing a compositional Automated Reasoning and Machine Learning architecture for cross-model retrieval‑augmented generation. It formally defines the Benefit Yield Function (BYF), measuring marginal performance gain per added model, and identifies an implosion threshold θ* where BYF becomes zero and ensemble performance degrades. The work highlights gaps in current LLM ensemble research, such as lack of performance analysis across full model universes, and connects these findings to implications for DoD AI acquisition policy and testing of AI‑enabled systems.

By Danielle Franklin, Vasu Raj Jain
Hugging Face Trending Papers
Sep 3

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

The paper argues that traditional semantic similarity fails to capture the true diversity of language models. It introduces a new metric—generative‑process diversity—measured via Normalised Compression Distance on raw outputs, which reveals hidden population structure among 38 models. This metric predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability.