arXiv Machine Learning

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

The paper introduces a new measure of generative‑process diversity for language models, using Normalised Compression Distance on raw outputs after controlling for permutation effects. Across 38 models, this metric uncovers population structure that semantic similarity misses and predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability. The authors argue that higher generative‑process diversity reduces correlated failures in multi‑model systems, offering a practical tool for safety‑relevant applications.

Hugging Face Trending Papers
Sep 3

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

The paper argues that traditional semantic similarity fails to capture the true diversity of language models. It introduces a new metric—generative‑process diversity—measured via Normalised Compression Distance on raw outputs, which reveals hidden population structure among 38 models. This metric predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability.

arXiv AI
Sep 3

Do Large Language Models Capture the Diversity in their Training Data?

The paper investigates whether large language models (LLMs) capture the full diversity of outputs present in their training data. Using an information‑theoretic approach, the authors compare the conditional entropy of model‑generated outputs with that of the training data, finding that LLMs consistently produce outputs with lower conditional entropy across various models, scales, and decoding strategies. They also extend the analysis to image and text‑conditioned generators, propose a post‑hoc correction method based on matrix‑entropy projection to increase conditional diversity, and provide theoretical guarantees and an efficient algorithm for this correction.

By Youqi Wu, Farzan Farnia
arXiv AI
Sep 10

Limits of Reliability and Scaling in Language Models

The paper argues that large language models cannot achieve perfect reliability for any task, even with unlimited scale. It establishes that each generative task has an inherent reliability ceiling set by how much output uncertainty can be resolved from observable context, with a resolvable part that can be improved by more context and a subjective part tied to task ambiguity. The authors derive a scaling law showing that performance is limited by the scarcer resource—either training data or model capacity—and explain how this law explains phenomena such as retrieval augmentation and catastrophic forgetting.

By Subhabrata Majumdar
arXiv AI
Aug 19

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

The paper argues that measuring diversity in AI-generated content using a single scalar score is inherently ambiguous and often misleading. It reviews existing diversity metrics, demonstrates their limitations through axiomatic and empirical analyses, and introduces diversity profiles—curve-valued, condition-aware summaries that evaluate diversity across a range of thresholds, scales, exponents, or orders. These profiles reveal whether comparisons are robust across resolutions or depend on arbitrary parameter choices, offering a more transparent framework for generative AI evaluation.

By Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, Jos\'e Miguel Hern\'andez-Lobato, Hao Zhang, Xue Liu