arXiv AI

Geographic Bias and Diversity in AI Evaluation

arXiv:2606. 05187v1 Announce Type: cross Abstract: Among the many challenges hindering the responsible development and deployment of AI, arguably none has faced more intense scrutiny than bias in its various forms.

arXiv AI
Aug 19

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

The paper argues that measuring diversity in AI-generated content using a single scalar score is inherently ambiguous and often misleading. It reviews existing diversity metrics, demonstrates their limitations through axiomatic and empirical analyses, and introduces diversity profiles—curve-valued, condition-aware summaries that evaluate diversity across a range of thresholds, scales, exponents, or orders. These profiles reveal whether comparisons are robust across resolutions or depend on arbitrary parameter choices, offering a more transparent framework for generative AI evaluation.

By Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, Jos\'e Miguel Hern\'andez-Lobato, Hao Zhang, Xue Liu
arXiv AI
Sep 23

Biased AI improves human performance but reduces perceived helpfulness

The study tests deliberately biased AI assistants and finds that such bias improves human performance on tasks like misinformation evaluation, financial investment, and graduate education compared to neutral AI. However, participants undervalue biased AI and overvalue neutral AI, even when performance is similar. When two AI biases flank a participant’s perspective, performance gains are maintained while reducing the perceived cost and one‑sided influence.

By Shiyang Lai, Jiwoong Choi, Junsol Kim, Nadav Kunievsky, Yujin Potter, James Evans
arXiv Machine Learning
Sep 4

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

The paper introduces a new measure of generative‑process diversity for language models, using Normalised Compression Distance on raw outputs after controlling for permutation effects. Across 38 models, this metric uncovers population structure that semantic similarity misses and predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability. The authors argue that higher generative‑process diversity reduces correlated failures in multi‑model systems, offering a practical tool for safety‑relevant applications.

By Ross Tieman, Evan Markou