The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.
By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva
arXiv:2609.35860v1 Announce Type: cross
Abstract: Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors a...
By Pranav Darshan, Pranav A, Sravan Karthick T, Minal Moharir, Ivan P. Yamshchikov
The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.
By Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty...
arXiv:2608.29266v1 Announce Type: cross
Abstract: Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments lik...
By An Duy Nguyen, Muhammad Aurangzeb Ahmad
arXiv:2510. 12857v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide.
By Robin Staab, Jasper Dekoninck, Maximilian Baader, Martin Vechev