The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.
By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva
arXiv:2609.22607v1 Announce Type: new
Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas,...
By Minwoo Kang, T\'ea Wright, Seun Eisape, Ayush Raj, Suhong Moon, Joseph Suh, Alane Suhr, David M. Chan, John Canny
arXiv:2607. 25292v1 Announce Type: new Abstract: Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution.
By Chaemin Jang, Dongman Lee, Jihee Kim
arXiv:2607. 10628v1 Announce Type: cross Abstract: We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models.
By Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan
Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual.
arXiv:2608.22438v1 Announce Type: new
Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response...
By Taehyeon An, Jaehyeong Park, Donghyuk Shin
The paper introduces a bias depth score to differentiate between stable model preferences (Deep biases) and prompt‑dependent responses (Shallow biases) in large language models. By analyzing 4,442 opinion prompts across four models, it finds that only about a quarter of concentrated preferences persist after scenario reframing, indicating that most are shallow. The study shows Deep biases are more often inherited from pretraining and harder to remove through fine‑tuning or prompt‑based debiasing, highlighting the need to distinguish learned biases from prompt artifacts.
By An Vo, Vy Tuong Dang, Khai-Nguyen Nguyen, Emilio Villa-Cueva, Thamar Solorio, Anh Totti Nguyen, Daeyoung Kim
arXiv:2608. 06115v1 Announce Type: new Abstract: Predicting how a population will answer a new question is a long-standing goal.
By Pranav Dahiya
The paper introduces Demographic Pluralism, an inference-time framework that estimates population-level opinion distributions without requiring training data or task-specific fine-tuning. It generates multiple perspectives within demographically grounded groups and, across four model backbones on GlobalOpinionQA and VITAL, reduces Jensen-Shannon distance by 8.4%–26.4% compared to Modular Pluralism. The study finds that equal weighting of group perspectives yields the best overall performance, while weighted aggregation performs worse due to increased group-level error.
By Meng-Chen Wu, Qipin Chen, Ansh Jain, Tess Wood, Zhe Du, Si-Chi Chin
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
The paper investigates how persona prompting—using short textual descriptions of individuals—to align large language models (LLMs) with human survey responses. It examines the impact of selecting different persona attributes and finds that not all attribute combinations improve performance, suggesting that the variation in human responses to survey questions may explain mixed results. The study evaluates multiple attribute selection methods across four social surveys, two countries, six LLMs, and twenty prediction tasks, offering guidance on when persona prompting is beneficial and which attribute choices are most effective.
By Leon Fr\"ohling, Jens Rupprecht, Markus Strohmaier, Claudia Wagner
The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.
By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du