The paper introduces a simulation framework that uses large language model (LLM) agents conditioned on data‑driven personas to predict A/B test outcomes. These personas are built from anonymized user behavioral patterns, engagement signals, and inferred demographics, offering a more realistic population model than synthetic or rule‑based personas. The authors evaluate question design, persona data source, behavioral depth versus diversity, and population subsampling, achieving 0.75–0.90 directional accuracy on 40 real A/B tests.
By Ziyad Benomar, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
The paper investigates how persona prompting—using short textual descriptions of individuals—to align large language models (LLMs) with human survey responses. It examines the impact of selecting different persona attributes and finds that not all attribute combinations improve performance, suggesting that the variation in human responses to survey questions may explain mixed results. The study evaluates multiple attribute selection methods across four social surveys, two countries, six LLMs, and twenty prediction tasks, offering guidance on when persona prompting is beneficial and which attribute choices are most effective.
By Leon Fr\"ohling, Jens Rupprecht, Markus Strohmaier, Claudia Wagner
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
arXiv:2608. 03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity.
By Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
arXiv:2606. 18263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks.
By Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah, Changyou Chen, Jitendra Ajmera
arXiv:2608.22438v1 Announce Type: new
Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response...
By Taehyeon An, Jaehyeong Park, Donghyuk Shin
arXiv:2512. 07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies.
By Xuan Zhang, Wenxuan Zhang, Anxu Wang, See-Kiong Ng, Yang Deng
arXiv:2607. 10628v1 Announce Type: cross Abstract: We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models.
By Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan
The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.
By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva
The paper introduces ASURRE, a benchmark dataset for detecting AI‑assisted responses in online surveys. It evaluates how different LLM usage strategies—ranging from full generation to persona‑grounded agentic completion—affect the performance of existing machine‑generated text detectors. The study finds that while naive AI usage is easily detected, more sophisticated persona‑grounded agents approach chance performance, yet still leave identifiable behavioural traces that can be aggregated to improve detection.
By Qizhou Wang, Bogdan Mamaev, Christopher Leckie
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
arXiv:2606. 28963v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated.
By Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi, Hong-En Chen, Bo-Ruei Huang