arXiv:2609.13148v1 Announce Type: cross
Abstract: Large language models are increasingly deployed as synthetic consumer panels, promising $97\%$ cost reductions over traditional surveys. Yet aggregat...
By Robson Tigre, Hugo Gobato Souto
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.
By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva
Researchers use synthetic survey respondents generated by large language models as substitutes for human samples, but current validation methods often compare them to human surveys in ways that may not reflect real-world consequential behaviour. The authors propose a new validation framework that requires explicit statements of how well synthetic data correspond to human behaviour, specifies which diagnostics are addressed, and demands subgroup-level validity claims to avoid misrepresentation. The framework operationalises distributional, procedural, and recognition justice dimensions and introduces within-persona counterfactual experiments, illustrated with a case study on electric vehicle charging tariffs and concluded with a reporting checklist for researchers.
By Florian Kutzner, Celina Kacperski, Laura de Moli\`ere, Edoardo Chidichimo, Min Jun Jung, Felix Patrick Sedgwick Wallis, James Kunling He
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
By Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
The article introduces the Artificial Societies Benchmark, a validation framework designed to evaluate synthetic populations used in research. It comprises eleven tests covering internal, construct, and external validity, drawing on twenty human data sources and comparing nine language models. The benchmark links specific research uses to the evidence required and assesses how results vary with different respondent information, revealing that strong performance in one domain does not guarantee fidelity in others.
By Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He