arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
By Olivier Jeunen
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
arXiv:2606. 13670v1 Announce Type: new Abstract: Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered.
By Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma, Frauke Kreuter, Markus Weinmann, Stefan Feuerriegel
The study evaluates whether large language models (LLMs) used as synthetic personas can predict real audience responses to marketing copy. Using thousands of headline A/B tests from the Upworthy Research Archive, the authors compare a ten-persona panel grounded in real audience demographics to a no-persona zero‑shot baseline that asks the model for a typical reader’s click likelihood. Results show that the no‑persona baseline outperforms the persona‑based approach, with higher predictive validity and top‑1 accuracy, indicating that forcing the model to role‑play specific personas introduces bias and noise.
By Alexandre Cristov\~ao Maiorano
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
By Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
arXiv:2609.25066v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead...
By Pei Wang, Lei Wang, Yuanzi Li, Xu Chen
Researchers use synthetic survey respondents generated by large language models as substitutes for human samples, but current validation methods often compare them to human surveys in ways that may not reflect real-world consequential behaviour. The authors propose a new validation framework that requires explicit statements of how well synthetic data correspond to human behaviour, specifies which diagnostics are addressed, and demands subgroup-level validity claims to avoid misrepresentation. The framework operationalises distributional, procedural, and recognition justice dimensions and introduces within-persona counterfactual experiments, illustrated with a case study on electric vehicle charging tariffs and concluded with a reporting checklist for researchers.
By Florian Kutzner, Celina Kacperski, Laura de Moli\`ere, Edoardo Chidichimo, Min Jun Jung, Felix Patrick Sedgwick Wallis, James Kunling He
arXiv:2604. 23904v3 Announce Type: replace-cross Abstract: Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference.
By Yichen Xu
arXiv:2606. 28963v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated.
By Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi, Hong-En Chen, Bo-Ruei Huang
The study audits 576 LLM-based social simulations from 350 papers using the PIMMUR framework, which evaluates agent profile, interaction, memory, minimal control, unawareness, and realism. Results show that PIMMUR principles are met more often than minimal control, unawareness, and realism, with frontier LLMs correctly identifying the underlying experiment in 65.2% of cases and half of prompts pre‑determining outcomes. Reproducing five experiments revealed that many reported collective phenomena disappear or reverse when PIMMUR principles are enforced, suggesting that apparent emergent behaviors may be methodological artifacts rather than genuine social dynamics.
By Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou, Man Ho Lam, Xintao Wang, Hao Zhu, Wenxuan Wang, Maarten Sap
arXiv:2605. 11954v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables that can enter standard empirical designs.
By Jinyuan Wang, Ningyuan Deng, Yi Yang