arXiv:2511. 04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences.
By Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal, Max Pellert
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
arXiv:2604. 02458v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible.
By Zonghan Li, Feng Ji
arXiv:2608.18768v2 Announce Type: replace
Abstract: Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differen...
By Fathin Difa Robbani
The article introduces the Artificial Societies Benchmark, a validation framework designed to evaluate synthetic populations used in research. It comprises eleven tests covering internal, construct, and external validity, drawing on twenty human data sources and comparing nine language models. The benchmark links specific research uses to the evidence required and assesses how results vary with different respondent information, revealing that strong performance in one domain does not guarantee fidelity in others.
By Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He
arXiv:2609.16436v1 Announce Type: cross
Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the...
By Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj
The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.
By Raad Bin Tareaf
The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.
arXiv:2606. 17657v1 Announce Type: new Abstract: People make decisions differently in strategic interactions.
By Zirui Cheng, Zeyu Shen, Thomas L. Griffiths, Peter Henderson
arXiv:2608. 03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity.
By Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
arXiv:2607. 18310v1 Announce Type: cross Abstract: Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent.
By Gurkan Ozkan
arXiv:2606. 22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure.
By Yujun Zhou, Christopher M. Ackerman