arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
arXiv:2608. 07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing.
By Ljubisa Bojic, Ljiljana Matic, Joerg Matthes, Milan Cabarkapa, Bojana Dinic, Jue Wang
arXiv:2607. 18310v1 Announce Type: cross Abstract: Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent.
By Gurkan Ozkan
The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.
By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
arXiv:2609.13148v1 Announce Type: cross
Abstract: Large language models are increasingly deployed as synthetic consumer panels, promising $97\%$ cost reductions over traditional surveys. Yet aggregat...
By Robson Tigre, Hugo Gobato Souto
arXiv:2411. 10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes.
By Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
The paper introduces a simulation framework that uses large language model (LLM) agents conditioned on data‑driven personas to predict A/B test outcomes. These personas are built from anonymized user behavioral patterns, engagement signals, and inferred demographics, offering a more realistic population model than synthetic or rule‑based personas. The authors evaluate question design, persona data source, behavioral depth versus diversity, and population subsampling, achieving 0.75–0.90 directional accuracy on 40 real A/B tests.
By Ziyad Benomar, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
arXiv:2605.29791v2 Announce Type: replace
Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions...
By Yutong Yang, Chenxi Miao, Weikang Li, Yunfang Wu
The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.
The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.
The study evaluates a Korean synthetic persona panel (NVIDIA Nemotron‑Personas‑Korea) conditioned on Gemini 3.5 Flash and EXAONE against the KISDI Korea Media Panel Survey. Across eight digital‑AI service‑use indicators and eight innovativeness constructs, the panels achieved mean absolute errors of 15–19 pp, with segment‑level errors up to 52.4 pp and correlation coefficients between 0.69 and 0.90. Holdout calibration reduced sex‑by‑age cell errors but still lagged behind direct real‑data estimation, indicating that synthetic panels are useful diagnostically but not as survey substitutes.
By Howard Kim, Keun Tae Cho
arXiv:2608.18768v2 Announce Type: replace
Abstract: Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differen...
By Fathin Difa Robbani