arXiv:2506. 12078v2 Announce Type: replace-cross Abstract: Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations.
By Haoxiang Guan, Jiyan He, Liyang Fan, Zhenzhen Ren, Shaobin He, Xin Yu, Yuan Chen, Xueyin Xu, Shuxin Zheng, Yan Gao, Enhong Chen, Tie-Yan Liu, Zhen Liu
The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.
By Farah Atif, Sougata Saha, Monojit Choudhury
arXiv:2509. 17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosing what an actor says or does, deciding what happens after an action, or both.
By Glenn Matlin, Isaac Song, Yixiong Hao, Parv Mahajan, Evan Montoya, Ryan Bard, Stuart R. Topp, Anthony Wen-Ming Zang, Mohammed Rehan Parwani, Soham Shetty, Mark Riedl
ExpertIVS is a framework that uses 14 sociological expert agents to interpret World Values Survey responses, reconstructing individual value systems in a coherent, internally consistent manner rather than simply concatenating survey answers. It introduces a multi‑agent debate mechanism to assess LLM alignment with these value profiles during dynamic interactions. Experiments on 480 individuals from 12 countries show a 90.78% value restoration fidelity and a 5.3% improvement in value generalization over baseline methods, while also demonstrating strong personality discriminability and behavioral consistency.
By Zhen Wang, Yuqi Ren, Yuehan Cui, Hongxiang Wang, Jianxiang Peng, Zhaoxia Zhang, Bingkun Zhu, Tongxuan Zhang, Dezhi Tong, Deyi Xiong
arXiv:2608. 03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity.
By Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
arXiv:2607. 26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions.
By Zihan Chen, Di Zhu, Lei Nico Zheng
arXiv:2606. 14715v1 Announce Type: cross Abstract: LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors.
By Yaoning Yu, Ye Yu, Haojing Luo, Haohan Wang
The paper investigates whether large language model (LLM) agents can simulate public deliberation by reflecting population opinion patterns and producing interaction-driven opinion change. Using census‑grounded Korean personas debating real policy questions, the study finds that persona agents fail to reliably reproduce population opinion patterns, often concentrating responses and reversing demographic differences. While deliberations generate reasoned, reciprocal arguments and some stance movement, much of this change occurs without peer exchange, and anchoring agents to population‑informed starting positions suppresses updating, indicating that population representation, argument generation, and interaction‑driven opinion change do not necessarily align.
By Chaemin Jang, Junsik Min, Jaewoo Choi, Donggyu Lee, Haiin Lee, Junyoung Park, Namhee Kim, Hyunwoo Kim, Jungwon Kim, Juho Kim, Nuri Kim, Jihee Kim
CityReal is a modular framework that uses large language model agents to simulate human-aligned urban behavior. It models agents as intention-driven decision makers who pursue coherent mobility and activity plans, learning habits and preferences over time. By training textual adapters to align agent decisions with observed population statistics, CityReal improves realism at both micro and macro levels and can scale to tens of thousands of agents for analyzing crowd density, place popularity, mobility flows, and well‑being under various urban scenarios.
By Nicolas Bougie, Xiaotong Ye, Narimasa Watanabe
WorldBench is a new multilingual benchmark that tests large language model agents on culturally grounded everyday workflows, offering 1,600 tasks in seven languages and eight cultures. The benchmark evaluates agents through structured sandbox actions and introduces Constrained Task Success (CTS), a metric that assesses task completion, minimal modification, and other complementary aspects via deterministic and LLM-as-a-Judge evaluations. Experiments show that even leading models achieve only 49.2% CTS, revealing significant gaps in correctness and state preservation across languages and cultures.
By Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
The study evaluates how well language‑model agents can simulate individual social media reactions by comparing predictions under different prompt conditions. Eight Serbian participants’ reactions to 68 posts were recorded, and four language models were asked to predict these reactions using prompts that varied in profile content and instruction style. The results show that prompts emphasizing attitudinal content and intuitive, immediate responses yield the highest fidelity, outperforming demographic backstories and a crowd baseline, and suggesting that such agents could act as general‑purpose simulated users.
By Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic, Jue Wang
arXiv:2608.10503v2 Announce Type: replace
Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. T...
By Davood Wadi, Mohsen Ghodrat, Matthew Philp