arXiv AI

Multilingual Agent-Based World Modeling for Social Science

arXiv:2512. 07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies.

arXiv AI
Sep 10

The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.

By Farah Atif, Sougata Saha, Monojit Choudhury
arXiv AI
Aug 3

Shall We Play a Game? Language Models for Open-ended Wargames

arXiv:2509. 17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosing what an actor says or does, deciding what happens after an action, or both.

By Glenn Matlin, Isaac Song, Yixiong Hao, Parv Mahajan, Evan Montoya, Ryan Bard, Stuart R. Topp, Anthony Wen-Ming Zang, Mohammed Rehan Parwani, Soham Shetty, Mark Riedl
arXiv AI
Aug 24

ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models

ExpertIVS is a framework that uses 14 sociological expert agents to interpret World Values Survey responses, reconstructing individual value systems in a coherent, internally consistent manner rather than simply concatenating survey answers. It introduces a multi‑agent debate mechanism to assess LLM alignment with these value profiles during dynamic interactions. Experiments on 480 individuals from 12 countries show a 90.78% value restoration fidelity and a 5.3% improvement in value generalization over baseline methods, while also demonstrating strong personality discriminability and behavioral consistency.

By Zhen Wang, Yuqi Ren, Yuehan Cui, Hongxiang Wang, Jianxiang Peng, Zhaoxia Zhang, Bingkun Zhu, Tongxuan Zhang, Dezhi Tong, Deyi Xiong
arXiv AI
Aug 5

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

arXiv:2608. 03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity.

By Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
arXiv AI
Sep 10

From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction

The paper investigates whether large language model (LLM) agents can simulate public deliberation by reflecting population opinion patterns and producing interaction-driven opinion change. Using census‑grounded Korean personas debating real policy questions, the study finds that persona agents fail to reliably reproduce population opinion patterns, often concentrating responses and reversing demographic differences. While deliberations generate reasoned, reciprocal arguments and some stance movement, much of this change occurs without peer exchange, and anchoring agents to population‑informed starting positions suppresses updating, indicating that population representation, argument generation, and interaction‑driven opinion change do not necessarily align.

By Chaemin Jang, Junsik Min, Jaewoo Choi, Donggyu Lee, Haiin Lee, Junyoung Park, Namhee Kim, Hyunwoo Kim, Jungwon Kim, Juho Kim, Nuri Kim, Jihee Kim
arXiv AI
Aug 19

CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents

CityReal is a modular framework that uses large language model agents to simulate human-aligned urban behavior. It models agents as intention-driven decision makers who pursue coherent mobility and activity plans, learning habits and preferences over time. By training textual adapters to align agent decisions with observed population statistics, CityReal improves realism at both micro and macro levels and can scale to tens of thousands of agents for analyzing crowd density, place popularity, mobility flows, and well‑being under various urban scenarios.

By Nicolas Bougie, Xiaotong Ye, Narimasa Watanabe
arXiv AI
Sep 2

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is a new multilingual benchmark that tests large language model agents on culturally grounded everyday workflows, offering 1,600 tasks in seven languages and eight cultures. The benchmark evaluates agents through structured sandbox actions and introduces Constrained Task Success (CTS), a metric that assesses task completion, minimal modification, and other complementary aspects via deterministic and LLM-as-a-Judge evaluations. Experiments show that even leading models achieve only 49.2% CTS, revealing significant gaps in correctness and state preservation across languages and cultures.

By Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
arXiv AI
6d ago

Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content

The study evaluates how well language‑model agents can simulate individual social media reactions by comparing predictions under different prompt conditions. Eight Serbian participants’ reactions to 68 posts were recorded, and four language models were asked to predict these reactions using prompts that varied in profile content and instruction style. The results show that prompts emphasizing attitudinal content and intuitive, immediate responses yield the highest fidelity, outperforming demographic backstories and a crowd baseline, and suggesting that such agents could act as general‑purpose simulated users.

By Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic, Jue Wang