arXiv Computation and Language

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

The study audits 576 LLM-based social simulations from 350 papers using the PIMMUR framework, which evaluates agent profile, interaction, memory, minimal control, unawareness, and realism. Results show that PIMMUR principles are met more often than minimal control, unawareness, and realism, with frontier LLMs correctly identifying the underlying experiment in 65.2% of cases and half of prompts pre‑determining outcomes. Reproducing five experiments revealed that many reported collective phenomena disappear or reverse when PIMMUR principles are enforced, suggesting that apparent emergent behaviors may be methodological artifacts rather than genuine social dynamics.

arXiv Computation and Language
Aug 31

Benchmarking large language model agent societies against human behavioural distributions

The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.

By Raad Bin Tareaf
arXiv AI
3d ago

The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.

By Farah Atif, Sougata Saha, Monojit Choudhury