arXiv AI

From the Fluency Fallacy to the Micro-to-Macro Validity Gap: Opportunities and Pitfalls of LLMs in Social Simulation

arXiv Computation and Language
2d ago

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

The study audits 576 LLM-based social simulations from 350 papers using the PIMMUR framework, which evaluates agent profile, interaction, memory, minimal control, unawareness, and realism. Results show that PIMMUR principles are met more often than minimal control, unawareness, and realism, with frontier LLMs correctly identifying the underlying experiment in 65.2% of cases and half of prompts pre‑determining outcomes. Reproducing five experiments revealed that many reported collective phenomena disappear or reverse when PIMMUR principles are enforced, suggesting that apparent emergent behaviors may be methodological artifacts rather than genuine social dynamics.

By Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou, Man Ho Lam, Xintao Wang, Hao Zhu, Wenxuan Wang, Maarten Sap
arXiv AI
3d ago

The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.

By Farah Atif, Sougata Saha, Monojit Choudhury
arXiv AI
Jul 16

A Survey on Hypergame Theory: Modelling Misaligned Perceptions and Nested Beliefs for Multi-Agent Systems

arXiv:2507. 19593v3 Announce Type: replace Abstract: Classical game-theoretic models typically assume rational agents, complete information, and common knowledge of payoffs - assumptions that are often violated in real-world MAS characterized by uncertainty, misaligned perceptions, and nested beliefs.

By Vince Trencsenyi, Agnieszka Mensfelt, Kostas Stathis
arXiv AI
Jun 9

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond

arXiv:2604. 22748v2 Announce Type: replace Abstract: As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck.

By Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin, Lingdong Kong, Jize Zhang, Teng Tu, Weijian Ma, Ziqi Huang, Senqiao Yang, Wei Huang, Yeying Jin, Zhefan Rao, Jinhui Ye, Xinyu Lin, Xichen Zhang, Qisheng Hu, Shuai Yang, Leyang Shen, Wei Chow, Yifei Dong, Fengyi Wu, Quanyu Long, Bin Xia, Shaozuo Yu, Mingkang Zhu, Wenhu Zhang, Jiehui Huang, Haokun Gui, Runyi Li, Shiyi Du, Xu Huang, Dong Huang, Rui Liu, Chenyu Tang, Xuhang Chen, Chengzu Li, Haoxuan Che, Long Chen, Qifeng Chen, Wenxuan Zhang, Wenya Wang, Xiaojuan Qi, Yang Deng, Yanwei Li, Mike Zheng Shou, Zhi-Qi Cheng, See-Kiong Ng, Ziwei Liu, Philip Torr, Jiaya Jia
arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin