The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.
By Farah Atif, Sougata Saha, Monojit Choudhury
The paper introduces Cultural Divergence Preservation (CDP), a new diagnostic for evaluating whether large language models (LLMs) preserve cross‑country differences when used as synthetic survey respondents. CDP uses a single human calibration to detect cultural flattening (reduced divergence) or caricature (increased divergence) and is shown to vary monotonically with cross‑country divergence, unlike conventional Jensen–Shannon divergence metrics. Experiments across multiple LLM backbones, prompting methods, and survey domains reveal that CDP uncovers systematic discrepancies with traditional fidelity metrics, highlighting that methods favored by those metrics can still produce strong flattening.
By Yeeun Chae, Yewon Choi, Seunghyun Lee, IL Im
arXiv:2608.28405v1 Announce Type: new
Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...
By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv:2609.10253v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situation...
By Bhuvan Arora, Devesh Saraogi, Sravya Varada, Dhruv Kumar
arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
arXiv:2601. 22396v2 Announce Type: replace-cross Abstract: Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain.
By Candida M. Greco, Lucio La Cava, Andrea Tagarelli