arXiv AI

CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

arXiv:2607. 05405v1 Announce Type: cross Abstract: To interact with users fairly and without stereotyping, AI models must display cultural competency, i.

arXiv AI
Sep 10

The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

The study introduces a World Values Survey–grounded simulation framework to test whether large language model agents can faithfully represent diverse human value systems. In about 4,000 conversations with 1,200 personas across three models, more than half of the agents failed to express their assigned value profiles from the start, and only 2–7% drifted over time. The results show systematic deviations from the intended value distributions and reveal that simulated dialogues differ from human discussions in their balance of stylistic consistency and semantic diversity.

By Farah Atif, Sougata Saha, Monojit Choudhury
arXiv AI
Sep 25

Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations

The paper introduces Cultural Divergence Preservation (CDP), a new diagnostic for evaluating whether large language models (LLMs) preserve cross‑country differences when used as synthetic survey respondents. CDP uses a single human calibration to detect cultural flattening (reduced divergence) or caricature (increased divergence) and is shown to vary monotonically with cross‑country divergence, unlike conventional Jensen–Shannon divergence metrics. Experiments across multiple LLM backbones, prompting methods, and survey domains reveal that CDP uncovers systematic discrepancies with traditional fidelity metrics, highlighting that methods favored by those metrics can still produce strong flattening.

By Yeeun Chae, Yewon Choi, Seunghyun Lee, IL Im
arXiv Computation and Language
Aug 31

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...

By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv AI
Jun 4

Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks

arXiv:2601. 22396v2 Announce Type: replace-cross Abstract: Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain.

By Candida M. Greco, Lucio La Cava, Andrea Tagarelli