arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
By Yiming Wang, Jiayuan Di
The study examines how multilingual large language models (LLMs) produce outputs that differ across sociocultural contexts, highlighting that identity labels and source-language cues can mislead assessments of cultural grounding. Using a human‑validated, multi‑agent audit on 89,253 outputs from 12 LLMs in English, French, and Chinese across 18 occupations and three task conditions, the authors find that bias representation varies systematically by language and task. Removing direct identity cues reduces identity‑label prediction in English and Chinese but not in French, and the source language’s cultural context consistently receives the highest relevance score, though this signal weakens after translation or name masking.
"whyItMatters":"The findings show that surface cues can obscure true cross‑cultural patterns, underscoring the need for careful audit designs to avoid misleading conclusions about bias in multilingual LLMs."
By Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha
We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grounded and regionally differentiated knowledge within France. While prior benchmarks focus on national-level cultural understanding, they largely overlook intra-country variation and the need to distinguish between closely related regional contexts.
arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.
By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.
By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.
By Sophie Wu, Andrew Piper
arXiv:2608.28405v1 Announce Type: new
Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...
By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
The paper investigates how large language models (LLMs) represent national cultural change over time, using more than two decades of World Values Survey data and the Inglehart‑Welzel cultural map. It finds that while LLMs generally place countries near their most recent surveyed positions, their representations lag behind current data, under‑capture the magnitude of change, introduce spurious movements, and rarely reproduce trajectory reversals. These temporal inaccuracies reveal a flattening effect that limits the models’ cultural awareness and raises concerns for evaluation, representational harms, and governance of culturally aware AI systems.
By Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp
arXiv:2607.02235v2 Announce Type: replace-cross
Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to short...
By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
arXiv:2405.06818v2 Announce Type: replace
Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...
By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
arXiv:2608. 02486v1 Announce Type: cross Abstract: Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently.
By Iaroslav Chelombitko, Ekaterina Chelombitko, Mika H\"am\"al\"ainen
arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.
By Thennal DK, Hans Ole Hatzel