arXiv:2508. 08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures.
By Haeun Yu, Arnav Arora Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
arXiv:2601. 14063v2 Announce Type: replace-cross Abstract: Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cultural contexts.
By Mohsinul Kabir, Tasnim Ahmed, Md Mezbaur Rahman, Shaoxiong Ji, Hassan Alhuzali, Yuechen Jiang, Jimin Huang, Sophia Ananiadou
CHARM is a multicultural benchmark that tests large language models’ ability to adopt a character’s style while respecting knowledge boundaries. It includes 40 real and fictional characters from five cultural-linguistic regions and evaluates two boundary types—Temporal and Cross-Universe—using abstention-enabled multiple-choice questions. The study finds that hallucinations mainly stem from compliance failures: models often recognize a query is out of scope yet still provide out-of-character answers, revealing systematic cultural variations in these errors.
By Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak
arXiv:2606. 07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed through English or the local language.
By Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis
The study investigates whether large language models (LLMs) are more prone to errors when they doubt the plausibility of input data, a phenomenon termed context‑memory conflict. Using non‑English and low‑resource language datasets, the authors generate text from factual, counterfactual, and fictional RDF triples in English, Czech, Slovak, and Upper Sorbian, and evaluate faithfulness with both human annotations and an LLM judge (Kimi K3). Contrary to expectations, the results show only a weak context‑memory conflict: counterfactual inputs receive slightly lower faithfulness scores than factual ones, and the choice of LLM judge can significantly affect perceived conflict strength.
By Peter Kochelka, Ale\v{s} Manuel Pap\'a\v{c}ek, Vojt\v{e}ch Dvo\v{r}\'ak, Ond\v{r}ej Du\v{s}ek
Camellia is a new benchmark that tests cultural bias in large language models (LLMs) across nine Asian languages and six Asian cultures. It contains 19,530 manually annotated entities linked to Asian or Western cultures and 2,173 masked social‑media contexts for these entities. Using Camellia, the authors evaluate four multilingual LLMs on cultural context adaptation, sentiment association, and entity extractive QA, finding that models struggle with cultural adaptation, exhibit differing biases across regions and families, and have difficulty understanding context in some Asian languages.
By Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu
arXiv:2609.08322v1 Announce Type: cross
Abstract: Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is...
By Ariun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen, Josef van Genabith, Koel Dutta Chowdhury
The paper "Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems" investigates how commercial text‑to‑image models silently modify user prompts before generating images, a step that is often hidden from users. Using the multilingual benchmark WORLDVIEW, the authors audit the revision layer in DALL‑E‑3, Imagen‑4, and GPT‑Image‑1.5, finding that non‑Western and non‑Anglophone contexts are disproportionately marked, reduced to narrow vocabularies, and stereotyped. The study demonstrates that the revision layer itself is a previously undocumented causal source of cultural stereotyping, underscoring the need to audit deployed systems rather than just the underlying models.
By Aleksandra Urman, Elsa Lichtenegger, Salima Jaoua, Azza Bouleimen, Robin Forsberg, Corinna Hertweck, Stefania Ionescu, Nicol\`o Pagan, Ancsa Hannak, Joachim Baumann
The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.
By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
The study investigates how large language models (LLMs) handle diverse Indian oral traditions, using the Rajasthani Pabuji epic, Tamil Sangam poetry, and Bengali folk tales as case studies. By prompting Claude Sonnet and Gemini with 54 generation requests across generic, culturally specific, and regional-language prompts, the authors measured reference drift and cross-tradition convergence using Sentence‑BERT embeddings. Results show that while outputs stay closer to their own tradition than to others, there is significant cross‑tradition similarity (0.52–0.66), indicating partial homogenisation; moreover, regional‑language prompting consistently reduced fidelity to authentic traditions.
By Paarth Singh Rathore
The study examines how multilingual large language models (LLMs) produce outputs that differ across sociocultural contexts, highlighting that identity labels and source-language cues can mislead assessments of cultural grounding. Using a human‑validated, multi‑agent audit on 89,253 outputs from 12 LLMs in English, French, and Chinese across 18 occupations and three task conditions, the authors find that bias representation varies systematically by language and task. Removing direct identity cues reduces identity‑label prediction in English and Chinese but not in French, and the source language’s cultural context consistently receives the highest relevance score, though this signal weakens after translation or name masking.
"whyItMatters":"The findings show that surface cues can obscure true cross‑cultural patterns, underscoring the need for careful audit designs to avoid misleading conclusions about bias in multilingual LLMs."
By Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha
arXiv:2609.00491v1 Announce Type: new
Abstract: Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect larg...
By Hangxiao Zhu, Suliu Qin, Zhuoyan Li, Ming Jiang, Yu Zhang, Meng Xia