arXiv Computation and Language By Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Read the original on arXiv Computation and Language →

Camellia is a new benchmark that tests cultural bias in large language models (LLMs) across nine Asian languages and six Asian cultures. It contains 19,530 manually annotated entities linked to Asian or Western cultures and 2,173 masked social‑media contexts for these entities. Using Camellia, the authors evaluate four multilingual LLMs on cultural context adaptation, sentiment association, and entity extractive QA, finding that models struggle with cultural adaptation, exhibit differing biases across regions and families, and have difficulty understanding context in some Asian languages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 2

IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

arXiv:2606. 01260v1 Announce Type: cross Abstract: Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and localized stereotypes within its uniquely vast, multilingual, and diverse sociocultural landscape.

By Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto
arXiv Computation and Language
4d ago

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.

By Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
arXiv Computation and Language
3d ago

CoCoA: Context-Conditional Cultural Alignment for Large Language Models

CoCoA (Context-Conditional Cultural Alignment) is a framework designed to mitigate cultural bias in large language models by learning context-conditional behavior. It trains on entity pairs under both culturally cued and neutral contexts, using a contrastive alignment objective combined with calibration, drift regularization, and goal-aware gradient reconciliation. Evaluations on CAMeL and Camellia across ten languages and four LLMs show that CoCoA reduces the Cultural Bias Score from 43 to 24 on average while keeping near-neutral preferences at 50.2, with minimal impact on general performance.

By Kyungdon Lee, Wei Xu, Alan Ritter, Dong-Ho Lee, JinYeong Bak
arXiv AI
Aug 25

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

The study examines how multilingual large language models (LLMs) produce outputs that differ across sociocultural contexts, highlighting that identity labels and source-language cues can mislead assessments of cultural grounding. Using a human‑validated, multi‑agent audit on 89,253 outputs from 12 LLMs in English, French, and Chinese across 18 occupations and three task conditions, the authors find that bias representation varies systematically by language and task. Removing direct identity cues reduces identity‑label prediction in English and Chinese but not in French, and the source language’s cultural context consistently receives the highest relevance score, though this signal weakens after translation or name masking. "whyItMatters":"The findings show that surface cues can obscure true cross‑cultural patterns, underscoring the need for careful audit designs to avoid misleading conclusions about bias in multilingual LLMs."

By Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha