arXiv Machine Learning

Probing Cultural Signals in Large Language Models through Author Profiling

arXiv AI
Jun 2

IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

arXiv:2606. 01260v1 Announce Type: cross Abstract: Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and localized stereotypes within its uniquely vast, multilingual, and diverse sociocultural landscape.

By Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto
arXiv Computation and Language
2d ago

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Camellia is a new benchmark that tests cultural bias in large language models (LLMs) across nine Asian languages and six Asian cultures. It contains 19,530 manually annotated entities linked to Asian or Western cultures and 2,173 masked social‑media contexts for these entities. Using Camellia, the authors evaluate four multilingual LLMs on cultural context adaptation, sentiment association, and entity extractive QA, finding that models struggle with cultural adaptation, exhibit differing biases across regions and families, and have difficulty understanding context in some Asian languages.

By Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu
arXiv Machine Learning
Jul 24

How Robust Is Homogeneity Bias in LLMs? Evidence Across Models, Decoding Settings, and Identity Signals

arXiv:2501. 02211v3 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias generalizes across models, is stable under different inference settings, or depends on how group identity is signaled remains unstudied.

By Messi H. J. Lee
arXiv Computation and Language
3d ago

CoCoA: Context-Conditional Cultural Alignment for Large Language Models

CoCoA (Context-Conditional Cultural Alignment) is a framework designed to mitigate cultural bias in large language models by learning context-conditional behavior. It trains on entity pairs under both culturally cued and neutral contexts, using a contrastive alignment objective combined with calibration, drift regularization, and goal-aware gradient reconciliation. Evaluations on CAMeL and Camellia across ten languages and four LLMs show that CoCoA reduces the Cultural Bias Score from 43 to 24 on average while keeping near-neutral preferences at 50.2, with minimal impact on general performance.

By Kyungdon Lee, Wei Xu, Alan Ritter, Dong-Ho Lee, JinYeong Bak