Camellia is a new benchmark that tests cultural bias in large language models (LLMs) across nine Asian languages and six Asian cultures. It contains 19,530 manually annotated entities linked to Asian or Western cultures and 2,173 masked social‑media contexts for these entities. Using Camellia, the authors evaluate four multilingual LLMs on cultural context adaptation, sentiment association, and entity extractive QA, finding that models struggle with cultural adaptation, exhibit differing biases across regions and families, and have difficulty understanding context in some Asian languages.
By Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu
The paper introduces WinoQueer-NL, a Dutch adaptation of the English WinoQueer benchmark, designed to assess anti‑queer bias in Dutch language models. After validating the dataset with 43 queer Dutch participants, the authors expanded it to 42,906 sentences and evaluated several Dutch and multilingual models, finding that while overall bias scores appeared neutral, specific identities—particularly transgender and non‑binary—were disproportionately favored in stereotypical sentences. The study underscores the need for culturally grounded datasets to identify and mitigate biases that affect marginalized groups in Dutch NLP systems.
By Jiska Beuk, Gerasimos Spanakis
The paper introduces WinoQueer‑NL, a Dutch adaptation of the English WinoQueer benchmark, designed to assess anti‑queer bias in Dutch language models. After validating the dataset with 43 queer Dutch participants, the authors released 42,906 sentences and evaluated several Dutch‑specific and multilingual models, finding that while overall bias scores were neutral, certain models disproportionately favored stereotypical statements for transgender and non‑binary identities. The study underscores the need for culturally grounded datasets to identify and mitigate biases that affect marginalized groups in Dutch NLP systems.
arXiv:2606. 09178v1 Announce Type: cross Abstract: Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that converts surface-level linguistic form while failing to reflect the cultural context embedded in threat scenarios, social norms, and legal frameworks.
By Hyeji Choi, Yongtaek Lim, Minwoo Kim
arXiv:2603. 13891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring.
By Petter T\"ornberg
Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that converts surface-level linguistic form while failing to reflect the cultural context embedded in threat scenarios, social norms, and legal frameworks. We construct paired DT and culturally-adapted (CA) datasets via 1:1 seed matching for four languages - Korean (KO), Japanese (JA), Thai (TH), and Khmer (KM) - and compare Attack Success Rate (ASR) and Cultural Realism scores across four open-source LLM.