Hugging Face Trending Papers

WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities

arXiv Computation and Language
6d ago

WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities

The paper introduces WinoQueer-NL, a Dutch adaptation of the English WinoQueer benchmark, designed to assess anti‑queer bias in Dutch language models. After validating the dataset with 43 queer Dutch participants, the authors expanded it to 42,906 sentences and evaluated several Dutch and multilingual models, finding that while overall bias scores appeared neutral, specific identities—particularly transgender and non‑binary—were disproportionately favored in stereotypical sentences. The study underscores the need for culturally grounded datasets to identify and mitigate biases that affect marginalized groups in Dutch NLP systems.

By Jiska Beuk, Gerasimos Spanakis
arXiv AI
Sep 1

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

The paper introduces a German-English benchmark dataset to evaluate anti‑LGBTQ biases in language models, combining community‑sourced stereotypes from German‑speaking queer individuals with a German translation of WinoQueer. Eight language models of varying sizes and architectures were assessed, revealing that they reproduce anti‑queer stereotypes with differences across identities and models. Fine‑tuning on community and progressive media content reduced bias on average, though the effect was not consistent across all models and identities.

By Melina Morch, Daniel Braun
arXiv AI
Jun 2

IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

arXiv:2606. 01260v1 Announce Type: cross Abstract: Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and localized stereotypes within its uniquely vast, multilingual, and diverse sociocultural landscape.

By Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto
arXiv Machine Learning
6d ago

GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models

The paper introduces GPTBIAS, a framework that uses powerful large language models like GPT‑4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the framework’s effectiveness and usability.

By Jiaxu Zhao, Meng Fang, Shirui Pan, Wenpeng Yin, Mykola Pechenizkiy
arXiv Machine Learning
Jul 24

How Robust Is Homogeneity Bias in LLMs? Evidence Across Models, Decoding Settings, and Identity Signals

arXiv:2501. 02211v3 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias generalizes across models, is stable under different inference settings, or depends on how group identity is signaled remains unstudied.

By Messi H. J. Lee
arXiv AI
Aug 25

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

The study examines how multilingual large language models (LLMs) produce outputs that differ across sociocultural contexts, highlighting that identity labels and source-language cues can mislead assessments of cultural grounding. Using a human‑validated, multi‑agent audit on 89,253 outputs from 12 LLMs in English, French, and Chinese across 18 occupations and three task conditions, the authors find that bias representation varies systematically by language and task. Removing direct identity cues reduces identity‑label prediction in English and Chinese but not in French, and the source language’s cultural context consistently receives the highest relevance score, though this signal weakens after translation or name masking. "whyItMatters":"The findings show that surface cues can obscure true cross‑cultural patterns, underscoring the need for careful audit designs to avoid misleading conclusions about bias in multilingual LLMs."

By Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha
arXiv Machine Learning
6d ago

Probing Cultural Signals in Large Language Models through Author Profiling

The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.

By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes