arXiv:2502.10577v2 Announce Type: replace-cross
Abstract: Instruct-based large language models (LLMs) have been shown to propagate and even amplify gender bias when prompted with contextually constra...
By Enzo Doyen, Amalia Todirascu
The paper introduces a unified framework that simultaneously measures intrinsic (encoded) and extrinsic (expressed) gender bias in large language models using identical neutral prompts. It finds a consistent link between latent gender information and output bias, but shows that alignment via supervised fine‑tuning reduces expressed bias while leaving internal gender associations largely intact and reactivatable by adversarial prompts. The study also demonstrates that debiasing gains on structured benchmarks may not transfer to realistic tasks such as story generation.
By Nour Bouchouchi, Thibault Laugel, Xavier Renard, Christophe Marsala, Marie-Jeanne Lesot, Marcin Detyniecki
The study introduces GAPA, a dataset of 316 physical attributes with 14,706 gender-association ratings from 304 US annotators, showing that such descriptions carry structured gender associations. It evaluates 16 LLMs, finding they partially mirror human ratings but exhibit biases such as compressed distributions, weaker alignment for men, and asymmetric abstention toward non‑binary identities. A proxy model trained on these data is released and applied to analyze character descriptions in LitBank, illustrating the persistence of gendered interpretations in ostensibly neutral language.
By Yingjia Wan, Lin Lin, Elisa Kreiss
arXiv:2609.16366v1 Announce Type: cross
Abstract: When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "sh...
By Yingjia Wan, Lin Lin, Elisa Kreiss
arXiv:2603. 23485v2 Announce Type: replace-cross Abstract: Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses.
By Sagar Kumar, Ariel Flint, Luca Maria Aiello, Andrea Baronchelli
arXiv:2609.38036v2 Announce Type: replace-cross
Abstract: Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support to...
By Edoardo Bolzoni, Valerio Capraro