The paper introduces a unified framework that simultaneously measures intrinsic (encoded) and extrinsic (expressed) gender bias in large language models using identical neutral prompts. It finds a consistent link between latent gender information and output bias, but shows that alignment via supervised fine‑tuning reduces expressed bias while leaving internal gender associations largely intact and reactivatable by adversarial prompts. The study also demonstrates that debiasing gains on structured benchmarks may not transfer to realistic tasks such as story generation.
By Nour Bouchouchi, Thibault Laugel, Xavier Renard, Christophe Marsala, Marie-Jeanne Lesot, Marcin Detyniecki
arXiv:2609.38036v2 Announce Type: replace-cross
Abstract: Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support to...
By Edoardo Bolzoni, Valerio Capraro
arXiv:2609.38036v1 Announce Type: cross
Abstract: Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with...
By Edoardo Bolzoni, Valerio Capraro
arXiv:2609.20838v1 Announce Type: new
Abstract: In this study, we examine how modern LLMs generate and detect fake news under controlled settings across four manipulation scenarios. These are open-en...
By Zeynep \"Ozdemir, Murat Osmano\u{g}lu, Sevgi Yi\u{g}it-Sert, \"Omer \"Ozg\"ur Tanr{\i}\"over, Y{\i}lmaz Ar
arXiv:2512. 15792v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have rapidly become indispensable tools for acquiring information and supporting human decision-making.
By Xulang Zhang, Rui Mao, Erik Cambria
FakeSpotter is a new tool that estimates the viral misinformation risk of textual content by measuring structural fingerprints of misinformation instead of directly judging truthfulness. It operates across linguistic, narrative, logical, and critical‑thinking dimensions, using repeated large language model assessments and domain‑specific logistic regression classifiers for both short and long texts. In a labeled corpus of 764 texts, FakeSpotter achieved macro F1 scores of 0.788 for short texts and 0.793 for long texts, and its interpretive layer offers explainable outputs such as feature‑based scores, signal agreement, and a caution index for social listening.
By Giovanni Spitale, Federico Germani