Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study
Read the original on arXiv Computation and Language →The paper introduces a unified framework that simultaneously measures intrinsic (encoded) and extrinsic (expressed) gender bias in large language models using identical neutral prompts. It finds a consistent link between latent gender information and output bias, but shows that alignment via supervised fine‑tuning reduces expressed bias while leaving internal gender associations largely intact and reactivatable by adversarial prompts. The study also demonstrates that debiasing gains on structured benchmarks may not transfer to realistic tasks such as story generation.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.