Hugging Face Trending Papers

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination, a phenomenon the authors term "harm laundering." By analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, they show that harmful content directed at women is transformed rather than removed, while men receive more positive representations. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily indicate reduced representational harm.

arXiv AI
Sep 18

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination by transforming harmful content rather than eliminating it, a phenomenon termed "harm laundering." Analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, the authors find that sexual violence content directed at women disappears while men receive more positive representations, with GPT‑5 showing stark disparities such as framing breast cancer as a men’s rights debate. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily reflect reduced representational harm.

By Sarah Wyer, Sue Black, Noura Al Moubayed
arXiv AI
Sep 17

Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations

The study evaluates gender representation in 8,000 images generated by four generations of the Stable Diffusion text‑to‑image model across 20 occupations and five prompt templates. It finds that 76.4% of the images depict male subjects, with 57.6% of historically female‑coded occupations also showing male subjects, and that newer model generations do not consistently reduce bias. Compared to U.S. Bureau of Labor Statistics data, the models underrepresent women by 20–46 percentage points, especially in near gender‑balanced fields such as scientists and cleaners.

By Shesh Narayan Gupta, Nik Bear Brown
arXiv AI
Sep 18

How Humans and LLMs Read Gender into "Gender-Neutral" Physical Descriptions

The study introduces GAPA, a dataset of 316 physical attributes with 14,706 gender-association ratings from 304 US annotators, showing that such descriptions carry structured gender associations. It evaluates 16 LLMs, finding they partially mirror human ratings but exhibit biases such as compressed distributions, weaker alignment for men, and asymmetric abstention toward non‑binary identities. A proxy model trained on these data is released and applied to analyze character descriptions in LitBank, illustrating the persistence of gendered interpretations in ostensibly neutral language.

By Yingjia Wan, Lin Lin, Elisa Kreiss