Hugging Face Trending Papers

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

Read the original on Hugging Face Trending Papers →

The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination, a phenomenon the authors term "harm laundering." By analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, they show that harmful content directed at women is transformed rather than removed, while men receive more positive representations. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily indicate reduced representational harm.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 18

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination by transforming harmful content rather than eliminating it, a phenomenon termed "harm laundering." Analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, the authors find that sexual violence content directed at women disappears while men receive more positive representations, with GPT‑5 showing stark disparities such as framing breast cancer as a men’s rights debate. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily reflect reduced representational harm.

By Sarah Wyer, Sue Black, Noura Al Moubayed
arXiv AI
Sep 17

Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations

The study evaluates gender representation in 8,000 images generated by four generations of the Stable Diffusion text‑to‑image model across 20 occupations and five prompt templates. It finds that 76.4% of the images depict male subjects, with 57.6% of historically female‑coded occupations also showing male subjects, and that newer model generations do not consistently reduce bias. Compared to U.S. Bureau of Labor Statistics data, the models underrepresent women by 20–46 percentage points, especially in near gender‑balanced fields such as scientists and cleaners.

By Shesh Narayan Gupta, Nik Bear Brown