arXiv AI

Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations

The study evaluates gender representation in 8,000 images generated by four generations of the Stable Diffusion text‑to‑image model across 20 occupations and five prompt templates. It finds that 76.4% of the images depict male subjects, with 57.6% of historically female‑coded occupations also showing male subjects, and that newer model generations do not consistently reduce bias. Compared to U.S. Bureau of Labor Statistics data, the models underrepresent women by 20–46 percentage points, especially in near gender‑balanced fields such as scientists and cleaners.

arXiv Computer Vision
Sep 24

Gender Bias in Vision-Language In-Context Learning

The paper investigates how in‑context learning (ICL) in large vision‑language models (LVLMs) can amplify gender bias. Using the VL‑BICLE framework, the authors show that gendered ICL demonstrations shift model bias toward the demonstrated gender, especially in tasks involving gendered language such as image captioning and pronoun prediction. They find that similarity‑based retrieval does not mitigate this bias and that replacing real images with synthetic ones from stable diffusion reduces bias without hurting caption quality.

By Tong Xiang, Noa Garcia, Yuta Nakashima
arXiv AI
Sep 18

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination by transforming harmful content rather than eliminating it, a phenomenon termed "harm laundering." Analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, the authors find that sexual violence content directed at women disappears while men receive more positive representations, with GPT‑5 showing stark disparities such as framing breast cancer as a men’s rights debate. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily reflect reduced representational harm.

By Sarah Wyer, Sue Black, Noura Al Moubayed
Hugging Face Trending Papers
Sep 17

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination, a phenomenon the authors term "harm laundering." By analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, they show that harmful content directed at women is transformed rather than removed, while men receive more positive representations. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily indicate reduced representational harm.

arXiv Computer Vision
Sep 1

Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration

arXiv:2608.30210v1 Announce Type: cross Abstract: AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to...

By Sunwhi Kim (Hwasung Medi-Science University, Dept. of Bio-Healthcare), Sunyul Kim (Yonsei University, Graduate School of Engineering, Dept. of Artificial Intelligence), Meounggun Jo (Hoseo University), Jini Tae (Gwangju Institute of Science and Technology, School of Humanities and Social Sciences)
arXiv AI
Sep 11

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

The study investigates how speech‑to‑speech (S2S) models handle gender, distinguishing between the acoustic voice and the content’s gender cues. Experiments across five models in English, Spanish, and Mandarin show that while the rendered voice remains unbiased, the models consistently attribute speaker gender based on textual content rather than voice. When content and voice disagree, misgendering rates soar to 90%, whereas agreement yields only 2% misgendering.

By Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji
arXiv Machine Learning
Sep 3

FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making

FairLens is a benchmark and evaluation framework that measures fairness and validity of vision‑language models (VLMs) in high‑stakes domains such as hiring, legal, and healthcare. It uses over 100,000 face‑image and question pairs covering gender, race, and age, and assesses responses through demographic parity, soundness, demographic association, and bias in free‑text generation. The study finds that VLMs often make unwarranted inferences from faces rather than abstaining, especially in legal and healthcare contexts, and that small parity gaps can still hide unsafe treatment across groups.

By Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza
arXiv Machine Learning
Sep 17

Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives

The paper audits vision‑language models (CLIP) for gender bias using 1,500 artworks from the Metropolitan Museum of Art, focusing on zero‑shot logit differences for prompts like "masterpiece," "quality," and "influence." Unadjusted results show no significant gender effect and statistical equivalence across models, but high residual variance suggests that global zero‑shot metrics are largely noise‑dominated and may miss fine‑grained biases. The study underscores the need for multivariate confound control, equivalence testing, and provenance auditing when evaluating AI fairness in cultural heritage data.

By Manpreet Singh, Rhythm Bhatia, Rahul Joshi