The paper investigates how safety evaluations for large language models may mask ongoing gender discrimination, a phenomenon the authors term "harm laundering." By analyzing 450,000 gender‑directed completions across GPT‑2 to GPT‑5, they show that harmful content directed at women is transformed rather than removed, while men receive more positive representations. The study introduces a formal test and detection protocol for harm laundering, demonstrating that reduced toxicity scores do not necessarily indicate reduced representational harm.
arXiv:2609.38036v2 Announce Type: replace-cross
Abstract: Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support to...
By Edoardo Bolzoni, Valerio Capraro
arXiv:2609.38036v1 Announce Type: cross
Abstract: Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with...
By Edoardo Bolzoni, Valerio Capraro
The study evaluates gender representation in 8,000 images generated by four generations of the Stable Diffusion text‑to‑image model across 20 occupations and five prompt templates. It finds that 76.4% of the images depict male subjects, with 57.6% of historically female‑coded occupations also showing male subjects, and that newer model generations do not consistently reduce bias. Compared to U.S. Bureau of Labor Statistics data, the models underrepresent women by 20–46 percentage points, especially in near gender‑balanced fields such as scientists and cleaners.
By Shesh Narayan Gupta, Nik Bear Brown
arXiv:2608. 14577v1 Announce Type: cross Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis.
By Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang
arXiv:2605. 05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.
By Alif Al Hasan, Sumon Biswas
arXiv:2606. 04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.
By Yanjing Ren, Reza Ebrahimi, TengTeng Ma
The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.
By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv:2601. 17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compliance with harmful ones.
By Zhihao Zhang, Liting Huang, Guanghao Wu, Preslav Nakov, Heng Ji, Usman Naseem
arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
arXiv:2609.16366v1 Announce Type: cross
Abstract: When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "sh...
By Yingjia Wan, Lin Lin, Elisa Kreiss
arXiv:2510. 21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational biases is critical.
By Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, David C. Anastasiu, Kai Lukoff