arXiv AI

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

arXiv:2510. 21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational biases is critical.

arXiv Machine Learning
Jul 24

How Robust Is Homogeneity Bias in LLMs? Evidence Across Models, Decoding Settings, and Identity Signals

arXiv:2501. 02211v3 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias generalizes across models, is stable under different inference settings, or depends on how group identity is signaled remains unstudied.

By Messi H. J. Lee
arXiv Machine Learning
Sep 22

Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes

The study audits demographic leakage in German-language resumes generated by large language models. Using ChatGPT, Gemini, and Qwen 3 variants, the authors generate resumes from anonymized profiles, varying only gender- and ethnicity-associated names while keeping qualifications constant. Even after anonymization and gender-neutralization, classifiers can reliably distinguish male- from female-generated resumes, driven by subtle differences in gender-neutral terminology rather than overtly gendered wording; ethnicity-related leakage remains weak.

By Charlotte Leininger, Helena Veit, Matthias A{\ss}enmacher, Andreas Bender
arXiv AI
Sep 17

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.

By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv AI
Sep 17

Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations

The study evaluates gender representation in 8,000 images generated by four generations of the Stable Diffusion text‑to‑image model across 20 occupations and five prompt templates. It finds that 76.4% of the images depict male subjects, with 57.6% of historically female‑coded occupations also showing male subjects, and that newer model generations do not consistently reduce bias. Compared to U.S. Bureau of Labor Statistics data, the models underrepresent women by 20–46 percentage points, especially in near gender‑balanced fields such as scientists and cleaners.

By Shesh Narayan Gupta, Nik Bear Brown