arXiv AI

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

arXiv:2507. 11548v3 Announce Type: replace-cross Abstract: The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment.

arXiv AI
Sep 10

PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset

PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.

By Sumin Yu, Juhyeon Park, Taesup Moon
arXiv AI
Aug 28

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv Machine Learning
Sep 22

Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes

The study audits demographic leakage in German-language resumes generated by large language models. Using ChatGPT, Gemini, and Qwen 3 variants, the authors generate resumes from anonymized profiles, varying only gender- and ethnicity-associated names while keeping qualifications constant. Even after anonymization and gender-neutralization, classifiers can reliably distinguish male- from female-generated resumes, driven by subtle differences in gender-neutral terminology rather than overtly gendered wording; ethnicity-related leakage remains weak.

By Charlotte Leininger, Helena Veit, Matthias A{\ss}enmacher, Andreas Bender
arXiv Machine Learning
Sep 3

FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making

FairLens is a benchmark and evaluation framework that measures fairness and validity of vision‑language models (VLMs) in high‑stakes domains such as hiring, legal, and healthcare. It uses over 100,000 face‑image and question pairs covering gender, race, and age, and assesses responses through demographic parity, soundness, demographic association, and bias in free‑text generation. The study finds that VLMs often make unwarranted inferences from faces rather than abstaining, especially in legal and healthcare contexts, and that small parity gaps can still hide unsafe treatment across groups.

By Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza
arXiv Computation and Language
Sep 16

Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

The study evaluates ten bias audit instruments across ten advanced language models on occupational gender, age, and socioeconomic status. While each tool reliably detects bias, their rankings of model performance are essentially random, indicating that different audits measure distinct constructs. The findings show that a single audit can identify bias direction within its own framework, but no audit can consistently rank models against one another.

By William Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, Jos\'e O. Gomes
arXiv AI
Sep 17

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.

By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv AI
Jul 14

BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts

arXiv:2601. 06861v2 Announce Type: replace-cross Abstract: Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions.

By William Guey, Wei Zhang, Pei-Luen Patrick Rau, Pierrick Bougault, Vitor D. de Moura, Bertan Ucar, Jose O. Gomes