arXiv:2407.20371v3 Announce Type: replace-cross
Abstract: Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the...
By Kyra Wilson, Aylin Caliskan
arXiv:2607. 20073v1 Announce Type: new Abstract: AI-based recruitment systems that rely on machine learning models trained on historical CV data, risk perpetuating and amplifying social biases.
By Farnaz Faramarzi Lighvan, Lynn Houthuys
Large language models (LLMs) are increasingly deployed in hiring workflows, yet most research on gender bias in LLM hiring decisions has focused on English-language, Western-format resumes. This study examines whether pro-female gender bias extends to a Japanese corporate context and evaluates two practical mitigation strategies.
arXiv:2601. 06861v2 Announce Type: replace-cross Abstract: Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions.
By William Guey, Wei Zhang, Pei-Luen Patrick Rau, Pierrick Bougault, Vitor D. de Moura, Bertan Ucar, Jose O. Gomes
arXiv:2607. 28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender.
By Martin Lukk (University of Toronto)
PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.
By Sumin Yu, Juhyeon Park, Taesup Moon
The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.
By Kosuke Kitahara, Nobuhiro Yamaguchi
The paper investigates whether language models still encode occupational biases even when they appear unbiased in behavioral tests. Using a causal framework, the authors separate bias into internal representations of user competence and observable outputs, deriving steering vectors that show these representations influence model behavior in question‑answering and hiring tasks. Across several open‑weight models, demographic factors such as gender, race, and socioeconomic status affect the models’ internal competence representations, revealing hidden bias that behavioral metrics alone may miss.
By Keren Fuentes, Aaron Mueller
arXiv:2603. 13891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring.
By Petter T\"ornberg
The paper introduces a unified framework that simultaneously measures intrinsic (encoded) and extrinsic (expressed) gender bias in large language models using identical neutral prompts. It finds a consistent link between latent gender information and output bias, but shows that alignment via supervised fine‑tuning reduces expressed bias while leaving internal gender associations largely intact and reactivatable by adversarial prompts. The study also demonstrates that debiasing gains on structured benchmarks may not transfer to realistic tasks such as story generation.
By Nour Bouchouchi, Thibault Laugel, Xavier Renard, Christophe Marsala, Marie-Jeanne Lesot, Marcin Detyniecki
The paper examines whether removing declared language fields from de‑identified résumés eliminates demographic leakage in large language models. By keeping language attributes identical and varying only unstructured prose across five ethnocultural groups and three cue‑salience levels, the authors find that non‑language text still allows target‑group recovery (average 0.757, reaching 1.000 under high salience). They also show that evaluation design—such as allowing or forbidding ties—dramatically affects LLM‑as‑a‑judge outcomes, underscoring the importance of evaluation protocol in bias audits.
By Qiangju Chen, Yang Xiao
arXiv:2510. 21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational biases is critical.
By Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, David C. Anastasiu, Kai Lukoff