arXiv AI By Siddharth Vohra, Manikandan Ravikiran

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

Read the original on arXiv AI →

The study examines how the design of audit questions influences perceived bias in large language models (LLMs). Using 40,726 requests across five models and three domains—hiring, lending, and medical triage—the authors find that demographic bias effects are not replicated when the audit is standardized. Instead, the audit’s construction, such as question phrasing and ordering, has a stronger impact on model responses than applicant demographics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
2d ago

Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

The study evaluates ten bias audit instruments across ten advanced language models on occupational gender, age, and socioeconomic status. While each tool reliably detects bias, their rankings of model performance are essentially random, indicating that different audits measure distinct constructs. The findings show that a single audit can identify bias direction within its own framework, but no audit can consistently rank models against one another.

By William Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, Jos\'e O. Gomes
arXiv AI
2d ago

Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits

The paper examines whether removing declared language fields from de‑identified résumés eliminates demographic leakage in large language models. By keeping language attributes identical and varying only unstructured prose across five ethnocultural groups and three cue‑salience levels, the authors find that non‑language text still allows target‑group recovery (average 0.757, reaching 1.000 under high salience). They also show that evaluation design—such as allowing or forbidding ties—dramatically affects LLM‑as‑a‑judge outcomes, underscoring the importance of evaluation protocol in bias audits.

By Qiangju Chen, Yang Xiao
arXiv AI
Sep 10

PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset

PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.

By Sumin Yu, Juhyeon Park, Taesup Moon