arXiv Computation and Language By William Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, Jos\'e O. Gomes

Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

Read the original on arXiv Computation and Language →

The study evaluates ten bias audit instruments across ten advanced language models on occupational gender, age, and socioeconomic status. While each tool reliably detects bias, their rankings of model performance are essentially random, indicating that different audits measure distinct constructs. The findings show that a single audit can identify bias direction within its own framework, but no audit can consistently rank models against one another.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 10

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

The study examines how the design of audit questions influences perceived bias in large language models (LLMs). Using 40,726 requests across five models and three domains—hiring, lending, and medical triage—the authors find that demographic bias effects are not replicated when the audit is standardized. Instead, the audit’s construction, such as question phrasing and ordering, has a stronger impact on model responses than applicant demographics.

By Siddharth Vohra, Manikandan Ravikiran
arXiv AI
Aug 28

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv AI
Jul 14

BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts

arXiv:2601. 06861v2 Announce Type: replace-cross Abstract: Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions.

By William Guey, Wei Zhang, Pei-Luen Patrick Rau, Pierrick Bougault, Vitor D. de Moura, Bertan Ucar, Jose O. Gomes