arXiv Machine Learning

What Must a Fairness Audit Report When Demographic Data Is Incomplete?

arXiv:2506. 23033v5 Announce Type: replace Abstract: Fairness audits are a key component of responsible machine-learning deployment.

arXiv AI
Sep 10

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

The study examines how the design of audit questions influences perceived bias in large language models (LLMs). Using 40,726 requests across five models and three domains—hiring, lending, and medical triage—the authors find that demographic bias effects are not replicated when the audit is standardized. Instead, the audit’s construction, such as question phrasing and ordering, has a stronger impact on model responses than applicant demographics.

By Siddharth Vohra, Manikandan Ravikiran
arXiv Machine Learning
Sep 24

When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations

The paper introduces FAPE, a four‑stage framework for evaluating the post‑processing fairness intervention ThresholdOptimizer across eight diverse domains, including criminal justice, finance, healthcare, and education. It reports that the intervention reduces disparity in most high‑disparity cases but can worsen fairness when baseline disparities are low, and that a single deployment‑time audit is unreliable without continuous monitoring and baseline‑disparity screening.

By Nithin Raghava Ramachandra Narla
arXiv AI
Sep 10

PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset

PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.

By Sumin Yu, Juhyeon Park, Taesup Moon
arXiv Machine Learning
Sep 22

CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

CleanScore is a black‑box audit framework that uses only scored outputs to assess whether models have been exposed to benchmark questions. It creates a public form and two fresh, independently written forms for each question, reports an interval for the public‑form advantage, and employs a private negative‑control bank with a transport radius to separate exposure from normal form mismatch. In a registered audit of five open models on 200 GSM8K and 200 ARC‑Challenge items, CleanScore found no exposure‑consistent advantage, bounding surface‑form inflation below five points, while also demonstrating how leaked items can inflate accuracy on unseen paraphrases and how planted advantages can be partially detected even after rewriting. whyItMatters":"The study shows that CleanScore can detect and quantify exposure effects in benchmark models, providing a more nuanced understanding of model performance beyond simple accuracy scores."

By Jeffery Opoku, David Banahene
arXiv Computation and Language
Sep 16

Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

The study evaluates ten bias audit instruments across ten advanced language models on occupational gender, age, and socioeconomic status. While each tool reliably detects bias, their rankings of model performance are essentially random, indicating that different audits measure distinct constructs. The findings show that a single audit can identify bias direction within its own framework, but no audit can consistently rank models against one another.

By William Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, Jos\'e O. Gomes
arXiv AI
Sep 17

Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.

By Vishnu Bindu Balachandran