arXiv Machine Learning By Yash Vardhan Tomar

What Must a Fairness Audit Report When Demographic Data Is Incomplete?

Read the original on arXiv Machine Learning →

arXiv:2506. 23033v5 Announce Type: replace Abstract: Fairness audits are a key component of responsible machine-learning deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

The study examines how the design of audit questions influences perceived bias in large language models (LLMs). Using 40,726 requests across five models and three domains—hiring, lending, and medical triage—the authors find that demographic bias effects are not replicated when the audit is standardized. Instead, the audit’s construction, such as question phrasing and ordering, has a stronger impact on model responses than applicant demographics.

By Siddharth Vohra, Manikandan Ravikiran
arXiv Machine Learning
Sep 24

When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations

The paper introduces FAPE, a four‑stage framework for evaluating the post‑processing fairness intervention ThresholdOptimizer across eight diverse domains, including criminal justice, finance, healthcare, and education. It reports that the intervention reduces disparity in most high‑disparity cases but can worsen fairness when baseline disparities are low, and that a single deployment‑time audit is unreliable without continuous monitoring and baseline‑disparity screening.

By Nithin Raghava Ramachandra Narla
arXiv AI
Sep 10

PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset

PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.

By Sumin Yu, Juhyeon Park, Taesup Moon