arXiv:2506. 23033v2 Announce Type: replace Abstract: Fairness audits are a key component of responsible machine-learning deployment.
By Yash Vardhan Tomar
arXiv:2601.03087v2 Announce Type: replace
Abstract: Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM...
By David Hartmann, Lena Pohlmann, Lelia Hanslik, Noah Gie{\ss}ing, Bettina Berendt, Pieter Delobelle
The study examines how the design of audit questions influences perceived bias in large language models (LLMs). Using 40,726 requests across five models and three domains—hiring, lending, and medical triage—the authors find that demographic bias effects are not replicated when the audit is standardized. Instead, the audit’s construction, such as question phrasing and ordering, has a stronger impact on model responses than applicant demographics.
By Siddharth Vohra, Manikandan Ravikiran
arXiv:2610.01005v1 Announce Type: new
Abstract: As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairnes...
By Jie Tang, Chuanlong Xie, Lixing Zhu
The paper introduces FAPE, a four‑stage framework for evaluating the post‑processing fairness intervention ThresholdOptimizer across eight diverse domains, including criminal justice, finance, healthcare, and education. It reports that the intervention reduces disparity in most high‑disparity cases but can worsen fairness when baseline disparities are low, and that a single deployment‑time audit is unreliable without continuous monitoring and baseline‑disparity screening.
By Nithin Raghava Ramachandra Narla
PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.
By Sumin Yu, Juhyeon Park, Taesup Moon