arXiv AI By Sumin Yu, Juhyeon Park, Taesup Moon

PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset

Read the original on arXiv AI →

PopResume is a population‑representative resume dataset designed for causal fairness auditing of large language model (LLM) and vision‑language model (VLM) resume screeners. It grounds fairness evaluation in real population statistics and preserves natural attribute relationships, enabling path‑specific effect (PSE) analysis that separates business‑necessity from redlining pathways. Using PopResume, the authors evaluated eight models on 60.8K resumes across five occupations and uncovered five discrimination patterns that aggregate metrics missed, demonstrating the value of causally‑grounded auditing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv AI
Sep 3

Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems

The paper introduces SCOPED‑Hiring, a process‑aware fairness diagnosis pipeline for large language model (LLM) based multi‑agent hiring systems. It generates controlled resume variants, runs role‑based hiring committees, and logs over 311,000 structured decision trajectories, converting them into quantitative fairness signals across six diagnostic lenses: final outcome, counterfactual, process, pathway, dynamic, and design effects. The study finds that balanced final hire rates can conceal hidden trajectory unfairness—such as career gaps, proxy cues, and identity cues—and demonstrates that targeted repairs guided by these diagnoses can reduce the total layered burden by 72.3% while only slightly altering the hire rate.

By Yiran Zhao, Lu Zhou, Liming Fang, Yufei Chen, Jiafei Wu, Zhe Liu, Xiaogang Xu