arXiv AI By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram

Counterfactual Bias Testing for Application Tracking System

Read the original on arXiv AI →

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
6d ago

Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints

The study evaluates how to reduce fabricated claims in multi‑stage large language model (LLM) hiring pipelines. Prompt guardrails alone cut fabrication density by 86 % but still left half of outputs containing false claims, while adding a human‑in‑the‑loop checkpoint after resume improvement eliminated all identity fabrications and significantly lowered overall fabrication rates. The results show that a layered approach—combining prompt guardrails with human checkpoints—provides stronger protection against severe failures without harming the quality of the final outputs.

By Hiroko Takano