The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.
By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
The study evaluates how to reduce fabricated claims in multi‑stage large language model (LLM) hiring pipelines. Prompt guardrails alone cut fabrication density by 86 % but still left half of outputs containing false claims, while adding a human‑in‑the‑loop checkpoint after resume improvement eliminated all identity fabrications and significantly lowered overall fabrication rates. The results show that a layered approach—combining prompt guardrails with human checkpoints—provides stronger protection against severe failures without harming the quality of the final outputs.
By Hiroko Takano
arXiv:2608. 06949v1 Announce Type: new Abstract: Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias.
By Paul-Peter Arslan
The paper introduces candidate‑fate accounting, an audit framework for transparent sensor diagnostic pipeline search that records every candidate, including invalid, pruned, or skipped ones, and assigns a terminal fate to each. It enhances traceability by hashing repeated observations, flagging illegal candidates, and documenting budget rationales. Experiments on three bearing‑diagnostic datasets demonstrate that the framework uncovers 30–41 omitted candidates and verifies complete accounting while preserving competitive performance.
By Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor