The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.
By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
The study evaluates how to reduce fabricated claims in multi‑stage large language model (LLM) hiring pipelines. Prompt guardrails alone cut fabrication density by 86 % but still left half of outputs containing false claims, while adding a human‑in‑the‑loop checkpoint after resume improvement eliminated all identity fabrications and significantly lowered overall fabrication rates. The results show that a layered approach—combining prompt guardrails with human checkpoints—provides stronger protection against severe failures without harming the quality of the final outputs.
By Hiroko Takano
arXiv:2608. 06949v1 Announce Type: new Abstract: Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias.
By Paul-Peter Arslan
The paper introduces candidate‑fate accounting, an audit framework for transparent sensor diagnostic pipeline search that records every candidate, including invalid, pruned, or skipped ones, and assigns a terminal fate to each. It enhances traceability by hashing repeated observations, flagging illegal candidates, and documenting budget rationales. Experiments on three bearing‑diagnostic datasets demonstrate that the framework uncovers 30–41 omitted candidates and verifies complete accounting while preserving competitive performance.
By Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
arXiv:2606. 16723v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.
By Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor
arXiv:2507. 11548v3 Announce Type: replace-cross Abstract: The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment.
By Kevin T Webster
arXiv:2607. 13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows.
By Keyur Gabani
LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question...
LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines.
arXiv:2608. 19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
By Bin Zhu, Yi Xie, Yanghui Rao
The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.
By Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig