The paper reviews the evolution of AI recruitment systems from simple profile matching to complex, multi‑stage workflows that retrieve evidence, compare candidates, and execute actions. It analyzes 40 representative works, highlighting transitions from similarity to reciprocal suitability, from single models to compound workflows, and from offline predictions to evidence‑aligned evaluation. The authors identify persistent gaps—such as confounded behavioral labels, limited data validity, and lack of privacy assessment—and propose a staged mapping for defensible evaluation and an agenda for auditable, evidence‑grounded systems.
By Ziyi Zhao, Guanzheng Wei
The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.
By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv:2610.08364v1 Announce Type: new
Abstract: Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions fr...
By Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas, Cozmin Ududec
The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.
By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv:2407.20371v3 Announce Type: replace-cross
Abstract: Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the...
By Kyra Wilson, Aylin Caliskan
arXiv:2605. 28969v2 Announce Type: replace-cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user.
By Aarik Gulaya