arXiv AI By Ziyi Zhao, Guanzheng Wei

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

Read the original on arXiv AI →

The paper reviews the evolution of AI recruitment systems from simple profile matching to complex, multi‑stage workflows that retrieve evidence, compare candidates, and execute actions. It analyzes 40 representative works, highlighting transitions from similarity to reciprocal suitability, from single models to compound workflows, and from offline predictions to evidence‑aligned evaluation. The authors identify persistent gaps—such as confounded behavioral labels, limited data validity, and lack of privacy assessment—and propose a staged mapping for defensible evaluation and an agenda for auditable, evidence‑grounded systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 3

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

The paper reviews the evolution of AI recruitment systems from simple profile matching to complex, multi‑stage workflows that retrieve evidence, compare candidates, and execute actions. It organizes 40 representative works, highlighting transitions from similarity to reciprocal suitability, from single models to compound workflows, and from offline prediction to evidence‑aligned evaluation. The authors identify persistent gaps such as confounded behavioral labels, limited data validity, hidden pipeline failures, and a lack of privacy assessment, and propose a staged mapping for defensible evaluation and an agenda for evidence‑grounded, auditable systems.

arXiv AI
Aug 28

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv AI
Sep 17

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.

By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv Computation and Language
4d ago

Building Interpretable Feature Representations for Resume-Vacancy Matching by Distilling Production LLM Signals

The paper presents a system for matching job candidates to vacancies that provides interpretable evidence rather than a single relevance score. It uses a two‑stage approach: an LLM‑based labeler refined through recruiter feedback and a distilled bi‑encoder that runs online on CPU. The model, trained on 168,772 labeled pairs, achieves 95.79% agreement with recruiter‑recorded decisions on a production‑feedback subset.

By Ilya Chekin (BroutonLab), Vyacheslav Malyugin (BroutonLab), Vladimir Chirkov (BroutonLab), Mikhail Yurushkin (Curately)