arXiv Computation and Language

Building Interpretable Feature Representations for Resume-Vacancy Matching by Distilling Production LLM Signals

The paper presents a system for matching job candidates to vacancies that provides interpretable evidence rather than a single relevance score. It uses a two‑stage approach: an LLM‑based labeler refined through recruiter feedback and a distilled bi‑encoder that runs online on CPU. The model, trained on 168,772 labeled pairs, achieves 95.79% agreement with recruiter‑recorded decisions on a production‑feedback subset.

arXiv AI
Sep 7

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

The paper reviews the evolution of AI recruitment systems from simple profile matching to complex, multi‑stage workflows that retrieve evidence, compare candidates, and execute actions. It analyzes 40 representative works, highlighting transitions from similarity to reciprocal suitability, from single models to compound workflows, and from offline predictions to evidence‑aligned evaluation. The authors identify persistent gaps—such as confounded behavioral labels, limited data validity, and lack of privacy assessment—and propose a staged mapping for defensible evaluation and an agenda for auditable, evidence‑grounded systems.

By Ziyi Zhao, Guanzheng Wei
Hugging Face Trending Papers
Sep 3

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

The paper reviews the evolution of AI recruitment systems from simple profile matching to complex, multi‑stage workflows that retrieve evidence, compare candidates, and execute actions. It organizes 40 representative works, highlighting transitions from similarity to reciprocal suitability, from single models to compound workflows, and from offline prediction to evidence‑aligned evaluation. The authors identify persistent gaps such as confounded behavioral labels, limited data validity, hidden pipeline failures, and a lack of privacy assessment, and propose a staged mapping for defensible evaluation and an agenda for evidence‑grounded, auditable systems.

arXiv Computation and Language
Aug 27

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

The paper introduces a scalable product‑linking system that uses a retrieve‑then‑match cascade. First, a lightweight text cross‑encoder auto‑resolves the majority of merchant‑catalog product pairs with high precision, while an agentic multimodal vision‑language model handles the remaining ambiguous cases by inspecting images and performing web searches. This approach balances computational cost and accuracy, improving overall link coverage from 68% to 77% without requiring fine‑tuning of the agent.

By Jian Wang, Steven Xu, Sanjyot Thete, Maryam Barouti, Tom Tang, Elaine Wu, Charu Sareen, Kyle MacDonald
arXiv AI
Aug 28

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv Computation and Language
Aug 31

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

The paper introduces LongJudgeBench, a benchmark designed to evaluate large language models (LLMs) acting as judges for long-form text generation. It highlights that long-form evaluation requires complex, document-level assessments beyond simple length, such as organization, coverage, depth, consistency, and scenario-specific quality. Experiments show a significant reliability gap among current LLM judges, indicating instability across scenarios and limited effectiveness of rubrics or references.

By Junjie Chen, Yuxi Dong, Haitao Li, Weihang Su, Yujia Zhou, Min Zhang, Yiqun Liu, Qingyao Ai