arXiv AI By Jian Gao, Hang Jiang

When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e Screening

Read the original on arXiv AI →

The study examines a two‑agent résumé screening process where both employer‑side and candidate‑side agents exchange evidence before deciding who advances, contrasting it with the traditional one‑call automated screening. Using GPT‑5.5 and Claude Opus 4.7 on 600 constructed résumé‑job pairs, the two‑agent method increased the proportion of applications advanced (up to 39.3% for GPT‑5.5) and raised pass rates for borderline cases from 4.5% to 26.2% (GPT‑5.5) and 6.5% to 16.1% (Opus 4.7). The results show that the screening procedure itself, rather than just the underlying model, determines which candidates reach human review and how consistently that access recurs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Jun 11

Search Discipline for Long-Horizon Research Agents

arXiv:2606. 11522v1 Announce Type: new Abstract: Autoresearch agents now propose, evaluate, and select scientific candidates against a metric, and that metric is usually an aggregate reduced over a heterogeneous space of regions, slices, or cohorts.

By Adithya Srinivasan, Devesh Paragiri
arXiv AI
3d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu