arXiv Computation and Language By Hiroko Takano

Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints

Read the original on arXiv Computation and Language →

The study evaluates how to reduce fabricated claims in multi‑stage large language model (LLM) hiring pipelines. Prompt guardrails alone cut fabrication density by 86 % but still left half of outputs containing false claims, while adding a human‑in‑the‑loop checkpoint after resume improvement eliminated all identity fabrications and significantly lowered overall fabrication rates. The results show that a layered approach—combining prompt guardrails with human checkpoints—provides stronger protection against severe failures without harming the quality of the final outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
6d ago

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.

By Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig
arXiv AI
Aug 11

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.

By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
arXiv AI
6d ago

Counterfactual Bias Testing for Application Tracking System

The paper proposes a scalable, automated method for auditing candidate‑job matching systems for demographic bias. It employs large‑language‑model agents to generate neutral resumes, injects controlled demographic variations, ranks candidates with a fine‑tuned embedding model, and evaluates nine fairness metrics across counterfactual, group‑fairness, and merit‑aware families, producing a composite risk report. Experiments on a small corpus show that single‑score audits miss nuanced issues, underscoring the need for multi‑metric evaluation and LLM‑generated audits as a low‑cost complement to human reviews.

By Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
arXiv AI
23h ago

Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems

The paper introduces SCOPED‑Hiring, a process‑aware fairness diagnosis pipeline for large language model (LLM) based multi‑agent hiring systems. It generates controlled resume variants, runs role‑based hiring committees, and logs over 311,000 structured decision trajectories, converting them into quantitative fairness signals across six diagnostic lenses: final outcome, counterfactual, process, pathway, dynamic, and design effects. The study finds that balanced final hire rates can conceal hidden trajectory unfairness—such as career gaps, proxy cues, and identity cues—and demonstrates that targeted repairs guided by these diagnoses can reduce the total layered burden by 72.3% while only slightly altering the hire rate.

By Yiran Zhao, Lu Zhou, Liming Fang, Yufei Chen, Jiafei Wu, Zhe Liu, Xiaogang Xu