Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is...
The paper introduces NAQD‑Env, a synthetic benchmark designed to test language agents’ ability to selectively withdraw and resume actions when new evidence, permissions, or stop instructions arise. It evaluates models against a deterministic reference policy across eleven dependency families, measuring policy agreement, task value, withdrawal, resumption, and event reporting. Experiments on 350 scenarios show low withdrawal recall, no valid resumption, and limited policy alignment, highlighting the need to treat selective withdrawal as a distinct reliability component.
By Mohamed Abouzahra
The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.
By Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig
The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.
By Jiapeng Li
arXiv:2609.36138v1 Announce Type: new
Abstract: Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or...
By Jiayi Li, Ruizhe Li
arXiv:2609.37472v1 Announce Type: cross
Abstract: Behavioral tests measure how a language model reads evidence. We ask whether those measurements help choose a recommendation interface. We evaluate s...
By Han Chen, Yingrui Li
arXiv:2607. 28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect.
By Xiaonan Xu, Wenjing Wu
arXiv:2608. 00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.
By William Caban
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
The paper presents a layered framework for evaluating conversational AI by aligning offline proxy signals with online A/B experiment outcomes. It introduces a three‑step alignment chain—behavioral label to product outcome, classifier to candidate behavior, and offline signal to experiment effect—alongside an audit protocol that compares confidence intervals and rankings. In a real‑world deployment, the composite proxy achieved 81.1% F1 versus 34.3% for the raw classifier, correctly predicting direction on all 113 contrasts and enabling efficient prioritization of candidate models before costly online testing.
By Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng
arXiv:2607. 12200v1 Announce Type: new Abstract: As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone.
By Rahul Gupta, Abhinav Mohanty, Payal Motwani, Venkatesh Saligrama, Satyapriya Krishna, Connor Harris, Gary Anthony Ackerman, Brandon Behlendorf, Tom Hobson, Theodore Wilson, Spyros Matsoukas
arXiv:2608. 13867v1 Announce Type: cross Abstract: AI coding agents are commonly evaluated as models but deployed as systems.
By Stephanie Jarmak