arXiv AI

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

arXiv:2603. 14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions.

arXiv AI
Sep 15

DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents

DynSTEER is a dynamic stage‑wise trajectory evaluation framework designed for large language model agents performing long‑horizon tasks. It segments rollouts into stages anchored by key actions, uses a path‑tolerant milestone graph to accommodate diverse valid strategies, and adaptively routes queries to multi‑tier judges while halting unrecoverable executions early. Experiments show it improves evaluation discriminability by 85.2% over native methods, separates all model pairs with statistical significance, and saves 34.51% of execution steps on failed rollouts.

By Zhichao Shi, Wenjie Zhang, Xuhui Jiang, Xiaojun Wu, Cehao Yang, Chengjin Xu, Jian Guo, Yuanzhuo Wang
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
arXiv Computation and Language
Sep 1

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

CAST is a critique‑aware training framework that transforms sparse task outcomes into action‑level supervision for both critique learning and policy optimization. By analyzing agent trajectories, CAST synthesizes structured rationales that explain action validity under partial observability, enabling the creation of richer training data. Fine‑tuned Qwen3‑family models trained with CAST show significant reliability gains, outperforming GPT‑OSS‑120B by over 10% on Retail tasks and improving Telehealth performance by 9% in an out‑of‑domain setting.

By Amir Saeidi, Zehua Zhang, Rishitosh Singh, Naman Ahuja, Vivek Gupta, Ali Payani, Gaowen Liu, Jayanth Srinivasa, Chitta Baral
arXiv AI
Sep 25

IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.

By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung
arXiv AI
Aug 28

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.

By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu
arXiv AI
Aug 11

$A^2E$ : An End-to-End Agent Auditing Engine

arXiv:2608. 07346v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
arXiv Computation and Language
Sep 3

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

The paper introduces RuVerBench, a benchmark with 2,458 instances for evaluating the reliability of Large Language Models acting as judges (LaaJ) in verifying rubric compliance within agentic scenarios such as deep research and agentic coding. It systematically meta‑evaluates frontier LLMs, revealing that even the most advanced models perform well yet still produce substantial noise. The study also examines how prompt design, batching, and majority voting affect verification accuracy, noting that weaker models are more prompt‑sensitive, batched verification trades accuracy for efficiency, and majority voting offers diminishing returns.

By Yangda Peng, Yunjia Qi, Haotian Xia, Guanzhong He, Xintong Shi, Richeng Xuan, Songyuanyi Lu, Yixian Liu, Zhichao Hu, Yuhong Liu, Hao Peng