arXiv AI

Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

arXiv:2606. 09748v1 Announce Type: new Abstract: Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their reports when guided by feedback?

arXiv Computation and Language
Sep 11

DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.

By Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao
Hugging Face Trending Papers
Sep 28

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

Dr.Credit introduces a rubric‑grounded credit assignment method that evaluates intermediate tool turns in deep research agents by comparing the information returned to the history of accepted support for each rubric. Unlike traditional approaches that rely on ground‑truth answers, Dr.Credit uses task requirements as a shared reference, distinguishing new support from previously seen evidence and recognizing partial rubric fulfillment. Experiments on four benchmarks show that Dr.Credit outperforms open deep research baselines across all primary metrics, achieving performance competitive with proprietary models while enabling more efficient evidence acquisition and higher‑quality reports under limited turn budgets.

arXiv Computation and Language
Sep 3

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

The paper introduces RuVerBench, a benchmark with 2,458 instances for evaluating the reliability of Large Language Models acting as judges (LaaJ) in verifying rubric compliance within agentic scenarios such as deep research and agentic coding. It systematically meta‑evaluates frontier LLMs, revealing that even the most advanced models perform well yet still produce substantial noise. The study also examines how prompt design, batching, and majority voting affect verification accuracy, noting that weaker models are more prompt‑sensitive, batched verification trades accuracy for efficiency, and majority voting offers diminishing returns.

By Yangda Peng, Yunjia Qi, Haotian Xia, Guanzhong He, Xintong Shi, Richeng Xuan, Songyuanyi Lu, Yixian Liu, Zhichao Hu, Yuhong Liu, Hao Peng
arXiv AI
Sep 25

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

The paper introduces AgentX-Model, a dual‑agent framework that links proposal development with model experimentation in industrial recommender systems. The Research Agent drafts proposals from literature and prior findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework iteratively selects starting implementations and formulates new research questions, organizing work into Reproduce, Follow‑up, Composition, and Diagnose actions. Across production evaluations, most experiments exceeded business baselines, with recent A/B tests showing significant gains in acquisition efficiency, advertising spend, and watch time while reducing computational cost.

By Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou
arXiv AI
Jun 30

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

arXiv:2606. 28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and substantial expert effort.

By Yen-Hsun Huang (Department of Education, Taipei Veterans General Hospital, Taipei, Taiwan), Yu-Shiou Lin (Department of Psychiatry, Taipei Veterans General Hospital, Taipei, Taiwan)
arXiv AI
Oct 1

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

DAGent introduces an Evaluate‑then‑Grow planning approach for deep research agents, building directed acyclic graphs incrementally based on confidence and uncertainty from completed tasks. The framework includes a hierarchical context layer for efficient query handling and a structural reinforcement learning component, DAGRPO, that rewards topology‑conditioned execution. Experiments on BrowseComp‑Plus, GAIA, and xbench‑DeepSearch show DAGent outperforming strong baselines across multiple backbones and scaling to large language models.

By Hanwen Liu, Yuanfu Sun, Qiaoyu Tan