Hugging Face Trending Papers

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

Read the original on Hugging Face Trending Papers →

Dr.Credit introduces a rubric‑grounded credit assignment method that evaluates intermediate tool turns in deep research agents by comparing the information returned to the history of accepted support for each rubric. Unlike traditional approaches that rely on ground‑truth answers, Dr.Credit uses task requirements as a shared reference, distinguishing new support from previously seen evidence and recognizing partial rubric fulfillment. Experiments on four benchmarks show that Dr.Credit outperforms open deep research baselines across all primary metrics, achieving performance competitive with proprietary models while enabling more efficient evidence acquisition and higher‑quality reports under limited turn budgets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 11

DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.

By Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao
arXiv Machine Learning
Jul 16

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

arXiv:2607. 13988v1 Announce Type: new Abstract: Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training.

By Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li