Dr.Credit introduces a rubric‑grounded credit assignment method that evaluates intermediate tool turns in deep research agents by comparing the information returned to the history of accepted support for each rubric. Unlike traditional approaches that rely on ground‑truth answers, Dr.Credit uses task requirements as a shared reference, distinguishing new support from previously seen evidence and recognizing partial rubric fulfillment. Experiments on four benchmarks show that Dr.Credit outperforms open deep research baselines across all primary metrics, achieving performance competitive with proprietary models while enabling more efficient evidence acquisition and higher‑quality reports under limited turn budgets.
The paper introduces VICT, a method that leverages the internal structure of verifiable tasks to perform fine‑grained credit assignment for long‑horizon LLM agents. VICT exposes executable or evidence‑backed atoms from a task’s terminal verifier and traces them back to actions via dependency‑valid proof edges, redistributing advantage only along these edges. This approach improves performance on ALFWorld and WebShop compared to outcome‑only training and matches recent fine‑grained credit methods without requiring additional critics, labels, or inference‑time verifier access.
By Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma
arXiv:2608. 05102v1 Announce Type: new Abstract: Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer.
By Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen
arXiv:2606. 05784v1 Announce Type: new Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform broadcast of trajectory-level advantages to all tokens causes valuable tool-use steps in failing trajectories to be penalized no differently from valueless ones.
By Chengqi Dong, Chuhuai Yue, Hang He, yandong liu, Fenghe Tang, S Kevin Zhou, Xiaohan Wang, Jiajun Chai, Guojun Yin
The paper introduces ICA, an evidence‑centric framework that represents information from web‑tool interactions as stable, rendered snapshots, enabling comparison across trajectories. It proposes Information‑Aware Credit Assignment, a post‑hoc reward propagation technique that estimates turn‑level utility from rollout success and assigns dense rewards to steps that provide high‑utility information. When combined with GSPO, ICA consistently improves performance on several web‑search benchmarks such as BrowseComp, GAIA, Xbench‑DS, and Seal‑0.
By Cong Pang, Xuyu Feng, Yujie Yi, Jiaqi Su, Zixuan Chen, Jiawei Hong, Tiankuo Yao, Nang Yuan, Jiapeng Luo, Lewei Lu, Xin Lou
arXiv:2606. 32017v1 Announce Type: cross Abstract: Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions.
By Yuanda Xu, Zhengze Zhou, Hejian Sang, Xiaomin Li, Jiaxin Zhang, Xinchen Du, Zhipeng Wang, Alborz Geramifard