arXiv AI By Cong Pang, Xuyu Feng, Yujie Yi, Jiaqi Su, Zixuan Chen, Jiawei Hong, Tiankuo Yao, Nang Yuan, Jiapeng Luo, Lewei Lu, Xin Lou

ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents

Read the original on arXiv AI →

The paper introduces ICA, an evidence‑centric framework that represents information from web‑tool interactions as stable, rendered snapshots, enabling comparison across trajectories. It proposes Information‑Aware Credit Assignment, a post‑hoc reward propagation technique that estimates turn‑level utility from rollout success and assigns dense rewards to steps that provide high‑utility information. When combined with GSPO, ICA consistently improves performance on several web‑search benchmarks such as BrowseComp, GAIA, Xbench‑DS, and Seal‑0.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 16

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

arXiv:2607. 13988v1 Announce Type: new Abstract: Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training.

By Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li