arXiv AI By Chengqi Dong, Chuhuai Yue, Hang He, yandong liu, Fenghe Tang, S Kevin Zhou, Xiaohan Wang, Jiajun Chai, Guojun Yin

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

Read the original on arXiv AI →

arXiv:2606. 05784v1 Announce Type: new Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform broadcast of trajectory-level advantages to all tokens causes valuable tool-use steps in failing trajectories to be penalized no differently from valueless ones.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

SLCA-GRPO addresses cross‑segment credit misattribution in tool‑calling reinforcement learning by introducing Segment‑Locked Credit Assignment (SLCA), which separates advantage estimation for tool‑invocation tokens and natural‑language summary tokens. The method leverages a Schema‑Guided LLM Simulator (SGLS) for scalable training and Hierarchical Rewards (HierR) to route execution and preference advantages appropriately. Experiments on a 7B backbone show that SLCA‑GRPO outperforms baseline methods, improving in‑domain accuracy by 2.53 pp, the Berkeley Function‑Calling Leaderboard by 1.36 pp, and $ au^2$‑Bench by 9.15 pp while reducing tool redundancy and costs.

By Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, Weizhou Pan, Jiahong Yan
arXiv AI
Jul 7

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

arXiv:2607. 03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals.

By Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li
arXiv AI
Sep 25

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

The paper introduces GRAFT, a Graph-based Faithful sTep-level credit-assignment framework that constructs a trajectory graph from rollout trajectories, recovers node state-values via Bellman iteration, and assigns step-level advantages based on node value differences. It also proposes Graph GAE to further reduce state-value estimation bias. Experiments on multi-turn agentic benchmarks demonstrate consistent improvements over GRPO and other recent agentic RL algorithms.

By Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang