SLCA-GRPO addresses cross‑segment credit misattribution in tool‑calling reinforcement learning by introducing Segment‑Locked Credit Assignment (SLCA), which separates advantage estimation for tool‑invocation tokens and natural‑language summary tokens. The method leverages a Schema‑Guided LLM Simulator (SGLS) for scalable training and Hierarchical Rewards (HierR) to route execution and preference advantages appropriately. Experiments on a 7B backbone show that SLCA‑GRPO outperforms baseline methods, improving in‑domain accuracy by 2.53 pp, the Berkeley Function‑Calling Leaderboard by 1.36 pp, and $ au^2$‑Bench by 9.15 pp while reducing tool redundancy and costs.
By Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, Weizhou Pan, Jiahong Yan
arXiv:2605. 17877v2 Announce Type: replace Abstract: A significant hurdle for current LLMs is the execution of complex, multi-stage tasks.
By Wonjoong Kim, Yeonjun In, Sangwu Park, Dongha Lee, Chanyoung Park
arXiv:2601. 15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving.
By Tianshi Xu, Yuteng Chen, Meng Li
arXiv:2607. 03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals.
By Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li
arXiv:2607. 11172v1 Announce Type: new Abstract: Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage.
By Ke Xu, Han Xu, Xinran Chen, Yuqian Wang, Zhixuan Li, Xiaojian Liu, Changwo Wu, Jianqiang Xia, Yuchen Li
The paper introduces GRAFT, a Graph-based Faithful sTep-level credit-assignment framework that constructs a trajectory graph from rollout trajectories, recovers node state-values via Bellman iteration, and assigns step-level advantages based on node value differences. It also proposes Graph GAE to further reduce state-value estimation bias. Experiments on multi-turn agentic benchmarks demonstrate consistent improvements over GRPO and other recent agentic RL algorithms.
By Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang