arXiv AI By Xucong Wang, Ziyu Ma, Yong Wang, Yuxiang Ji, Shidong Yang, Guanhua Chen, Pengkun Wang, Xiangxiang Chu

APPO: Agentic Procedural Policy Optimization

Read the original on arXiv AI →

arXiv:2606. 12384v1 Announce Type: cross Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Contrastive Branch Policy Optimization

Contrastive Branch Policy Optimization (CBPO) is a reinforcement learning method that separates the allocation of a fixed rollout budget from the translation of branch outcomes into token-level credit. It uses generation entropy to screen branch positions, path- and node-level decay to distribute the budget, and Contrastive Branch Value (CBV) to estimate local decision sensitivity without changing reward signs. CBPO partitions trajectories into non-overlapping credit segments, preventing duplicated gradients and enabling fine-grained credit assignment using only outcome rewards.

By Ying Wang, Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun, Jingli Yang
arXiv AI
3d ago

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

arXiv:2609.36178v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment o...

By Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier
arXiv Machine Learning
Aug 26

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

The paper introduces Influence-Aware Policy Optimization (IAPO), a method that models multi‑turn agent rollouts as typed influence‑dependency graphs to better assign credit to actions based on how information and errors flow through user and tool interactions. IAPO transforms the structure of support and failure usage into routing weights that redistribute trajectory‑level advantage, enabling more effective learning from sparse final rewards. Experiments with Qwen3‑4B and Qwen3‑8B on three service‑agent benchmarks show that IAPO outperforms existing multi‑turn reinforcement learning baselines without harming function‑calling performance.

By Bo Ren, Yirong Mao, Yi Yang, Wenhui Que
arXiv AI
Jun 10

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.

By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
arXiv AI
Jul 7

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy

arXiv:2607. 03126v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult.

By Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Zeyu Chen, Quan Chen, Yanhua Cheng, Peng Jiang, Yadong Mu