arXiv Machine Learning By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

Read the original on arXiv Machine Learning →

AdaStep introduces an adaptive step-credit weighting technique for agentic reinforcement learning, addressing the coarse granularity of trajectory-level objectives in long-horizon LLM agents. By formulating the weighting as a mean-squared-error estimation problem and deriving an optimal per-state shrinkage coefficient, AdaStep selectively preserves local credit when return variation is due to the chosen action and suppresses it when downstream randomness dominates. The method requires only lightweight scalar computations, no critic or extra rollouts, and demonstrates consistent performance gains across three model backbones on ALFWorld, WebShop, and ScienceWorld.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 14

Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

The paper introduces GACA, a critic‑free reinforcement learning estimator that adapts credit assignment granularity based on a step‑level uncertainty proxy. GACA assigns higher weight to fine‑grained signals for steps with above‑average negative log‑likelihood, while relying on episode‑level signals for less uncertain steps, improving task success on ALFWorld and WebShop for 1.5B and 7B language models. The authors provide a risk decomposition, a conditional bound on action‑value variation, and an error‑projection analysis to justify the method’s effectiveness.

By Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang, Yingzong Min, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong
arXiv AI
Sep 3

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

The paper introduces Potential-Guided Policy Optimization (PGPO), a method for multi-turn agentic tasks that improves credit assignment by estimating empirical state potentials from anchor-state-group return statistics. PGPO derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation and finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop demonstrate strong performance compared to recent group-based reinforcement learning methods, with negligible training overhead.

By Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou
arXiv AI
Sep 25

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

The paper introduces GRAFT, a Graph-based Faithful sTep-level credit-assignment framework that constructs a trajectory graph from rollout trajectories, recovers node state-values via Bellman iteration, and assigns step-level advantages based on node value differences. It also proposes Graph GAE to further reduce state-value estimation bias. Experiments on multi-turn agentic benchmarks demonstrate consistent improvements over GRPO and other recent agentic RL algorithms.

By Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang