arXiv AI

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents

arXiv:2608. 06861v1 Announce Type: new Abstract: Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions.

arXiv AI
5d ago

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

arXiv:2609.36178v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment o...

By Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier
arXiv Machine Learning
Sep 14

Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

The paper introduces GACA, a critic‑free reinforcement learning estimator that adapts credit assignment granularity based on a step‑level uncertainty proxy. GACA assigns higher weight to fine‑grained signals for steps with above‑average negative log‑likelihood, while relying on episode‑level signals for less uncertain steps, improving task success on ALFWorld and WebShop for 1.5B and 7B language models. The authors provide a risk decomposition, a conditional bound on action‑value variation, and an error‑projection analysis to justify the method’s effectiveness.

By Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang, Yingzong Min, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong
arXiv AI
Sep 25

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

The paper introduces GRAFT, a Graph-based Faithful sTep-level credit-assignment framework that constructs a trajectory graph from rollout trajectories, recovers node state-values via Bellman iteration, and assigns step-level advantages based on node value differences. It also proposes Graph GAE to further reduce state-value estimation bias. Experiments on multi-turn agentic benchmarks demonstrate consistent improvements over GRPO and other recent agentic RL algorithms.

By Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang
arXiv AI
Sep 3

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

The paper introduces Potential-Guided Policy Optimization (PGPO), a method for multi-turn agentic tasks that improves credit assignment by estimating empirical state potentials from anchor-state-group return statistics. PGPO derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation and finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop demonstrate strong performance compared to recent group-based reinforcement learning methods, with negligible training overhead.

By Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou
arXiv Machine Learning
Sep 4

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

TIGPO (Temporal Instance-Graph Policy Optimization) extends graph-based credit assignment for long-horizon LLM agents by maintaining a persistent transition graph per task across policy updates. It allocates rollout budgets to both new exploration and revisiting past tasks, pairing current rollouts with earlier ones to create cross‑temporal references that stabilize advantage estimation. Experiments on ALFWorld and WebShop show TIGPO consistently outperforms previous group‑based and graph‑based policy optimization methods.

By Jinwei Gan