arXiv AI

Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents

The paper introduces Milestone Viability Potential Policy Optimization (MVPO), a reinforcement learning algorithm designed for long‑horizon large language model agents. MVPO addresses the zero‑credit failure problem by learning from viable failure prefixes, estimating prefix potential over Union‑Find viability regions, and repairing zero‑credit groups with potential‑difference advantages. Experiments on Qwen2.5‑1.5B‑Instruct demonstrate that MVPO outperforms eight strong baselines, improving success rates on ALFWorld and WebShop with minimal overhead.

arXiv AI
Sep 3

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

The paper introduces Potential-Guided Policy Optimization (PGPO), a method for multi-turn agentic tasks that improves credit assignment by estimating empirical state potentials from anchor-state-group return statistics. PGPO derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation and finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop demonstrate strong performance compared to recent group-based reinforcement learning methods, with negligible training overhead.

By Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou
Hugging Face Trending Papers
Sep 2

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks proposes a new reinforcement learning approach that estimates empirical state potentials from anchor-state-group return statistics within each rollout group. By deriving action advantages from potential differences between adjacent states, PGPO enables cross‑trajectory credit propagation, providing finer‑grained step‑level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop demonstrate strong overall performance compared to recent group‑based RL methods, with negligible training overhead.

Hugging Face Trending Papers
Jun 24

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps receive opposite credit depending on whether their trajectory eventually succeeded or failed.

arXiv AI
4d ago

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

arXiv:2609.36178v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment o...

By Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier