arXiv Machine Learning By Jinwei Gan

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

Read the original on arXiv Machine Learning →

TIGPO (Temporal Instance-Graph Policy Optimization) extends graph-based credit assignment for long-horizon LLM agents by maintaining a persistent transition graph per task across policy updates. It allocates rollout budgets to both new exploration and revisiting past tasks, pairing current rollouts with earlier ones to create cross‑temporal references that stabilize advantage estimation. Experiments on ALFWorld and WebShop show TIGPO consistently outperforms previous group‑based and graph‑based policy optimization methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

Tracking the Moving Frontier: Long-Short Term Advantage Estimator

The paper introduces Long-Short Term Advantage Estimator (LSTAE), a single‑stream reinforcement learning algorithm that replaces repeated within‑iteration trajectory sampling with historical experience for advantage estimation. LSTAE tracks each task anchor with a drift‑aware long‑term baseline and a short‑term state‑experience buffer, converting accumulated experience into multi‑granular credit signals while requiring only one rollout per anchor. Experiments on agentic and mathematical reasoning benchmarks show that LSTAE matches or surpasses strong group‑based baselines while significantly reducing rollout cost.

By Xinhao Yao, Lu Yu, Changhao Wang, Fengwei Teng, Yuyao Zhang, Qing Cui, Jun Zhou, Yong Liu
arXiv AI
Jun 2

Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

arXiv:2605. 26684v2 Announce Type: replace-cross Abstract: Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agentic tasks.

By Xin Cheng, Shuo He, Lang Feng, HaiYang Xu, Ming Yan, Lei Feng, Bo An
Hugging Face Trending Papers
Aug 20

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks.