arXiv:2609.39634v1 Announce Type: cross
Abstract: Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribu...
By Nima H. Siboni
arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
The paper introduces Anchored Neighborhood Optimization (ANO), a new policy‑optimization method that directly designs a smooth, bounded gain field for the probability‑ratio surrogate objective. ANO anchors the identity map at a ratio of one, peaks at a specified trust‑region boundary, and limits the influence of extreme off‑policy samples while providing a bounded, redescending pull on outliers. Empirical results show ANO consistently outperforms existing methods on Atari and MuJoCo benchmarks, and it remains robust under aggressive learning‑rate settings.
By Yiheng Zhang, Yiming Wang, Kaiyan Zhao, Zhenglin Wan, Jiayu Chen, Leong Hou U
arXiv:2606. 23932v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning.
By Riccardo Colletti, Robin Holzinger
arXiv:2609.36802v1 Announce Type: new
Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning...
By Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez
The paper investigates how reusing past samples can improve the sample efficiency of Proximal Policy Optimization (PPO). Two variants, wPPO-U and wPPO-BH, are introduced within a multiple importance weighting framework, each reusing data from recent iterations while preserving core PPO mechanics. The authors derive theoretical policy improvement bounds for both variants and empirically evaluate their impact on continuous control tasks.
By Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli