arXiv:2609.39634v1 Announce Type: cross
Abstract: Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribu...
By Nima H. Siboni
arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
The paper introduces Anchored Neighborhood Optimization (ANO), a new policy‑optimization method that directly designs a smooth, bounded gain field for the probability‑ratio surrogate objective. ANO anchors the identity map at a ratio of one, peaks at a specified trust‑region boundary, and limits the influence of extreme off‑policy samples while providing a bounded, redescending pull on outliers. Empirical results show ANO consistently outperforms existing methods on Atari and MuJoCo benchmarks, and it remains robust under aggressive learning‑rate settings.
By Yiheng Zhang, Yiming Wang, Kaiyan Zhao, Zhenglin Wan, Jiayu Chen, Leong Hou U
arXiv:2606. 23932v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning.
By Riccardo Colletti, Robin Holzinger
arXiv:2609.36802v1 Announce Type: new
Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning...
By Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez
The paper investigates how reusing past samples can improve the sample efficiency of Proximal Policy Optimization (PPO). Two variants, wPPO-U and wPPO-BH, are introduced within a multiple importance weighting framework, each reusing data from recent iterations while preserving core PPO mechanics. The authors derive theoretical policy improvement bounds for both variants and empirically evaluate their impact on continuous control tasks.
By Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli
arXiv:2505. 15201v5 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) algorithms sample multiple n>1 solution attempts for each problem and reward them independently.
By Christian Walder, Deep Karkhanis
arXiv:2610.02198v1 Announce Type: cross
Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic....
By Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
arXiv:2602. 04879v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorithm.
By Penghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang, Chao Du, Min Lin, Wee Sun Lee
Group Adaptive Clipping Policy Optimization (GAPO) is a plug‑in modification to GRPO methods that adapts the importance‑sampling clipping boundary based on rollout advantage. By allowing rollouts with larger learning signals to receive proportionally greater update headroom, GAPO addresses the limitation of fixed clipping that suppresses rare but informative rollouts. Experiments on Qwen and Llama models show that GAPO consistently improves Pass@1 and Pass@k on math reasoning and coding benchmarks where base model pass rates are low.
By Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan, Rein Houthooft
BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.
By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen