Proximal Policy Optimization
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance.
Related stories
Transductive Off-policy Proximal Policy Optimization
arXiv:2406.03894v2 Announce Type: replace Abstract: Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due...
Reflective Policy Optimization
arXiv:2406.03678v2 Announce Type: replace Abstract: On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensi...
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
arXiv:2512. 06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting.
Lipschitz-Regularized Critics Lead to Policy Robustness Against Transition Dynamics Uncertainty
arXiv:2404. 13879v5 Announce Type: replace Abstract: Uncertainties in transition dynamics pose a critical challenge in reinforcement learning (RL), often resulting in performance degradation of trained policies when deployed on hardware.
OpenAI Baselines: DQN
We’re open-sourcing OpenAI Baselines, our internal effort to reproduce reinforcement learning algorithms with performance on par with published results. We’ll release the algorithms over upcoming months; today’s release includes DQN and three of its variants.
EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning
arXiv:2606. 26327v1 Announce Type: cross Abstract: In actor-critic reinforcement learning, network architectures are typically manually designed.
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.
Towards Optimal Policy Improvement
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
KLip-PPO: A per-sample KL perspective on PPO-Clip
arXiv:2606. 23932v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning.
Reinforcement learning to choose optimizers
The paper introduces a reinforcement learning framework that selects among a portfolio of gradient‑based and derivative‑free optimizers during a run. At each decision point a recurrent policy reads the current run state and chooses both the next optimizer and its usage duration, passing the best solution and step size forward. The method is trained with a decoupled actor‑critic using the same runtime distribution metric as evaluation, and on unseen problems it outperforms all individual portfolio optimizers except at the smallest budgets, remaining robust to distribution shift.
Learning Montezuma’s Revenge from a single demonstration
We’ve trained an agent to achieve a high score of 74,500 on Montezuma’s Revenge from a single human demonstration, better than any previously published result. Our algorithm is simple: the agent plays a sequence of games starting from carefully chosen states from the demonstration, and learns from them by optimizing the game score using PPO, the same reinforcement learning algorithm that underpins OpenAI Five.