We’ve trained an agent to achieve a high score of 74,500 on Montezuma’s Revenge from a single human demonstration, better than any previously published result. Our algorithm is simple: the agent plays a sequence of games starting from carefully chosen states from the demonstration, and learns from them by optimizing the game score using PPO, the same reinforcement learning algorithm that underpins OpenAI Five.
arXiv:2503. 14833v2 Announce Type: replace-cross Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models.
By Zihao Liu, Xing Liu, Yuhang Dong, Haitao Chang, Zhengxiong Liu, Panfeng Huang
arXiv:2604. 18701v3 Announce Type: replace-cross Abstract: Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction error across all visited transitions.
By Vin Bhaskara, Haicheng Wang
arXiv:2503. 13077v2 Announce Type: replace Abstract: Multi-agent reinforcement learning has shown promise in learning cooperative behaviors in team-based environments.
By Amir Baghi, Jens Sj\"olund, Joakim Bergdahl, Linus Gissl\'en, Alessandro Sestini
arXiv:2609.07575v1 Announce Type: cross
Abstract: This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic...
By Mikel Malag\'on, Jon Vadillo, Josu Ceberio, Michael Bowling, Jose A. Lozano
arXiv:2607. 29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process.
By Bumgeun Park, Donghwan Lee
The paper introduces Gradient‑Momentum Coupling (GMC), a method that quantifies learning progress by measuring how strongly a sample influences changes in the parameter space, using the normalized absolute product of its gradient and the momentum of previous gradients. GMC filters out noise by accumulating consistent directions of change while canceling random fluctuations, leading to a more uniform prioritization across tasks with varying noise levels and better ranking of learnable tasks by improvement speed. Experiments on MiniGrid MultiRoom tasks show that replacing prediction error with GMC in the Intrinsic Curiosity Module restores exploration capabilities that were lost to unpredictable observations.
By Samuel Blad, Martin L\"angkvist, Amy Loutfi
We’re releasing an experimental metalearning approach called Evolved Policy Gradients, a method that evolves the loss function of learning agents, which can enable fast training on novel tasks. Agents trained with EPG can succeed at basic tasks at test time that were outside their training regime, like learning to navigate to an object on a different side of the room from where it was placed during training.
The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.
By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
We’re launching a transfer learning contest that measures a reinforcement learning algorithm’s ability to generalize from previous experience.
We’ve found that adding adaptive noise to the parameters of reinforcement learning algorithms frequently boosts performance. This exploration method is simple to implement and very rarely decreases performance, so it’s worth trying on any problem.