We’ve trained an agent to achieve a high score of 74,500 on Montezuma’s Revenge from a single human demonstration, better than any previously published result. Our algorithm is simple: the agent plays a sequence of games starting from carefully chosen states from the demonstration, and learns from them by optimizing the game score using PPO, the same reinforcement learning algorithm that underpins OpenAI Five.
arXiv:2503. 14833v2 Announce Type: replace-cross Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models.
By Zihao Liu, Xing Liu, Yuhang Dong, Haitao Chang, Zhengxiong Liu, Panfeng Huang
arXiv:2604. 18701v3 Announce Type: replace-cross Abstract: Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction error across all visited transitions.
By Vin Bhaskara, Haicheng Wang
arXiv:2503. 13077v2 Announce Type: replace Abstract: Multi-agent reinforcement learning has shown promise in learning cooperative behaviors in team-based environments.
By Amir Baghi, Jens Sj\"olund, Joakim Bergdahl, Linus Gissl\'en, Alessandro Sestini
arXiv:2607. 29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process.
By Bumgeun Park, Donghwan Lee