← Back to all news
OpenAI Blog November 9, 2016

RL²: Fast reinforcement learning via slow reinforcement learning

Read the original on OpenAI Blog →

The Flow has not summarised this story yet — read it at OpenAI Blog.

  • reinforcement-learning

Related stories

arXiv Machine Learning
Jun 2

Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks

arXiv:2605. 22305v2 Announce Type: replace Abstract: We analytically solve the Mountain Car problem, a canonical benchmark in RL, and derive an optimal control solution, closing a gap after 36 years.

By Stefan Huber, Hannes Unger, Georg Sch\"afer, Jakob Rehrl
agentsreinforcement-learningbenchmarks
More like this →
OpenAI Blog
Dec 13, 2019

Dota 2 with large scale deep reinforcement learning

reinforcement-learning
More like this →
Hugging Face Trending Papers
Jun 25

Heavy-Ball Q-Learning with Residual Weighting Correction

This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning.

reinforcement-learning
More like this →
arXiv AI
Jun 26

Heavy-Ball Q-Learning with Residual Weighting Correction

arXiv:2606. 27112v1 Announce Type: cross Abstract: This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence.

By Donghwan Lee
reinforcement-learning
More like this →
arXiv AI
Jul 21

From Classical to Quantum Reinforcement Learning and Its Applications in Quantum Control: A Beginner's Tutorial

arXiv:2601. 08662v3 Announce Type: replace Abstract: This tutorial is designed to make reinforcement learning (RL) more accessible to undergraduate students by offering clear, example-driven explanations.

By Abhijit Sen, Sonali Panda, Mahima Arya, Subhajit Patra, Zizhan Zheng, Denys I. Bondar
reinforcement-learning
More like this →
arXiv Machine Learning
Jun 17

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach

arXiv:2510. 19528v2 Announce Type: replace-cross Abstract: We investigate the fundamental problem of leveraging offline data to accelerate online reinforcement learning - a direction with strong potential but limited theoretical grounding.

By Sebastian Reboul, H\'el\`ene Halconruy
reinforcement-learningfine-tuning
More like this →