An Introduction to Deep Reinforcement Learning
Related stories
Stochastic Neural Networks for hierarchical reinforcement learning
An Introduction to Q-Learning Part 1
Introducing ⚔️ AI vs. AI ⚔️ a deep reinforcement learning multi-agents competition system
An Introduction to Q-Learning Part 2/2
Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning
arXiv:2606. 20411v1 Announce Type: new Abstract: Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms.
Quantifying generalization in reinforcement learning
We’re releasing CoinRun, a training environment which provides a metric for an agent’s ability to transfer its experience to novel situations and has already helped clarify a longstanding puzzle in reinforcement learning. CoinRun strikes a desirable balance in complexity: the environment is simpler than traditional platformer games like Sonic the Hedgehog but still poses a worthy generalization challenge for state of the art algorithms.
Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System
arXiv:2607. 03125v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its "black-box" exploration risks violating strict hardware safety limits.
Benchmarking safe exploration in deep reinforcement learning
RL²: Fast reinforcement learning via slow reinforcement learning
Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks
arXiv:2605. 22305v2 Announce Type: replace Abstract: We analytically solve the Mountain Car problem, a canonical benchmark in RL, and derive an optimal control solution, closing a gap after 36 years.
Pareto Q-Learning with Reward Machines
arXiv:2606. 19134v1 Announce Type: cross Abstract: We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward structure is specified by a set of reward machines (RMs).