The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.
arXiv:2607. 29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process.
By Bumgeun Park, Donghwan Lee
The paper proposes Instant Episode Repetition (IER), a method that immediately repeats action sequences from high‑reward episodes during training to boost sample efficiency. Unlike passive replay techniques, IER actively shapes data collection by re‑executing successful behaviors for a set number of subsequent episodes. Experiments on MuJoCo, the DeepMind Control Suite, and a robotic manipulation task show that IER improves learning performance over standard SAC, TD3, and self‑imitation baselines.
By Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, Henry Williams
arXiv:2607. 17760v1 Announce Type: cross Abstract: Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations.
By Ziyi Liu, Grace Zhang
arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.
By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv:2609.36473v1 Announce Type: new
Abstract: Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target...
By Akhil Bagaria, Anita De Mello Koch, George Konidaris
The paper introduces Uncertainty-Driven Replay Memory (UDRM), a new experience replay buffer for reinforcement learning that prioritizes storing transitions with high uncertainty estimates. Unlike traditional buffers that rely on temporal difference error or transition distributions, UDRM updates its contents based on uncertainty derived from the RL model during training. Experiments show that this uncertainty-aware buffer leads to higher rewards during training compared to other uncertainty-aware RL frameworks.
By Sheeraja Rajakrishnan, Alexander G. Ororbia, Travis Desell, Daniel E. Krutz
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibit substantial natural variations (e.
arXiv:2606. 04492v1 Announce Type: new Abstract: Cooperative Multi-Agent Reinforcement Learning (MARL) frequently suffers from severe reward sparsity and exploration bottlenecks.
By Zicheng Zhao, Yu Lan, Chengzhengxu Li, Zhaohan Zhang, Xiaoming Liu
arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.
By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
Efficient Diversity-based Experience Replay (EDER) is introduced to enhance learning efficiency in deep reinforcement learning. It uses a determinantal point process to model sample diversity, prioritizes replay based on this diversity, and incorporates Cholesky decomposition and rejection sampling to handle large state spaces. Experiments on MuJoCo, Atari, and Habitat show that EDER significantly improves learning efficiency and performance in high‑dimensional, realistic environments.
By Kaiyan Zhao, Yiming Wang, Yuyang Chen, Yan Li, Leong Hou U, Xiaoguang Niu