arXiv AI

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper proposes Novelty and Surprise Prioritized Experience Replay (NSPER) for image-based reinforcement learning, combining novelty to highlight underrepresented states and surprise to reveal gaps in the agent’s knowledge. An extended version, NSPER+R, also uses these signals as intrinsic rewards to enhance both replay quality and exploration. Experiments on DeepMind Control Suite tasks demonstrate that NSPER and NSPER+R accelerate training and improve convergence compared to existing methods.

Hugging Face Trending Papers
Aug 18

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.

arXiv Machine Learning
Aug 19

Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning

The paper proposes Instant Episode Repetition (IER), a method that immediately repeats action sequences from high‑reward episodes during training to boost sample efficiency. Unlike passive replay techniques, IER actively shapes data collection by re‑executing successful behaviors for a set number of subsequent episodes. Experiments on MuJoCo, the DeepMind Control Suite, and a robotic manipulation task show that IER improves learning performance over standard SAC, TD3, and self‑imitation baselines.

By Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, Henry Williams
arXiv AI
Jun 2

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.

By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv Machine Learning
Sep 1

Uncertainty-Driven Replay Memory for Reinforcement Learning

The paper introduces Uncertainty-Driven Replay Memory (UDRM), a new experience replay buffer for reinforcement learning that prioritizes storing transitions with high uncertainty estimates. Unlike traditional buffers that rely on temporal difference error or transition distributions, UDRM updates its contents based on uncertainty derived from the RL model during training. Experiments show that this uncertainty-aware buffer leads to higher rewards during training compared to other uncertainty-aware RL frameworks.

By Sheeraja Rajakrishnan, Alexander G. Ororbia, Travis Desell, Daniel E. Krutz
arXiv Machine Learning
Jul 3

Optimizing Visual Generative Models via Distribution-wise Rewards

arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.

By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv AI
Sep 11

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

Efficient Diversity-based Experience Replay (EDER) is introduced to enhance learning efficiency in deep reinforcement learning. It uses a determinantal point process to model sample diversity, prioritizes replay based on this diversity, and incorporates Cholesky decomposition and rejection sampling to handle large state spaces. Experiments on MuJoCo, Atari, and Habitat show that EDER significantly improves learning efficiency and performance in high‑dimensional, realistic environments.

By Kaiyan Zhao, Yiming Wang, Yuyang Chen, Yan Li, Leong Hou U, Xiaoguang Niu