Hugging Face Trending Papers

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.

arXiv AI
Aug 19

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper proposes Novelty and Surprise Prioritized Experience Replay (NSPER) for image-based reinforcement learning, combining novelty to highlight underrepresented states and surprise to reveal gaps in the agent’s knowledge. An extended version, NSPER+R, also uses these signals as intrinsic rewards to enhance both replay quality and exploration. Experiments on DeepMind Control Suite tasks demonstrate that NSPER and NSPER+R accelerate training and improve convergence compared to existing methods.

By Hoda Yamani, Henry Williams, Bruce A. MacDonald
arXiv Machine Learning
Aug 19

Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning

The paper proposes Instant Episode Repetition (IER), a method that immediately repeats action sequences from high‑reward episodes during training to boost sample efficiency. Unlike passive replay techniques, IER actively shapes data collection by re‑executing successful behaviors for a set number of subsequent episodes. Experiments on MuJoCo, the DeepMind Control Suite, and a robotic manipulation task show that IER improves learning performance over standard SAC, TD3, and self‑imitation baselines.

By Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, Henry Williams
arXiv AI
Jun 2

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.

By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv Machine Learning
Sep 1

Uncertainty-Driven Replay Memory for Reinforcement Learning

The paper introduces Uncertainty-Driven Replay Memory (UDRM), a new experience replay buffer for reinforcement learning that prioritizes storing transitions with high uncertainty estimates. Unlike traditional buffers that rely on temporal difference error or transition distributions, UDRM updates its contents based on uncertainty derived from the RL model during training. Experiments show that this uncertainty-aware buffer leads to higher rewards during training compared to other uncertainty-aware RL frameworks.

By Sheeraja Rajakrishnan, Alexander G. Ororbia, Travis Desell, Daniel E. Krutz
arXiv AI
Sep 11

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

Efficient Diversity-based Experience Replay (EDER) is introduced to enhance learning efficiency in deep reinforcement learning. It uses a determinantal point process to model sample diversity, prioritizes replay based on this diversity, and incorporates Cholesky decomposition and rejection sampling to handle large state spaces. Experiments on MuJoCo, Atari, and Habitat show that EDER significantly improves learning efficiency and performance in high‑dimensional, realistic environments.

By Kaiyan Zhao, Yiming Wang, Yuyang Chen, Yan Li, Leong Hou U, Xiaoguang Niu
arXiv Machine Learning
4d ago

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning

The paper introduces Space-sampled Value Decay (SsVD), a forgetting mechanism designed for non-stationary reinforcement learning where the environment can drift at every timestep. SsVD selectively pulls value estimates of randomly chosen state-space elements toward a baseline, discarding outdated information without requiring reset or change-point detection. Integrated into Soft Actor Critic and Deep Q-Networks, SsVD outperforms its base algorithms across six non-stationary environments and can also promote optimism in hard-exploration tasks.

By Felix St\"orck, Philipp Hartmann, Fabian Hinder, Klaus Neumann, Barbara Hammer
arXiv AI
Jun 30

Dual-Flow Reinforcement Learning with State-Aware Exploration

arXiv:2606. 29820v1 Announce Type: cross Abstract: In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging.

By Qijun Li, Zheng Fu, Qi Song, Yifei He, Weitao Zhou, Kun Jiang, Diange Yang