The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.
arXiv:2608. 19684v1 Announce Type: new Abstract: Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL.
By Tanachai Anakewat, Takayuki Osa, Tatsuya Harada
The paper proposes Novelty and Surprise Prioritized Experience Replay (NSPER) for image-based reinforcement learning, combining novelty to highlight underrepresented states and surprise to reveal gaps in the agent’s knowledge. An extended version, NSPER+R, also uses these signals as intrinsic rewards to enhance both replay quality and exploration. Experiments on DeepMind Control Suite tasks demonstrate that NSPER and NSPER+R accelerate training and improve convergence compared to existing methods.
By Hoda Yamani, Henry Williams, Bruce A. MacDonald
arXiv:2506. 01568v4 Announce Type: replace Abstract: Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima.
By Cornelius V. Braun, Sayantan Auddy, Marc Toussaint
arXiv:2602. 12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the representational strengths of model-based approaches, without incurring planning overhead.
By Jashaswimalya Acharjee, Balaraman Ravindran
arXiv:2505. 03296v2 Announce Type: replace-cross Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy representation and imitation learning in robot manipulation.
By Jan Ole von Hartz, Adrian R\"ofer, Joschka Boedecker, Abhinav Valada
The paper proposes Instant Episode Repetition (IER), a method that immediately repeats action sequences from high‑reward episodes during training to boost sample efficiency. Unlike passive replay techniques, IER actively shapes data collection by re‑executing successful behaviors for a set number of subsequent episodes. Experiments on MuJoCo, the DeepMind Control Suite, and a robotic manipulation task show that IER improves learning performance over standard SAC, TD3, and self‑imitation baselines.
By Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, Henry Williams
arXiv:2606. 20411v1 Announce Type: new Abstract: Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms.
By Hsiao-Ru Pan, Bernhard Sch\"olkopf
arXiv:2602. 05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization.
By Hua Zheng, Wei Xie, M. Ben Feng, Keilung Choy
arXiv:2501. 04426v2 Announce Type: replace-cross Abstract: Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, improving robustness to distribution shift without additional environment interaction.
By Pavel Kolev, Marin Vlastelica, Georg Martius
arXiv:2606. 01151v1 Announce Type: new Abstract: Behavior cloning with high-capacity generative policies achieves strong imitation performance, but is often limited by demonstration coverage and distribution shift.
By Hikmet Simsir, Ozgur S. Oguz
arXiv:2606. 09758v1 Announce Type: cross Abstract: Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment.
By Quinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal, Siddhartha Srinivasa, Abhishek Gupta