arXiv AI By Kaiyan Zhao, Yiming Wang, Yuyang Chen, Yan Li, Leong Hou U, Xiaoguang Niu

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

Read the original on arXiv AI →

Efficient Diversity-based Experience Replay (EDER) is introduced to enhance learning efficiency in deep reinforcement learning. It uses a determinantal point process to model sample diversity, prioritizes replay based on this diversity, and incorporates Cholesky decomposition and rejection sampling to handle large state spaces. Experiments on MuJoCo, Atari, and Habitat show that EDER significantly improves learning efficiency and performance in high‑dimensional, realistic environments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 18

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.

arXiv AI
Aug 19

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

The paper proposes Novelty and Surprise Prioritized Experience Replay (NSPER) for image-based reinforcement learning, combining novelty to highlight underrepresented states and surprise to reveal gaps in the agent’s knowledge. An extended version, NSPER+R, also uses these signals as intrinsic rewards to enhance both replay quality and exploration. Experiments on DeepMind Control Suite tasks demonstrate that NSPER and NSPER+R accelerate training and improve convergence compared to existing methods.

By Hoda Yamani, Henry Williams, Bruce A. MacDonald