arXiv Machine Learning By Jai Malegaonkar, Rohan Patil, Henrik I. Christensen

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 05111v1 Announce Type: new Abstract: In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 23

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

The paper introduces ELEMENT, a framework that combines episodic and lifelong entropy maximization to drive reward-free exploration in reinforcement learning. It addresses two key limitations of existing entropy-based methods: the vanishing intrinsic reward after a state is visited and the computational cost of estimating entropy over large datasets. ELEMENT achieves this by deriving an average episodic state entropy reward and employing a k‑NN graph‑based estimator for lifelong entropy, leading to superior state coverage and unsupervised pre‑training performance compared to current baselines.

By Hongming Li, Zhao Yang, Xiaoxuan Liang, Shujian Yu, Jose C. Principe