arXiv Machine Learning
Sep 21

Rollout Total Correlation for Deep Reinforcement Learning

The paper proposes a method for learning task-relevant representations in deep reinforcement learning by maximizing rollout total correlation, which captures the correlation among all learned representations and actions across entire trajectories. It introduces two complementary lower bounds—one generative and one discriminative—along with chunk‑wise mini‑batching to improve this objective, and also proposes an intrinsic reward derived from the learned representation to enhance exploration. Experiments on challenging image‑based simulated control tasks demonstrate improved sample efficiency and robustness to white noise and natural video backgrounds compared to leading baselines.

By Bang You, Huaping Liu, Jan Peters, Oleg Arenz
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
arXiv Machine Learning
Sep 17

DSD: Learning Diverse and Reusable Motor Skills via Diffusion Skill Discovery

The paper introduces Diffusion Skill Discovery (DSD), a method that employs a diffusion model to approximate the entropy gradient of policy-induced state distributions via score matching. This approach encourages the learning of a diverse repertoire of motor skills for high‑dimensional humanoid control, overcoming limitations of prior mutual‑information based methods that rely on indirect state entropy estimates. The discovered skills are effectively reused in hierarchical control and zero‑shot tasks, yielding more complex and agile behaviors than previous skill discovery techniques.

By Sun Woo Kim, Xue Bin Peng