arXiv AI By Lorenzo Steccanella, Joshua B. Evans, \"Ozg\"ur \c{S}im\c{s}ek, Anders Jonsson

Learning The Minimum Action Distance

Read the original on arXiv AI →

arXiv:2506. 09276v4 Announce Type: replace-cross Abstract: This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectories, requiring neither reward signals nor the actions executed by the agent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

Rollout Total Correlation for Deep Reinforcement Learning

The paper proposes a method for learning task-relevant representations in deep reinforcement learning by maximizing rollout total correlation, which captures the correlation among all learned representations and actions across entire trajectories. It introduces two complementary lower bounds—one generative and one discriminative—along with chunk‑wise mini‑batching to improve this objective, and also proposes an intrinsic reward derived from the learned representation to enhance exploration. Experiments on challenging image‑based simulated control tasks demonstrate improved sample efficiency and robustness to white noise and natural video backgrounds compared to leading baselines.

By Bang You, Huaping Liu, Jan Peters, Oleg Arenz
arXiv Machine Learning
Aug 27

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.

By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White