arXiv:2608.30640v1 Announce Type: new
Abstract: While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open q...
By Michal Korniak, Kamil Dybek, Benjamin Eysenbach, Marco Bagatella, Micha{\l} Bortkiewicz
arXiv:2606. 10979v1 Announce Type: new Abstract: Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints.
By Yi Chen (Lucy), Rushuai Yang (Lucy), Qiang Chen (Lucy), Dongyan (Lucy), Huo
arXiv:2607. 18004v1 Announce Type: new Abstract: Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by reward and transition similarity.
By Daegyeong Roh, Juho Bae, Han-Lim Choi
arXiv:2603. 08558v3 Announce Type: replace Abstract: Learning compact state representations in Markov Decision Processes (MDPs) has proven crucial for addressing the curse of dimensionality in large-scale reinforcement learning (RL) problems.
By Tommaso Giorgi, Pierriccardo Olivieri, Keyue Jiang, Laura Toni, Matteo Papini
The paper proposes a method for learning task-relevant representations in deep reinforcement learning by maximizing rollout total correlation, which captures the correlation among all learned representations and actions across entire trajectories. It introduces two complementary lower bounds—one generative and one discriminative—along with chunk‑wise mini‑batching to improve this objective, and also proposes an intrinsic reward derived from the learned representation to enhance exploration. Experiments on challenging image‑based simulated control tasks demonstrate improved sample efficiency and robustness to white noise and natural video backgrounds compared to leading baselines.
By Bang You, Huaping Liu, Jan Peters, Oleg Arenz
The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.
By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White