arXiv:2603. 08558v3 Announce Type: replace Abstract: Learning compact state representations in Markov Decision Processes (MDPs) has proven crucial for addressing the curse of dimensionality in large-scale reinforcement learning (RL) problems.
By Tommaso Giorgi, Pierriccardo Olivieri, Keyue Jiang, Laura Toni, Matteo Papini
arXiv:2609.37677v1 Announce Type: cross
Abstract: Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, th...
By Feiyang Wu, Chenxiao Gao, Chen Yang, Ye Zhao, Bo Dai, Anqi Wu
arXiv:2607. 00811v1 Announce Type: new Abstract: Unsupervised pre-training on large-scale datasets has demonstrated significant potential for improving the sample efficiency and performance of Reinforcement Learning (RL).
By Jinwen Wang, Youfang Lin, Xiaobo Hu, Siyu Yang, Sheng Han, Shuo Wang, Kai Lv
arXiv:2506. 09276v4 Announce Type: replace-cross Abstract: This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectories, requiring neither reward signals nor the actions executed by the agent.
By Lorenzo Steccanella, Joshua B. Evans, \"Ozg\"ur \c{S}im\c{s}ek, Anders Jonsson
arXiv:2605. 26012v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value and policy structure may be intrinsically low-dimensional.
By Aleksandar Todorov, Matthia Sabatelli
arXiv:2607. 21670v1 Announce Type: cross Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies.
By Chaoqi Liu, Yue Zhao, Haonan Chen, Xiaoshen Han, Jiawei Gao, Ehsan Adeli, Yilun Du
arXiv:2607. 00808v1 Announce Type: new Abstract: Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging.
By Jinwen Wang, Youfang Lin, Xiaobo Hu, Shuo Wang, Kai Lv
arXiv:2608.30640v1 Announce Type: new
Abstract: While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open q...
By Michal Korniak, Kamil Dybek, Benjamin Eysenbach, Marco Bagatella, Micha{\l} Bortkiewicz
arXiv:2606. 03017v1 Announce Type: cross Abstract: Reward transfer in Inverse Reinforcement Learning (IRL) is unreliable when policies must generalize to unseen combinations of environment dynamics and task goals.
By Yikang Gui, Bikramjit Banerjee, Prashant Doshi
arXiv:2606. 20104v1 Announce Type: cross Abstract: Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions.
By Petr Ivashkov, Randall Balestriero, Bernhard Sch\"olkopf
The paper introduces Commute-Time-Preserving World Models (CTWMs), which learn latent representations that reflect commute-times in an environment by using a latent displacement predictor and a log-determinant regularizer. This approach addresses the issue that existing self-supervised methods degrade the necessary eigenvalue-dependent scaling for accurate commute-time representation. In experiments, CTWMs outperform the task-agnostic baseline LeWM on several continuous goal-reaching benchmarks while using only half the parameters.
By Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke
The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.
By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White