arXiv:2606. 20104v1 Announce Type: cross Abstract: Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions.
By Petr Ivashkov, Randall Balestriero, Bernhard Sch\"olkopf
arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.
By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv:2602. 05031v2 Announce Type: replace Abstract: Planning with a learned model remains a key challenge in model-based reinforcement learning (RL).
By Dikshant Shehmar, Matthew Schlegel, Matthew E. Taylor, Marlos C. Machado
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
By Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
arXiv:2608.24855v1 Announce Type: new
Abstract: Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online traje...
By Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu, Jenq-Neng Hwang
The Representation World Model (RWM) learns states, transitions, and executable plans directly within a representation space, bypassing traditional explicit dynamics models and action-space search. It uses inverse-dynamics supervision along latent paths to shape the representation geometry, enabling direct planning by constructing a latent path between current and goal states and recovering actions via inverse dynamics. Experiments on continuous-control benchmarks and robotic manipulation tasks demonstrate RWM’s effectiveness and potential for complex embodied control.
By Yijun Yuan, Weicheng Zheng, Weibang Wang, Minghui Qin, Chang Sun, Junhao Huang, Kenan Li, Anmin Liu, Yicheng Yao, Hang Zhao
arXiv:2603. 12231v2 Announce Type: replace Abstract: Learning good representations is essential for latent planning with world models.
By Ying Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, Mengye Ren
arXiv:2608. 16287v1 Announce Type: new Abstract: Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself.
By Jiaming Hu, Yan Zheng, Tian Wang
Joint-embedding predictive architectures (JEPAs) learn latent dynamics for planning and avoid representation collapse by matching features to maximum-entropy distributions such as isotropic Gaussians,...
arXiv:2606. 27014v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input level.
By Jingyi Cui, Qi Zhang, Hongwei Wen, Yisen Wang
The paper introduces ATLAS, a training objective that preserves relational geometry in latent world models while calibrating the global latent distribution. By transferring normalized pairwise structure from an informative encoder to the planning latent and applying Wasserstein embedding matching, ATLAS improves goal‑reaching success on tasks such as PushT, TwoRoom, and OGBench‑Cube, especially on higher‑novelty episodes. Diagnostics show stronger novelty‑related structure, better marginal calibration, and lower multi‑step prediction error in the planning latent.
By Ke Fang, Yupu Yao, Lu Cheng