arXiv Machine Learning

CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization

CF-JEPA is a JEPA-style latent world model that separates the latent space into controllable and uncontrollable subspaces, allowing distractor information to be captured in the uncontrollable region. This factorization prevents latent collapse and maintains performance across 2D and 3D control tasks, even under distracted conditions. The model is validated on a simulated robot task, demonstrating its practical applicability.

arXiv AI
Aug 10

TaskSense: Focusing on What Matters in World Models

arXiv:2608. 06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input.

By SM Mazharul Islam, Manfred Huber
arXiv Machine Learning
Sep 22

Contrastive World Models

Contrastive World Models propose a new method for learning latent dynamics without pixel reconstruction. By replacing observation reconstruction with a Deep InfoMax-like objective that maximizes mutual information between state-action sequences and local patch features of future observations, the approach encourages state representations to retain predictive information while ignoring visually irrelevant details. Experiments show that this method matches existing baselines in simple settings and significantly outperforms them when distractors or natural video backgrounds are present, while also training more efficiently by eliminating the pixel decoder.

By Bonnie Li
arXiv Machine Learning
3d ago

JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models

JEPA-Bisim introduces a bisimulation encoder to joint-embedding predictive world models, ensuring that states with similar transition dynamics are mapped to nearby latent representations while suppressing irrelevant slow features such as background changes and distractors. The approach improves robustness on navigation (PointMaze) and manipulation (PushT) tasks under varied test-time visual conditions, achieving up to tenfold smaller latent spaces than DINO-WM. It remains effective across different pre-trained visual encoders, including DINOv2, SimDINOv2, and iBOT.

By Leonardo F. Toso, Davit Shadunts, Yunyang Lu, Gloria Geng, Nihal Sharma, Donglin Zhan, Nam H. Nguyen, James Anderson
arXiv AI
Aug 10

Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection

arXiv:2608. 06706v1 Announce Type: cross Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the training loss keeps improving.

By Jiazhuo Li, Yiming Fei, Zhiruo Zhou, Heikichi Hayashi
arXiv AI
Jun 16

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.

By Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu
arXiv Computer Vision
Sep 28

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM introduces a self‑supervised attention routing mechanism that selectively applies reconstruction only to dynamically relevant regions of visual input. By combining a cross‑attention module with a dual‑stream decoder and stop‑gradient barriers, it balances faithful environmental dynamics capture with abstraction of irrelevant content. Experiments on DeepMind Control show that StarWM outperforms both reconstruction‑based and reconstruction‑free baselines, especially under distractor conditions, and preserves state attributes over long‑horizon imagination.

By Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald, Arne Peter Raulf, Daniel Alexander Braun
arXiv AI
1d ago

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

The paper introduces VIGOR, a framework for zero‑shot visual generalization in model‑based reinforcement learning. VIGOR enforces latent‑space consistency through asymmetric weak‑to‑strong augmentations, dynamics‑level consistency, and encoder‑level stabilization, allowing the agent to handle unseen visual distractions while maintaining sample efficiency. Experiments on the DeepMind Control Suite and Robosuite demonstrate that VIGOR outperforms state‑of‑the‑art baselines, achieving significant gains in both environments.

By Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee