arXiv Machine Learning By Bonnie Li

Contrastive World Models

Read the original on arXiv Machine Learning →

Contrastive World Models propose a new method for learning latent dynamics without pixel reconstruction. By replacing observation reconstruction with a Deep InfoMax-like objective that maximizes mutual information between state-action sequences and local patch features of future observations, the approach encourages state representations to retain predictive information while ignoring visually irrelevant details. Experiments show that this method matches existing baselines in simple settings and significantly outperforms them when distractors or natural video backgrounds are present, while also training more efficiently by eliminating the pixel decoder.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
4d ago

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM introduces a self‑supervised attention routing mechanism that selectively applies reconstruction only to dynamically relevant regions of visual input. By combining a cross‑attention module with a dual‑stream decoder and stop‑gradient barriers, it balances faithful environmental dynamics capture with abstraction of irrelevant content. Experiments on DeepMind Control show that StarWM outperforms both reconstruction‑based and reconstruction‑free baselines, especially under distractor conditions, and preserves state attributes over long‑horizon imagination.

By Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald, Arne Peter Raulf, Daniel Alexander Braun
arXiv AI
Aug 10

TaskSense: Focusing on What Matters in World Models

arXiv:2608. 06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input.

By SM Mazharul Islam, Manfred Huber
arXiv AI
Jun 16

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.

By Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu
arXiv AI
Jul 1

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.

By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu