arXiv:2608. 06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input.
By SM Mazharul Islam, Manfred Huber
Contrastive World Models propose a new method for learning latent dynamics without pixel reconstruction. By replacing observation reconstruction with a Deep InfoMax-like objective that maximizes mutual information between state-action sequences and local patch features of future observations, the approach encourages state representations to retain predictive information while ignoring visually irrelevant details. Experiments show that this method matches existing baselines in simple settings and significantly outperforms them when distractors or natural video backgrounds are present, while also training more efficiently by eliminating the pixel decoder.
By Bonnie Li
World in World introduces a training‑free inference interface that lets users control autoregressive video world models from new viewpoints. By converting diverse control signals—source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual tokens, the system uses a frozen causal video model’s self‑attention to maintain synchronization, complete unseen regions, and recover past appearances. The method supports camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer while preserving perceptual quality, temporal consistency, and camera‑following accuracy.
World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.
By Chenxi Song, Yanming Yang, Chi Zhang
arXiv:2605.15618v2 Announce Type: replace-cross
Abstract: Self-supervised video models are increasingly framed as world models, yet they are still evaluated almost entirely on clean video and reporte...
By Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar
arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
By Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim
arXiv:2606. 12217v1 Announce Type: cross Abstract: World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions.
By Lu Qiu, Yizhuo Li, Yi Chen, Yuying Ge, Yixiao Ge, Xihui Liu
arXiv:2608. 07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions.
By Yuehao Huang, Yunzi Wu, Xiaotao Zhang, Xinhai Li, Jiankun Dong, Jiajun Lv, Chi Zhang, Chenjia Bai, Yong Liu, Xuelong Li
RECAP-Forcing is a new method for long autoregressive video generation that addresses the memory challenge by organizing memory based on appearance novelty rather than recency. The approach retains key-value caches for newly appearing content—such as entering subjects, disoccluded regions, and new scenes—at the moment they first appear, ensuring consistent identities over time. It combines an attention sink for the initial scene with an optical-flow-based novelty bank for later frames, improving visual quality and semantic fidelity without adding learnable parameters.
By Haiyang Xu, Zheng Ding, Zhuowen Tu
World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions.
WorldCrafter is a video world model that introduces a camera‑queryable implicit 3D‑aware memory to improve long‑horizon consistency and viewpoint control. The model compresses multi‑view evidence into a limited token budget shaped by the requested viewpoint, integrating historical observations via a memory encoder and pose‑conditioned readout before denoising. Experiments on static and dynamic scenes demonstrate significant gains in consistency and camera‑control accuracy while maintaining visual quality during minute‑scale exploration.
By Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao, Yingmin Luo, Ying Shan
arXiv:2608.20974v1 Announce Type: cross
Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...
By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang