arXiv Computer Vision

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM introduces a self‑supervised attention routing mechanism that selectively applies reconstruction only to dynamically relevant regions of visual input. By combining a cross‑attention module with a dual‑stream decoder and stop‑gradient barriers, it balances faithful environmental dynamics capture with abstraction of irrelevant content. Experiments on DeepMind Control show that StarWM outperforms both reconstruction‑based and reconstruction‑free baselines, especially under distractor conditions, and preserves state attributes over long‑horizon imagination.

arXiv AI
Aug 10

TaskSense: Focusing on What Matters in World Models

arXiv:2608. 06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input.

By SM Mazharul Islam, Manfred Huber
arXiv Machine Learning
Sep 22

Contrastive World Models

Contrastive World Models propose a new method for learning latent dynamics without pixel reconstruction. By replacing observation reconstruction with a Deep InfoMax-like objective that maximizes mutual information between state-action sequences and local patch features of future observations, the approach encourages state representations to retain predictive information while ignoring visually irrelevant details. Experiments show that this method matches existing baselines in simple settings and significantly outperforms them when distractors or natural video backgrounds are present, while also training more efficiently by eliminating the pixel decoder.

By Bonnie Li
Hugging Face Trending Papers
Sep 10

World in World: Explore the World with World Models

World in World introduces a training‑free inference interface that lets users control autoregressive video world models from new viewpoints. By converting diverse control signals—source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual tokens, the system uses a frozen causal video model’s self‑attention to maintain synchronization, complete unseen regions, and recover past appearances. The method supports camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer while preserving perceptual quality, temporal consistency, and camera‑following accuracy.

arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang
arXiv AI
Aug 10

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

arXiv:2608. 07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions.

By Yuehao Huang, Yunzi Wu, Xiaotao Zhang, Xinhai Li, Jiankun Dong, Jiajun Lv, Chi Zhang, Chenjia Bai, Yong Liu, Xuelong Li
arXiv Computer Vision
Aug 28

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

RECAP-Forcing is a new method for long autoregressive video generation that addresses the memory challenge by organizing memory based on appearance novelty rather than recency. The approach retains key-value caches for newly appearing content—such as entering subjects, disoccluded regions, and new scenes—at the moment they first appear, ensuring consistent identities over time. It combines an attention sink for the initial scene with an optical-flow-based novelty bank for later frames, improving visual quality and semantic fidelity without adding learnable parameters.

By Haiyang Xu, Zheng Ding, Zhuowen Tu
arXiv Computer Vision
Sep 22

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter is a video world model that introduces a camera‑queryable implicit 3D‑aware memory to improve long‑horizon consistency and viewpoint control. The model compresses multi‑view evidence into a limited token budget shaped by the requested viewpoint, integrating historical observations via a memory encoder and pose‑conditioned readout before denoising. Experiments on static and dynamic scenes demonstrate significant gains in consistency and camera‑control accuracy while maintaining visual quality during minute‑scale exploration.

By Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao, Yingmin Luo, Ying Shan
arXiv AI
Aug 24

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

arXiv:2608.20974v1 Announce Type: cross Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...

By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang