arXiv AI

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

The paper introduces VIGOR, a framework for zero‑shot visual generalization in model‑based reinforcement learning. VIGOR enforces latent‑space consistency through asymmetric weak‑to‑strong augmentations, dynamics‑level consistency, and encoder‑level stabilization, allowing the agent to handle unseen visual distractions while maintaining sample efficiency. Experiments on the DeepMind Control Suite and Robosuite demonstrate that VIGOR outperforms state‑of‑the‑art baselines, achieving significant gains in both environments.

arXiv AI
Aug 10

TaskSense: Focusing on What Matters in World Models

arXiv:2608. 06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input.

By SM Mazharul Islam, Manfred Huber
arXiv AI
6d ago

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv AI
Jul 1

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.

By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv Machine Learning
6d ago

Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

arXiv:2609.37165v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific f...

By Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, Gi-Cheon Kang, Byoung-Tak Zhang
arXiv AI
Jun 3

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

arXiv:2606. 02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control.

By Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota