arXiv:2607. 11964v1 Announce Type: new Abstract: Autonomous driving requires long-horizon closedloop decision making in dynamic traffic environments.
By Yongzhi Liu, Yang Xiao, Zhong Cao, Zeng Kang, Sunan Zhang, Zhaozhi Dong, Guojun Yu, Weichao Zhuang
Autonomous driving requires long-horizon closedloop decision making in dynamic traffic environments. Latent world models offer an effective framework for this problem by enabling imagination-based decision making in compact latent spaces.
The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.
By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv:2606. 20104v1 Announce Type: cross Abstract: Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions.
By Petr Ivashkov, Randall Balestriero, Bernhard Sch\"olkopf
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
By Wancong Zhang, Basile Terver, Artem Zholus, Soham Chitnis, Harsh Sutaria, Mido Assran, Randall Balestriero, Amir Bar, Adrien Bardes, Yann LeCun, Nicolas Ballas
arXiv:2604. 25416v2 Announce Type: replace Abstract: Model-based reinforcement learning distinguishes between dynamics models operating on proprioceptive states and latent dynamics models typically operating on high-dimensional image observations.
By Julia Berger, Bernd Frauenknecht, Sebastian Trimpe, Bastian Leibe
The paper investigates the LeWorldModel (LeWM) and its Sketched Isotropic Gaussian Regularizer (SIGReg), showing that the original Raw LeWM objective biases variance toward temporally persistent components, which suppresses residual variance and hampers robot state decodability. By applying SIGReg specifically to temporally centered residuals, the authors decouple persistent and residual variance allocation, improving representation quality. On the LIBERO benchmark, this adjustment boosts downstream policy success on the Goal suite by 1.66× and raises overall success rates from 63.6% to 83.8%, outperforming Diffusion Policy and pretrained OpenVLA without external pretraining.
By Chang Liu, Fei Suo, Yanzhou Jin, Zeyu Ping, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
arXiv:2609.13845v1 Announce Type: cross
Abstract: World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet pl...
By Saksham Bansal, Om Naphade, Chayan Aggarwal, Vrishin M
arXiv:2608. 04471v1 Announce Type: cross Abstract: Time series in real-world applications are often generated by nonlinear dynamical systems, making accurate forecasting challenging.
By Mengzhou Gao, Huangqian Yu, Pengfei Jiao
HaM-World introduces a structured world model that combines history-conditioned selective memory with a Soft‑Hamiltonian latent dynamics prior. The model decomposes the latent state into a canonical (q,p) subspace governed by an energy‑derived Hamiltonian vector field and a context subspace c capturing non‑conservative factors, while Mamba selective state‑space memory conditions the transition used for prediction, reward, value estimation, and planning. Across six DeepMind Control Suite tasks, HaM-World achieves top rankings on four tasks, improves average AUC, reduces imagined‑rollout error by 45% on short‑to‑medium horizons, and outperforms baselines under 12 out‑of‑distribution perturbations.
By Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang
arXiv:2607. 27138v1 Announce Type: cross Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change.
By Zuojin Tang, Feifan Luo, Haoyun Liu, Botai Yuan, Dekang Qi, Ronghan Chen, Yandan Yang, Tong Lin, Xinyuan Chang, Mu Xu, Bin Liu, De Ma, Zhiheng Ma
Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.
By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do