The paper investigates why latent‑flow world models that use a frozen self‑supervised latent space lose the ability to manipulate motion. It shows that the pretrained flow does not move the manipulated object and that training with latent‑only losses only produces stillness or teleport‑like motion. The authors introduce Decode‑Augmented Rollout Training (DART), which keeps the representation frozen but retrains the flow using decode‑path supervision, restoring temporal motion structure and improving prediction quality, even closing much of the gap to an oracle‑informed reference. The study also notes that pixel error alone can favor frozen predictions.
By Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu
arXiv:2607. 07763v1 Announce Type: new Abstract: World models are typically trained to predict discrete-time physical dynamics with a fixed step size baked into the model weights, preventing prediction at variable temporal resolutions.
By Eli Laird, Corey Clark
Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.
By Everest Yang
Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow ne...
arXiv:2607. 28362v1 Announce Type: cross Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models.
By Jin Cao, Zian Meng, Kaipeng Zhang
The paper demonstrates that learned simulators can fail in two distinct ways when conditions change: long‑horizon drift due to accumulated errors and incorrect responses to interventions on physical parameters. By adding a symplectic integrator to preserve conservative dynamics, rollouts remain stable for up to 100× the training horizon, while encoding physical coupling via explicit linear factorization allows the model to generalize to unseen signs of that coupling. The study shows that stability and counterfactual generalization arise from separate structural choices, enabling designers to impose each property independently.
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv:2609.10464v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning,...
By Andy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero, John Sous
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2606. 24945v1 Announce Type: new Abstract: We ask a representation-learning question about physical world models: when does a conservation law remain certifiable after a model learns a latent representation?
By Hongbo Wang
arXiv:2607. 27036v1 Announce Type: cross Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time.
By Taiye Chen, Qi Zhang, Yisen Wang
arXiv:2602. 13977v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interaction prevents direct deployment on physical robots.
By Zhennan Jiang, Shangqing Zhou, Yutong Jiang, Zefang Huang, Mingjie Wei, Yuhui Chen, Tianxing Zhou, Zhen Guo, Hao Lin, Quanlu Zhang, Yu Wang, Haoran Li, Chao Yu, Dongbin Zhao
JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi