arXiv AI

Correcting a learned physical invariant improves world-model rollouts

arXiv AI
Sep 24

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

The paper investigates why latent‑flow world models that use a frozen self‑supervised latent space lose the ability to manipulate motion. It shows that the pretrained flow does not move the manipulated object and that training with latent‑only losses only produces stillness or teleport‑like motion. The authors introduce Decode‑Augmented Rollout Training (DART), which keeps the representation frozen but retrains the flow using decode‑path supervision, restoring temporal motion structure and improving prediction quality, even closing much of the gap to an oracle‑informed reference. The study also notes that pixel error alone can favor frozen predictions.

By Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu
arXiv Machine Learning
Sep 17

Changepoint-Aware World Models: Detecting Dynamics Shifts and Recovering by Forgetting Stale Replay in Model-Based RL

Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.

By Everest Yang
arXiv Machine Learning
Sep 18

Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models

The paper demonstrates that learned simulators can fail in two distinct ways when conditions change: long‑horizon drift due to accumulated errors and incorrect responses to interventions on physical parameters. By adding a symplectic integrator to preserve conservative dynamics, rollouts remain stable for up to 100× the training horizon, while encoding physical coupling via explicit linear factorization allows the model to generalize to unseen signs of that coupling. The study shows that stability and counterfactual generalization arise from separate structural choices, enabling designers to impose each property independently.

By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv AI
Jun 30

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

arXiv:2602. 13977v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interaction prevents direct deployment on physical robots.

By Zhennan Jiang, Shangqing Zhou, Yutong Jiang, Zefang Huang, Mingjie Wei, Yuhui Chen, Tianxing Zhou, Zhen Guo, Hao Lin, Quanlu Zhang, Yu Wang, Haoran Li, Chao Yu, Dongbin Zhao
arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi