JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv:2608.24044v1 Announce Type: new
Abstract: Latent world models plan by predicting how candidate actions transform learned representations. In self-predictive models, however, the encoder and pre...
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv:2606. 16076v1 Announce Type: cross Abstract: Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evolution.
By Weizhi Nie, Weichao Liu, Honglin Guo, Yuting Su
A robot policy trained with either absolute joint targets or delta‑relative actions inherits the chosen action parameterization in its world model, leading to catastrophic failures when the model is exposed to the alternate encoding. Experiments on three robot datasets and two morphologies show retrieval performance drops 2.6–13.4×, goal‑conditioned action selection plummets from 53% to 15%, and predictions for the same future become nearly orthogonal. The issue is not a loss of information—both encodings are highly reconstructible—but a lack of invariance in the action channel, which existing visual‑model invariance research does not address. Averaging over the two encodings partially restores performance, yet the worst‑case disagreement remains high, indicating that the defect persists in certain scenarios.
By Ahmed Karim, Leon Chlon
arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.
By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
The paper demonstrates that learned simulators can fail in two distinct ways when conditions change: long‑horizon drift due to accumulated errors and incorrect responses to interventions on physical parameters. By adding a symplectic integrator to preserve conservative dynamics, rollouts remain stable for up to 100× the training horizon, while encoding physical coupling via explicit linear factorization allows the model to generalize to unseen signs of that coupling. The study shows that stability and counterfactual generalization arise from separate structural choices, enabling designers to impose each property independently.
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv:2609.37378v1 Announce Type: cross
Abstract: Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occur...
By Hossein Resani, Javen Qinfeng Shi
arXiv:2607. 18715v1 Announce Type: new Abstract: Latent world models underpin much of modern model-based control, yet current action-conditioned formulations supervise the next-latent transition with a single, undifferentiated target, forcing a monolithic learning signal to absorb every source of state change.
By Yi-Ge Zhang, Tianqi Du, Qi Zhang, Yisen Wang
arXiv:2606. 28455v1 Announce Type: cross Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics.
By Yang Liu, Yuming Chen
arXiv:2608. 00591v2 Announce Type: replace Abstract: A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches.
By Yibin Dong
The paper introduces a formal framework and benchmark for time‑series world models (TSWMs) that separates state, actions, and exogenous inputs, and defines a new metric called mechanism consistency to evaluate whether model predictions move in the expected direction when actions change. Experiments on eight public datasets show that using a frozen latent prediction space and gated output fusion improves prediction accuracy, while prediction error and mechanism consistency often diverge, with the best‑performing models sometimes failing to exhibit consistent directional responses. Adding a directional supervision loss significantly boosts mechanism consistency without affecting mean‑absolute error, providing a practical recipe for building more reliable TSWMs.
By Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen
arXiv:2607. 06640v1 Announce Type: cross Abstract: A learned world model is usually judged by how faithfully it reconstructs its observations or predicts reward, as though quality were something the model simply has or lacks.
By Donna Vakalis