JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv:2608.24044v1 Announce Type: new
Abstract: Latent world models plan by predicting how candidate actions transform learned representations. In self-predictive models, however, the encoder and pre...
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv:2606. 16076v1 Announce Type: cross Abstract: Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evolution.
By Weizhi Nie, Weichao Liu, Honglin Guo, Yuting Su
A robot policy trained with either absolute joint targets or delta‑relative actions inherits the chosen action parameterization in its world model, leading to catastrophic failures when the model is exposed to the alternate encoding. Experiments on three robot datasets and two morphologies show retrieval performance drops 2.6–13.4×, goal‑conditioned action selection plummets from 53% to 15%, and predictions for the same future become nearly orthogonal. The issue is not a loss of information—both encodings are highly reconstructible—but a lack of invariance in the action channel, which existing visual‑model invariance research does not address. Averaging over the two encodings partially restores performance, yet the worst‑case disagreement remains high, indicating that the defect persists in certain scenarios.
By Ahmed Karim, Leon Chlon
arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.
By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
The paper demonstrates that learned simulators can fail in two distinct ways when conditions change: long‑horizon drift due to accumulated errors and incorrect responses to interventions on physical parameters. By adding a symplectic integrator to preserve conservative dynamics, rollouts remain stable for up to 100× the training horizon, while encoding physical coupling via explicit linear factorization allows the model to generalize to unseen signs of that coupling. The study shows that stability and counterfactual generalization arise from separate structural choices, enabling designers to impose each property independently.
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling