arXiv Machine Learning By Kaizhen Tan (New York University, Carnegie Mellon University), Xin Xu (Carnegie Mellon University), Siru Tao (Carnegie Mellon University), Hanzhe Hong (Carnegie Mellon University), Yang Feng (Columbia University), Heqing Du (Columbia University)

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

Read the original on arXiv Machine Learning →

arXiv:2607. 27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv Machine Learning
Sep 22

Robot World Models Are Not Invariant to How the Actions Are Written

A robot policy trained with either absolute joint targets or delta‑relative actions inherits the chosen action parameterization in its world model, leading to catastrophic failures when the model is exposed to the alternate encoding. Experiments on three robot datasets and two morphologies show retrieval performance drops 2.6–13.4×, goal‑conditioned action selection plummets from 53% to 15%, and predictions for the same future become nearly orthogonal. The issue is not a loss of information—both encodings are highly reconstructible—but a lack of invariance in the action channel, which existing visual‑model invariance research does not address. Averaging over the two encodings partially restores performance, yet the worst‑case disagreement remains high, indicating that the defect persists in certain scenarios.

By Ahmed Karim, Leon Chlon
arXiv Machine Learning
Jul 14

A Control Theory of Predictability in Latent World Models

arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.

By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
arXiv Machine Learning
Sep 18

Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models

The paper demonstrates that learned simulators can fail in two distinct ways when conditions change: long‑horizon drift due to accumulated errors and incorrect responses to interventions on physical parameters. By adding a symplectic integrator to preserve conservative dynamics, rollouts remain stable for up to 100× the training horizon, while encoding physical coupling via explicit linear factorization allows the model to generalize to unseen signs of that coupling. The study shows that stability and counterfactual generalization arise from separate structural choices, enabling designers to impose each property independently.

By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling