JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv:2608.24044v1 Announce Type: new
Abstract: Latent world models plan by predicting how candidate actions transform learned representations. In self-predictive models, however, the encoder and pre...
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.
arXiv:2608. 09876v1 Announce Type: cross Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics.
By Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang
The paper investigates the LeWorldModel (LeWM) and its Sketched Isotropic Gaussian Regularizer (SIGReg), showing that the original Raw LeWM objective biases variance toward temporally persistent components, which suppresses residual variance and hampers robot state decodability. By applying SIGReg specifically to temporally centered residuals, the authors decouple persistent and residual variance allocation, improving representation quality. On the LIBERO benchmark, this adjustment boosts downstream policy success on the Goal suite by 1.66× and raises overall success rates from 63.6% to 83.8%, outperforming Diffusion Policy and pretrained OpenVLA without external pretraining.
By Chang Liu, Fei Suo, Yanzhou Jin, Zeyu Ping, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that replaces the Gaussian regularizer used in Joint‑Embedding Predictive Architectures (JEPAs) with a contrastive inverse‑dynamics head. AC‑MTM trains a forward latent‑prediction model while an auxiliary inverse‑dynamics task forces the encoder to distinguish actions from latent transitions, preventing collapse without requiring a target network or reconstruction loss. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM matches or surpasses the performance of the Gaussian‑based SIGReg regularizer, achieving up to 20–24 point improvements on the OGBench Visual Scene benchmark.
By Jack Boylan, Chris Hokamp
The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to improve goal‑conditioned robotic planning. By preventing latent collapse and grounding representations in physical configuration, the model achieves top success rates on tasks such as TwoRoom, PushT, and OGBench‑Cube, outperforming the baseline LeWorldModel. Ablation studies confirm that state alignment consistently boosts planning success over inverse dynamics alone across all four benchmark tasks.
By Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)
arXiv:2606. 23444v2 Announce Type: replace-cross Abstract: Accurate dynamics models are critical for informed decision-making in robotic systems, particularly for agile aerial vehicles operating under uncertainty.
By Pratyaksh Rao, Wancong Zhang, Randall Balestriero, Yann LeCun, Giuseppe Loianno
arXiv:2607. 27924v1 Announce Type: new Abstract: In the physical world we inhabit, space and time are fundamentally continuous.
By Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao, Xirui Kang, Sangli Teng, Koushil Sreenath, Xianyuan Zhan
Joint-embedding predictive architectures (JEPAs) learn latent dynamics for planning and avoid representation collapse by matching features to maximum-entropy distributions such as isotropic Gaussians,...
arXiv:2608.20974v1 Announce Type: cross
Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...
By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
arXiv:2606. 16076v1 Announce Type: cross Abstract: Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evolution.
By Weizhi Nie, Weichao Liu, Honglin Guo, Yuting Su