arXiv Machine Learning By Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Read the original on arXiv Machine Learning →

arXiv:2608. 12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models

JEPA-Bisim introduces a bisimulation encoder to joint-embedding predictive world models, ensuring that states with similar transition dynamics are mapped to nearby latent representations while suppressing irrelevant slow features such as background changes and distractors. The approach improves robustness on navigation (PointMaze) and manipulation (PushT) tasks under varied test-time visual conditions, achieving up to tenfold smaller latent spaces than DINO-WM. It remains effective across different pre-trained visual encoders, including DINOv2, SimDINOv2, and iBOT.

By Leonardo F. Toso, Davit Shadunts, Yunyang Lu, Gloria Geng, Nihal Sharma, Donglin Zhan, Nam H. Nguyen, James Anderson
arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv Machine Learning
Jun 26

Fast LeWorldModel

arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.

By Yuntian Gao, Xiangyu Xu
arXiv AI
Sep 17

PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation

PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.

By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding