arXiv AI

FF-JEPA: Long-Horizon Planning in World Models with Latent Planners

arXiv:2606. 09311v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM).

arXiv AI
Sep 3

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?

The paper investigates Joint-Embedding Predictive World Models (JEPA-WMs), a class of methods that perform planning in a learned representation space rather than raw input space. It systematically studies how model architecture, training objectives, and planning algorithms influence success across simulated and real‑world robotic tasks, and proposes a JEPA-WM variant that surpasses established baselines in navigation and manipulation. The authors provide code, data, and checkpoints for reproducibility.

By Basile Terver, Tsung-Yen Yang, Jean Ponce, Adrien Bardes, Yann LeCun
arXiv AI
4d ago

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv Machine Learning
4d ago

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

FlexiWorld is a JEPA-based latent world model that learns variable‑length action chunks across multiple time scales for goal‑directed planning. It jointly trains a causal action encoder and an autoregressive actor, using mixed‑span goal supervision and Student Forcing to reduce exposure bias. In experiments on four benchmarks, FlexiWorld with the Actor‑Residual Cross‑Entropy Method (ARCEM) achieves higher mean success rates than the strongest baseline and supports flexible planning chunk lengths without retraining.

By Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu
arXiv Machine Learning
Jun 26

Fast LeWorldModel

arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.

By Yuntian Gao, Xiangyu Xu
Hugging Face Trending Papers
Sep 3

Latent Energy Action Planning with World Models

Latent Energy Action Planning (LEAP) improves model predictive control by treating the entire action horizon as a differentiable variable and optimizing it using a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal-window state energy, ensuring both the predicted terminal latent and the decoder-predicted terminal descriptor align with the goal. In four control domains, LEAP raises mean success from 77.5% (LeWM+CEM) to 94.8%, a 17.3‑percentage‑point improvement while keeping the frozen LeWM representation.

arXiv AI
Jul 1

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.

By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv AI
Sep 4

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to improve goal‑conditioned robotic planning. By preventing latent collapse and grounding representations in physical configuration, the model achieves top success rates on tasks such as TwoRoom, PushT, and OGBench‑Cube, outperforming the baseline LeWorldModel. Ablation studies confirm that state alignment consistently boosts planning success over inverse dynamics alone across all four benchmark tasks.

By Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)