arXiv AI By Sergi Masip, Jonathan Swinnen, Yutong Hu, Renaud Detry, Tinne Tuytelaars

FF-JEPA: Long-Horizon Planning in World Models with Latent Planners

Read the original on arXiv AI →

arXiv:2606. 09311v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?

The paper investigates Joint-Embedding Predictive World Models (JEPA-WMs), a class of methods that perform planning in a learned representation space rather than raw input space. It systematically studies how model architecture, training objectives, and planning algorithms influence success across simulated and real‑world robotic tasks, and proposes a JEPA-WM variant that surpasses established baselines in navigation and manipulation. The authors provide code, data, and checkpoints for reproducibility.

By Basile Terver, Tsung-Yen Yang, Jean Ponce, Adrien Bardes, Yann LeCun
arXiv AI
4d ago

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv Machine Learning
4d ago

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

FlexiWorld is a JEPA-based latent world model that learns variable‑length action chunks across multiple time scales for goal‑directed planning. It jointly trains a causal action encoder and an autoregressive actor, using mixed‑span goal supervision and Student Forcing to reduce exposure bias. In experiments on four benchmarks, FlexiWorld with the Actor‑Residual Cross‑Entropy Method (ARCEM) achieves higher mean success rates than the strongest baseline and supports flexible planning chunk lengths without retraining.

By Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu