arXiv:2608.29029v3 Announce Type: replace-cross
Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free...
By Yanchen Huo, Ziying Song, Yadan Luo
The paper introduces Commute-Time-Preserving World Models (CTWMs), which learn latent representations that reflect commute-times in an environment by using a latent displacement predictor and a log-determinant regularizer. This approach addresses the issue that existing self-supervised methods degrade the necessary eigenvalue-dependent scaling for accurate commute-time representation. In experiments, CTWMs outperform the task-agnostic baseline LeWM on several continuous goal-reaching benchmarks while using only half the parameters.
By Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke
The paper introduces ATLAS, a training objective that preserves relational geometry in latent world models while calibrating the global latent distribution. By transferring normalized pairwise structure from an informative encoder to the planning latent and applying Wasserstein embedding matching, ATLAS improves goal‑reaching success on tasks such as PushT, TwoRoom, and OGBench‑Cube, especially on higher‑novelty episodes. Diagnostics show stronger novelty‑related structure, better marginal calibration, and lower multi‑step prediction error in the planning latent.
By Ke Fang, Yupu Yao, Lu Cheng
Orthogonal JEPA introduces a latent world‑modeling framework that factorizes predictive states into orthogonal components. By learning basis matrices and dedicated prediction branches, the method reduces redundancy and improves gradient signals for less dominant predictive structures. The factorized states can be synthesized into complete latent representations for downstream tasks such as decoding, planning, or autoregressive rollout, and are evaluated across vision, biology, health, control, and molecular dynamics domains.
arXiv:2608. 20065v1 Announce Type: new Abstract: World models construct latent states that support prediction, planning, and reasoning about an underlying system.
By Taoyong Cui, Pheng Ann Heng, Wanli Ouyang
arXiv:2607. 23337v1 Announce Type: new Abstract: Neural operators provide data-driven mappings for modeling dynamical systems.
By Zituo Chen, Qiaofeng Li, Jiaxin Hu, Sili Deng
arXiv:2606. 28712v1 Announce Type: cross Abstract: Classical SLAM estimates metric poses and a geometric map but produces no actionable predictive model for planning.
By Guanqun Cao, Liang Chen
arXiv:2606.28712v2 Announce Type: replace-cross
Abstract: Classical simultaneous localization and mapping (SLAM) estimates metric poses and a geometric map but does not provide an action-conditioned...
By Guanqun Cao, Liang Chen
The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.
By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.
By Yanchen Huo, Ziying Song, Yadan Luo
arXiv:2606. 09311v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM).
By Sergi Masip, Jonathan Swinnen, Yutong Hu, Renaud Detry, Tinne Tuytelaars
The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.
By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White