arXiv Machine Learning

A Generalization Theory for JEPA-Based World Models

arXiv:2606. 27014v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input level.

arXiv Machine Learning
1d ago

Learning Commute-Time-Preserving World Models for Planning

The paper introduces Commute-Time-Preserving World Models (CTWMs), which learn latent representations that reflect commute-times in an environment by using a latent displacement predictor and a log-determinant regularizer. This approach addresses the issue that existing self-supervised methods degrade the necessary eigenvalue-dependent scaling for accurate commute-time representation. In experiments, CTWMs outperform the task-agnostic baseline LeWM on several continuous goal-reaching benchmarks while using only half the parameters.

By Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke
arXiv AI
4d ago

ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning

The paper introduces ATLAS, a training objective that preserves relational geometry in latent world models while calibrating the global latent distribution. By transferring normalized pairwise structure from an informative encoder to the planning latent and applying Wasserstein embedding matching, ATLAS improves goal‑reaching success on tasks such as PushT, TwoRoom, and OGBench‑Cube, especially on higher‑novelty episodes. Diagnostics show stronger novelty‑related structure, better marginal calibration, and lower multi‑step prediction error in the planning latent.

By Ke Fang, Yupu Yao, Lu Cheng
Hugging Face Trending Papers
Aug 20

Orthogonal JEPA: Factorized Predictive States for Latent World Models

Orthogonal JEPA introduces a latent world‑modeling framework that factorizes predictive states into orthogonal components. By learning basis matrices and dedicated prediction branches, the method reduces redundancy and improves gradient signals for less dominant predictive structures. The factorized states can be synthesized into complete latent representations for downstream tasks such as decoding, planning, or autoregressive rollout, and are evaluated across vision, biology, health, control, and molecular dynamics domains.

arXiv Machine Learning
5d ago

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.

By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv AI
Sep 1

Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models

Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.

By Yanchen Huo, Ziying Song, Yadan Luo
arXiv Machine Learning
Aug 27

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.

By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White