arXiv Machine Learning By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

Read the original on arXiv Machine Learning →

The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 27

On the Identifiability of Controlled World Models

arXiv:2607. 22430v1 Announce Type: new Abstract: Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control.

By Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
arXiv AI
Sep 1

Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models

Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.

By Yanchen Huo, Ziying Song, Yadan Luo
arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi