Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.
By Yanchen Huo, Ziying Song, Yadan Luo
arXiv:2608.24855v1 Announce Type: new
Abstract: Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online traje...
By Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu, Jenq-Neng Hwang
arXiv:2608.20974v1 Announce Type: cross
Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...
By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.
By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv:2606. 29059v1 Announce Type: cross Abstract: World modeling requires forecasting uncertain futures while preserving information useful for downstream perception.
By Francois Porcher, Nicolas Carion, Karteek Alahari, Shizhe Chen
JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.
By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv:2603. 22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.
By Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu
arXiv:2607. 27924v1 Announce Type: new Abstract: In the physical world we inhabit, space and time are fundamentally continuous.
By Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao, Xirui Kang, Sangli Teng, Koushil Sreenath, Xianyuan Zhan
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv:2608.29434v1 Announce Type: cross
Abstract: JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction...
By Fabio F. Oberweger, Michael Schwingshackl
In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in capturing the dynamics of physical world.