arXiv:2610.01942v1 Announce Type: new
Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
By Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
arXiv:2608.29029v3 Announce Type: replace-cross
Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free...
By Yanchen Huo, Ziying Song, Yadan Luo
arXiv:2605.12957v2 Announce Type: replace
Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of do...
By Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen
Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.
By Yanchen Huo, Ziying Song, Yadan Luo
Variational Streaming Flow (VSF) extends the efficient Streaming Flow (SF) framework by learning a latent distribution conditioned on system dynamics, enabling probabilistic forecasting in physical time. Unlike SF’s deterministic velocity field, VSF produces multiple plausible future trajectories, improving predictive accuracy and distributional fidelity across deterministic and stochastic dynamical systems. The method supports long‑horizon rollouts over 1,000 steps, handles bifurcating dynamics, and can be integrated as a plug‑and‑play predictor into Joint‑Embedding Predictive Architecture (JEPA) world models to enhance navigation, motion planning, and manipulation tasks.
By Hans Hao-Hsun Hsu, Minseon Gwak, Soon Hoe Lim, Pan Li, N. Benjamin Erichson
arXiv:2607. 16314v1 Announce Type: cross Abstract: World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments.
By Usman M. Khan
The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.
By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
arXiv:2608.20974v1 Announce Type: cross
Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...
By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
arXiv:2608. 10544v1 Announce Type: cross Abstract: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations.
By Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling.