The paper introduces a lightweight Fourier auxiliary head to enforce physically-informed structuring of latent states in JEPA-style world models, addressing a newly identified failure mode called physical representation laziness that hampers planning in dynamic environments. Experiments show that this auxiliary supervision improves planning success rates, enhances latent space correlations with key physical properties, and boosts data efficiency, even when the baseline model does not exhibit laziness.
By Penghao Zhu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.
By Yuntian Gao, Xiangyu Xu
Subspace-Decomposed JEPAs (SD-JEPA) split the latent space of Joint-Embedding Predictive Architectures into two orthogonal subspaces: a low-dimensional progression subspace trained with a cosine-margin triplet loss and a high-dimensional content subspace regularised by SIGReg. The authors prove that the anti-collapse forces act on disjoint coordinates, allowing additive composition rather than competition. SD-JEPA outperforms the LeWM baseline on most control benchmarks and the strongest non-LeWM JEPA baseline on Push‑T, with a subspace-ablation confirming the split as essential. The 1‑D angular progression coordinate serves as a scene-aware compass, advancing with task progress, regressing on backtracking, and relocalising under perturbations to separate surprise from meaning.
By Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet
arXiv:2610.01942v1 Announce Type: new
Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
By Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
arXiv:2607. 26924v1 Announce Type: new Abstract: Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse.
By Chang Liu, Fei Suo, Yanzhou Jin, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. While effective and elegant in single-task settings, this recipe does not extend reliably to multi-task training, leading to substantially worse downstream behavior-cloning performance.
arXiv:2608. 07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions.
By Xinyi Li, Zaishuo Xia, Chenjie Hao, Yubei Chen
arXiv:2601. 06212v2 Announce Type: replace-cross Abstract: We present Akasha 2, a state-of-the-art multimodal architecture that integrates Hamiltonian State Space Duality (H-SSD) with Visual-Language Joint Embedding Predictive Architecture (VL-JEPA).
By Yani Meziani
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv:2608. 09876v1 Announce Type: cross Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics.
By Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang
The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.
By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
Contrastive World Models propose a new method for learning latent dynamics without pixel reconstruction. By replacing observation reconstruction with a Deep InfoMax-like objective that maximizes mutual information between state-action sequences and local patch features of future observations, the approach encourages state representations to retain predictive information while ignoring visually irrelevant details. Experiments show that this method matches existing baselines in simple settings and significantly outperforms them when distractors or natural video backgrounds are present, while also training more efficiently by eliminating the pixel decoder.
By Bonnie Li