The paper introduces a lightweight Fourier auxiliary head to enforce physically-informed structuring of latent states in JEPA-style world models, addressing a newly identified failure mode called physical representation laziness that hampers planning in dynamic environments. Experiments show that this auxiliary supervision improves planning success rates, enhances latent space correlations with key physical properties, and boosts data efficiency, even when the baseline model does not exhibit laziness.
By Penghao Zhu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.
By Yuntian Gao, Xiangyu Xu
Subspace-Decomposed JEPAs (SD-JEPA) split the latent space of Joint-Embedding Predictive Architectures into two orthogonal subspaces: a low-dimensional progression subspace trained with a cosine-margin triplet loss and a high-dimensional content subspace regularised by SIGReg. The authors prove that the anti-collapse forces act on disjoint coordinates, allowing additive composition rather than competition. SD-JEPA outperforms the LeWM baseline on most control benchmarks and the strongest non-LeWM JEPA baseline on Push‑T, with a subspace-ablation confirming the split as essential. The 1‑D angular progression coordinate serves as a scene-aware compass, advancing with task progress, regressing on backtracking, and relocalising under perturbations to separate surprise from meaning.
By Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet
arXiv:2610.01942v1 Announce Type: new
Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
By Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
arXiv:2607. 26924v1 Announce Type: new Abstract: Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse.
By Chang Liu, Fei Suo, Yanzhou Jin, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. While effective and elegant in single-task settings, this recipe does not extend reliably to multi-task training, leading to substantially worse downstream behavior-cloning performance.