arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.
By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
arXiv:2606. 30068v1 Announce Type: new Abstract: Joint-embedding predictive (JEPA-style) objectives learn representations by predicting future latents.
By Ayan Pendharkar
JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv:2609.13845v1 Announce Type: cross
Abstract: World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet pl...
By Saksham Bansal, Om Naphade, Chayan Aggarwal, Vrishin M
The paper introduces the Latent Generative Solver (LGS), a neural PDE solver that combines a Physics VAE, a Pyramidal Flow-Forcing Transformer, and input noising to achieve generalization across twelve PDE families and stable long-term rollouts. LGS matches or surpasses deterministic baselines on one-step predictions, outperforms them on 5- and 10-step rollouts, and significantly reduces long-horizon error while cutting compute costs. It also adapts efficiently to unseen higher-resolution systems, demonstrating strong empirical performance on 2D regular-grid PDE simulations.
By Zituo Chen, Sili Deng
HaM-World introduces a structured world model that combines history-conditioned selective memory with a Soft‑Hamiltonian latent dynamics prior. The model decomposes the latent state into a canonical (q,p) subspace governed by an energy‑derived Hamiltonian vector field and a context subspace c capturing non‑conservative factors, while Mamba selective state‑space memory conditions the transition used for prediction, reward, value estimation, and planning. Across six DeepMind Control Suite tasks, HaM-World achieves top rankings on four tasks, improves average AUC, reduces imagined‑rollout error by 45% on short‑to‑medium horizons, and outperforms baselines under 12 out‑of‑distribution perturbations.
By Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang
arXiv:2609.10464v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning,...
By Andy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero, John Sous
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
By Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2609.39888v2 Announce Type: new
Abstract: Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls a...
By Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis
arXiv:2608.24044v1 Announce Type: new
Abstract: Latent world models plan by predicting how candidate actions transform learned representations. In self-predictive models, however, the encoder and pre...
By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.