arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.
By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv:2607. 11270v1 Announce Type: cross Abstract: Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities.
By Peijun Tang, Shangjin Xie, Baifu Huang, Binyan Sun, Haotian Yang, Kuncheng Luo, Weiqi Jin, Shilin Fang, Jianan Wang
arXiv:2606.27504v2 Announce Type: replace
Abstract: World Action Models (WAMs) unify future environment prediction with action generation for autonomous driving, yet existing approaches optimize only...
By Tianze Xia, Lijun Zhou, Kaixin Xiong, Jingfeng Yao, Zhenxin Zhu, Haiyang Sun, Bing Wang, Guang Chen, Wenyu Liu, Hangjun Ye, Xinggang Wang
arXiv:2605. 00412v3 Announce Type: replace Abstract: World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning.
By Sen Cui, Jingheng Ma
The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to better support goal‑conditioned robotic planning. By incorporating inverse dynamics, the model prevents latent collapse and encodes action information, while state alignment ties consecutive latent states to their physical configurations and motions. Experiments on four benchmark tasks show the model achieves top success rates on TwoRoom, PushT, and OGBench‑Cube, and its state alignment consistently improves planning performance over inverse dynamics alone.
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
By Wancong Zhang, Basile Terver, Artem Zholus, Soham Chitnis, Harsh Sutaria, Mido Assran, Randall Balestriero, Amir Bar, Adrien Bardes, Yann LeCun, Nicolas Ballas
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to improve goal‑conditioned robotic planning. By preventing latent collapse and grounding representations in physical configuration, the model achieves top success rates on tasks such as TwoRoom, PushT, and OGBench‑Cube, outperforming the baseline LeWorldModel. Ablation studies confirm that state alignment consistently boosts planning success over inverse dynamics alone across all four benchmark tasks.
By Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
The paper introduces Commute-Time-Preserving World Models (CTWMs), which learn latent representations that reflect commute-times in an environment by using a latent displacement predictor and a log-determinant regularizer. This approach addresses the issue that existing self-supervised methods degrade the necessary eigenvalue-dependent scaling for accurate commute-time representation. In experiments, CTWMs outperform the task-agnostic baseline LeWM on several continuous goal-reaching benchmarks while using only half the parameters.
By Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.
arXiv:2609.23881v1 Announce Type: new
Abstract: Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction....
By Markus Karmann, Shile Li, Christian Intern\`o, Bruno Andreis, David Klindt, Randall Balestriero, Jindong Gu, Philip Torr, Qi Zhang, Peng-Tao Jiang, Hao Zhang, Bo Li, Onay Urfalioglu