DeepJEPA is a weight‑tied joint‑embedding predictive world model that treats transition depth as an inner test‑time scaling axis, learning when additional recurrent updates are worthwhile for each candidate and rollout step. Unlike traditional planners that uniformly deepen every transition, DeepJEPA concentrates extra computation on decision‑critical events such as contact onset and sustained object interaction, achieving comparable or better performance with only 1.00–1.26 updates per transition across five visual‑control settings. The approach demonstrates that improved planning does not require uniformly better object‑state decodability, but rather targeted internal computation where it can alter the planner’s elite set and action selection.
By Zijian Jin, Yunbei Zhang, Yuanzhe Liu, Ming Liu, Baian Chen, Weirui Ye, Shilong Liu, Marco Pavone
arXiv:2609.10506v1 Announce Type: cross
Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...
By Nisarga Nilavadi, Ralf R\"omer, Moritz Reuss, Michael Krawez, Tobias J\"ulg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and ro...
arXiv:2608.29434v1 Announce Type: cross
Abstract: JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction...
By Fabio F. Oberweger, Michael Schwingshackl
arXiv:2609.39235v1 Announce Type: cross
Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
By Ali Alrasheed, Basim Azam, Naveed Akhtar
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.
By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding
arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.
By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.
By Yuntian Gao, Xiangyu Xu
JEPA‑TTT is a method that continuously adapts the latent dynamics predictor of a pretrained Joint‑Embedding Predictive Architecture (JEPA) world model during test time. It performs self‑supervised updates across episodes while keeping the visual encoder and reward head fixed, using dense replay to sample prediction windows from a growing buffer. In experiments on eight dynamics shifts across four continuous‑control environments, JEPA‑TTT reduces latent prediction error by 83% and improves planning performance by 153% compared to the frozen model.
By Zheyuan Zhang, Suyu Ye, Nakul Agarwal, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Daniel Khashabi, Tianmin Shu, Vaishnav Tadiparthi
The paper introduces UniWM, a unified, memory‑augmented world model that merges egocentric visual foresight and planning into a single multimodal autoregressive backbone. By grounding action selection in visually imagined outcomes and using a hierarchical memory to fuse short‑term perception with long‑term trajectory context, UniWM aligns prediction with control and improves navigation stability. Experiments on four challenging benchmarks and the 1X Humanoid Dataset show up to 30% higher success rates, reduced trajectory errors, zero‑shot generalization to unseen datasets, and scalability to high‑dimensional humanoid navigation.
By Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng
arXiv:2608. 12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance.
By Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian