DA‑WAM is a framework that integrates predictive representation learning, action‑conditioned future modeling, and trajectory scoring into a single decision‑making objective for autonomous driving. It uses an online encoder with a stable momentum target to keep future representations aligned with the driving task, generating a distinct future latent for each trajectory candidate. A future‑latent‑conditioned scorer evaluates these latents, with expert‑matched trajectories supervised by observed futures and safety‑critical hard negatives providing additional guidance, achieving state‑of‑the‑art results on NAVSIM‑v1 and NAVSIM‑v2.
By Ruiguo Zhong, Benshan Ma, Xiaolong Chen, Lang Zhang, Mingyue Feng, Yaonong Wang, Pei Liu, Jun Ma
arXiv:2608.20974v1 Announce Type: cross
Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...
By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv:2608. 12854v1 Announce Type: cross Abstract: Autonomous driving requires planning under both semantic constraints and predictive dynamics.
By Bing Zhan, Shuyao Shang, Jiahao Gu, Shuo Lu, Yuan Xu, Zhao Wang, Yida Wang, Xueyang Zhang, Kun Zhan, Lue Fan, Zhaoxiang Zhang
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
By Chen Yang, Yuhao Wei, Ze Xu, Ziheng Zou, Shuang Liang, Delin Ouyang, Lingfeng Qi, Jie Li, Guofa Li
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling.
arXiv:2608. 14125v1 Announce Type: new Abstract: LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal.
By Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang
ForeDrive introduces a planning-relevant latent world model that is asymmetrically coupled to a Diffusion Transformer planner. The model learns multi‑horizon latent futures with a JEPA‑style world model, while planning gradients update the shared encoder and stop‑gradient routing trains the predictor with forecasting losses only. Gated visual fusion, future‑status injection, and Trajectory‑Adaptive Bias are used to guide trajectory generation without overriding current observations, achieving high performance on NAVSIM benchmarks using only front‑view images and pure imitation learning.
By Sinuo Wang, Zichong Gu, Yuhan Huang, Wenxin Wen, Xun Yang, Yiqing Zhang, Xingyu Zhang, Ningyu Che, Jie Ling, Qiankun Yu, Wei Liu, Jing Xu, Xinggang Wang
arXiv:2608. 09876v1 Announce Type: cross Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics.
By Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang
arXiv:2606.27504v2 Announce Type: replace
Abstract: World Action Models (WAMs) unify future environment prediction with action generation for autonomous driving, yet existing approaches optimize only...
By Tianze Xia, Lijun Zhou, Kaixin Xiong, Jingfeng Yao, Zhenxin Zhu, Haiyang Sun, Bing Wang, Guang Chen, Wenyu Liu, Hangjun Ye, Xinggang Wang
The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.
By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
By Hoang Nguyen, Xiaohao Xu, Xiaonan Huang