arXiv Computer Vision By Shuai Liu, Hechangle Gong, Hao Jiang, Runlin He, Junxiang Zhan, Kai Huang, Sheng Yang, Shaoqing Ren

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

Read the original on arXiv Computer Vision →

MM-Future is a world-action model for autonomous driving that generates multiple paired scene-action hypotheses and captures bidirectional interaction within each pair. It initializes each hypothesis from a structured action prior and an independent future scene source, then co-evolves them using a modality-aware diffusion Transformer. The model compresses multi-view video into planning-oriented MM-Tokens and uses a future-conditioned proposal scorer to rank trajectory candidates, achieving strong performance on NAVSIM and HUGSIM benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 24

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

arXiv:2608.20974v1 Announce Type: cross Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...

By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
arXiv Computer Vision
Sep 4

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Drive‑HWM introduces a hierarchical slow‑fast world modeling framework for autonomous driving. The slow model predicts multi‑step future representations, while the fast model jointly predicts the next frame and immediate action using a lightweight multimodal backbone and an autoregressive expert. Dynamic‑Aware Latents, learned through optical‑flow prediction, explicitly capture motion dynamics, and experiments on NAVSIM v1 and v2 show strong driving performance with validated ablation studies.

By Zhaoxin Fan, Tianbao Zhang, Wenjun Wu, Xiaofeng Wang, Yeying Jin, Jian Zhao, Zheng Zhu, Shuicheng Yan
arXiv Computer Vision
2d ago

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

PhysWAM is a unified world-action model for autonomous driving that jointly denoises multiview video, metric depth, and ego motion using a flow‑matching transformer. It introduces Coupled Point Projection (CPP), a geometric constraint that aligns generated depth points with LiDAR data after applying the predicted SE(3) ego motion, thereby enforcing physical consistency. At inference, trajectory selection uses a simple label‑free consensus rule, and the model demonstrates strong planning performance, zero‑shot transfer to unseen environments, and accurate, temporally coherent depth and video predictions.

By Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang
arXiv AI
Aug 20

DA-WAM: Decision-Aligned Future Latents for Driving World Models

DA‑WAM is a framework that integrates predictive representation learning, action‑conditioned future modeling, and trajectory scoring into a single decision‑making objective for autonomous driving. It uses an online encoder with a stable momentum target to keep future representations aligned with the driving task, generating a distinct future latent for each trajectory candidate. A future‑latent‑conditioned scorer evaluates these latents, with expert‑matched trajectories supervised by observed futures and safety‑critical hard negatives providing additional guidance, achieving state‑of‑the‑art results on NAVSIM‑v1 and NAVSIM‑v2.

By Ruiguo Zhong, Benshan Ma, Xiaolong Chen, Lang Zhang, Mingyue Feng, Yaonong Wang, Pei Liu, Jun Ma