The paper introduces a layered evaluation protocol for generative scenario models used in autonomous driving, focusing on physical consistency and plausibility. It examines internal representations through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis, and then tests outputs against vehicle dynamics constraints such as lateral jerk thresholds. The protocol is applied to a VAE-based scenario generator and other generative models, revealing deeper insights than standard output-level metrics.
By Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner
arXiv:2606. 03159v1 Announce Type: cross Abstract: As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck.
By NVIDIA, :, Aarti Basant, Amlan Kar, Despoina Paschalidou, Fangyin Wei, Francesco Ferroni, Guillermo Garcia Cobo, Haithem Turki, Huan Ling, Jaewoo Seo, James Lucas, Jay Zhangjie Wu, Jialiang Wang, Jonathan Lorraine, Jun Gao, Kai He, Katarina Tothova, Kevin Xie, Micha{\l} Tyszkiewicz, Qi Wu, Riccardo de Lutio, Ruilong Li, Sanja Fidler, Seung Wook Kim, Tianchang Shen, Tianshi Cao, Tobias Pfaff, William Lew, Xindi Wu, Xuanchi Ren, Yifan Lu, Yuxuan Zhang, Zan Gojcic, Zian Wang
As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions dynamically update the simulator state and directly influence the next set of generated sensor observations.
arXiv:2512. 08991v3 Announce Type: replace-cross Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems.
By Yuang Geng, Zhongzheng Zhang, Chengzhen Jiang, Yanru Li, Xinyang Wang, Zhuoyang Zhou, Hoang-Dung Tran, Ivan Ruchkin
arXiv:2409. 16663v5 Announce Type: replace-cross Abstract: We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving.
By Alexander Popov, Alperen Degirmenci, David Wehr, Shashank Hegde, Ryan Oldja, Alexey Kamenev, Bertrand Douillard, David Nist\'er, Urs Muller, Ruchi Bhargava, Stan Birchfield, Nikolai Smolyanskiy
arXiv:2609.36851v1 Announce Type: new
Abstract: End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their...
By Hongbin Lin, Chaoda Zheng, Yiming Yang, Xiangyu Li, Shijia Chen, Jinhao Deng, Kangjie Chen, Dongbin Zhang, Jie Feng, Yu Zhang, Xianming Liu, Shuguang Cui, Boyang Wang, Zhen Li
Contrastive World Models propose a new method for learning latent dynamics without pixel reconstruction. By replacing observation reconstruction with a Deep InfoMax-like objective that maximizes mutual information between state-action sequences and local patch features of future observations, the approach encourages state representations to retain predictive information while ignoring visually irrelevant details. Experiments show that this method matches existing baselines in simple settings and significantly outperforms them when distractors or natural video backgrounds are present, while also training more efficiently by eliminating the pixel decoder.
By Bonnie Li
arXiv:2609.22762v1 Announce Type: new
Abstract: Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expe...
By Fengcheng Yu, Dhruv Parikh, Junjie Ye, Maulik Bhatt, Thang Vu, Igor Vasiljevic, Vitor Guizilini, Yue Wang
SV-WAM is a surround‑view world‑action model that keeps all six camera views while enabling efficient inference by discarding the video branch at deployment. It uses future‑video prediction as dense training supervision and an action‑centered causal mask to prevent action tokens from attending to future‑video tokens during joint denoising. A differentiable drivable‑area compliance regularizer penalizes vehicle‑footprint corners near or crossing drivable boundaries, improving safety and boundary awareness. Experiments on NAVSIMv2 and nuScenes show state‑of‑the‑art planning performance with low latency and strong zero‑shot transfer.
By Jinyang Wang, Shiwei Li, Junjian Wang, Zhiqiang Deng, Jianbin Gao, Yihang Zhao, Liu Liu, Yongjia Zhao, Jinlong Chen, Huirui Xu, Yifeng Pan, Kangwei Liu, Fan Ren, Ji Tao, Minghao Yang
arXiv:2609.26792v1 Announce Type: cross
Abstract: Faithfully evaluating end-to-end driving policies in simulation requires observations that are not merely photo-realistic, but preserve the scene fea...
By Ziyang Leng, Sicheng Mo, Seth Z. Zhao, Haoyuan Cai, Yu Zeng, Rowan McAllister, Bolei Zhou
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2506. 06006v3 Announce Type: replace-cross Abstract: Can unified vision-language models (VLMs) perform forward dynamics prediction (FDP), i.
By Yifu Qiu, Yftah Ziser, Anna Korhonen, Shay B. Cohen, Edoardo M. Ponti