End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems remains largely opaque, due to the complexity of traffic scenes.
TrafficImag is the first benchmark designed to evaluate counterfactual roadside traffic video generation, combining a large roadside dataset with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is encoded as an actor-level program specifying target actor, intended behavior, legal route, interaction order, and temporal constraints, allowing a unified evaluation across diverse foundation models. The benchmark assesses four validity dimensions—initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation—and reports that the best models achieve 80.4% macro F1 for reasoning and 55.0% end-to-end success when using a complete condition interface.
By Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.
By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
arXiv:2608. 19380v1 Announce Type: new Abstract: While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpret traffic accidents.
By Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich
arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.
By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy
arXiv:2606. 17362v1 Announce Type: cross Abstract: Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent.
By Xinglong Sun, Kevin Xie, Jenny Schmalfuss, Despoina Paschalidou, Xiuming Zhang, Sanja Fidler, Kashyap Chitta, Jose M. Alvarez
The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.
By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations.
CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.
By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
arXiv:2511.20022v3 Announce Type: replace-cross
Abstract: Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their...
By Seungjun Yu, Seonho Lee, Namho Kim, Jaeyo Shin, Junsung Park, Wonjeong Ryu, Raehyuk Jung, Hyunjung Shim
arXiv:2605. 06264v2 Announce Type: replace Abstract: End-to-end autonomous driving models generate future trajectories from multi-view inputs, improving system integration but introducing opaque decisions and hard-to-localize risks.
By Le Yang, Haijun Liu, Jiawei Liang, ShangQuan Sun, Xiaochun Cao
arXiv:2607. 00283v1 Announce Type: cross Abstract: Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view.
By Amirhosein Chahe, Tyler Naes, Jovin D'sa, Faizan M. Tariq, Sangjae Bae, Lifeng Zhou, David Isele