arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.
By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita
arXiv:2510. 03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastructure-based measures are often insufficient in dynamic urban environments.
By Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Muhammad Monjurul Karim, Kehua Chen, Chenxi Liu, Mehrdad Nasri, Yinhai Wang
arXiv:2511. 14592v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns.
By Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao
LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.
By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
arXiv:2511.20022v3 Announce Type: replace-cross
Abstract: Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their...
By Seungjun Yu, Seonho Lee, Namho Kim, Jaeyo Shin, Junsung Park, Wonjeong Ryu, Raehyuk Jung, Hyunjung Shim
AccidentSim is a framework that generates physically realistic vehicle collision videos by extracting physical clues from real-world accident reports. It uses a reliable physical simulator to replicate post-collision trajectories, builds a trajectory dataset, fine‑tunes a language model to predict consistent trajectories from user prompts, and finally renders high‑quality videos with Neural Radiance Fields. The resulting videos show strong visual and physical authenticity compared to existing methods.
By Xiangwen Zhang, Qian Zhang, Longfei Han, Qiang Qu, Xiaoming Chen, Weidong Cai
HERMES is a holistic end‑to‑end multimodal driving framework that incorporates long‑tail semantic knowledge into trajectory planning for autonomous vehicles. It uses a foundation‑model‑assisted annotation pipeline to build Long‑Tail Scene Context and Long‑Tail Planning Context, capturing hazard‑centric scene information, maneuver intent, and risk‑aware guidance. A Tri‑Modal Driving Module then fuses multi‑view visual observations, historical ego‑motion, and long‑tail semantic instructions to generate intent‑ and risk‑aware trajectories, achieving consistent performance gains on a large‑scale real‑world long‑tail driving benchmark.
By Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran
arXiv:2607. 07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety.
By Heye Huang, Jingguang Li, Zhiyuan Zhou, Paul Liang, Mingyu Wu, Kitae Jang, Jianqiang Wang
arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.
By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
The paper introduces a new criticality metric specifically designed for vulnerable road users (VRUs) and a scenario‑independent prediction framework that applies to all traffic participants. The VRU‑centric metric improves pedestrian criticality classification by up to 50 %, while the prediction framework surpasses state‑of‑the‑art metrics by 275 %, achieving an F1‑score of 0.96 on the DeepAccident dataset. These advances enable more accurate, scenario‑agnostic safety assessments for autonomous driving systems.
By J\"org Gamerdinger, Victor Schwarzenberger, Philipp Schmid, Sven Teufel, Oliver Bringmann
Qwen-Drive-1.0 is a vision‑language foundation model tailored for autonomous driving that builds on a pretrained VLM architecture. It incorporates a bird’s‑eye‑view perception head for 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert that generates future ego trajectories from shared representations. Experiments show strong 3D perception, driving scene understanding, and competitive motion‑planning performance while largely preserving general vision‑language capabilities.
By Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
arXiv:2603. 29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment.
By Qiucheng Yu, Ruijie Xu, Mingang Chen Jianfeng Dong, Xin Tan