arXiv AI

LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality

The paper introduces an automated pipeline that converts non‑critical driving scenes into safety‑critical scenarios by integrating computer vision, Large Language Models (LLMs), and Augmented Reality (AR). It detects and tracks road users, extracts safety features such as distance, velocity, motion direction, and Time‑to‑Collision (TTC), and evaluates scene criticality. Safe scenes are then modified by an LLM, which generates realistic collision‑inducing objects and behaviors that are overlaid onto the original scene using AR, achieving 97.52% safety classification accuracy on the nuScenes dataset and producing realistic scenarios like pedestrian crossings, rear overtaking vehicles, and sudden‑stop events.

arXiv AI
Jul 10

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.

By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita
arXiv AI
Jul 28

From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

arXiv:2510. 03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastructure-based measures are often insufficient in dynamic urban environments.

By Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Muhammad Monjurul Karim, Kehua Chen, Chenxi Liu, Mehrdad Nasri, Yinhai Wang
arXiv AI
Jul 21

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

arXiv:2511. 14592v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns.

By Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao
arXiv AI
Sep 4

LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving

LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.

By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
arXiv AI
Sep 7

AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports

AccidentSim is a framework that generates physically realistic vehicle collision videos by extracting physical clues from real-world accident reports. It uses a reliable physical simulator to replicate post-collision trajectories, builds a trajectory dataset, fine‑tunes a language model to predict consistent trajectories from user prompts, and finally renders high‑quality videos with Neural Radiance Fields. The resulting videos show strong visual and physical authenticity compared to existing methods.

By Xiangwen Zhang, Qian Zhang, Longfei Han, Qiang Qu, Xiaoming Chen, Weidong Cai
arXiv AI
Sep 21

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

HERMES is a holistic end‑to‑end multimodal driving framework that incorporates long‑tail semantic knowledge into trajectory planning for autonomous vehicles. It uses a foundation‑model‑assisted annotation pipeline to build Long‑Tail Scene Context and Long‑Tail Planning Context, capturing hazard‑centric scene information, maneuver intent, and risk‑aware guidance. A Tri‑Modal Driving Module then fuses multi‑view visual observations, historical ego‑motion, and long‑tail semantic instructions to generate intent‑ and risk‑aware trajectories, achieving consistent performance gains on a large‑scale real‑world long‑tail driving benchmark.

By Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran
arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
arXiv Machine Learning
Sep 14

Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving

The paper introduces a new criticality metric specifically designed for vulnerable road users (VRUs) and a scenario‑independent prediction framework that applies to all traffic participants. The VRU‑centric metric improves pedestrian criticality classification by up to 50 %, while the prediction framework surpasses state‑of‑the‑art metrics by 275 %, achieving an F1‑score of 0.96 on the DeepAccident dataset. These advances enable more accurate, scenario‑agnostic safety assessments for autonomous driving systems.

By J\"org Gamerdinger, Victor Schwarzenberger, Philipp Schmid, Sven Teufel, Oliver Bringmann
arXiv Computer Vision
Sep 2

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Qwen-Drive-1.0 is a vision‑language foundation model tailored for autonomous driving that builds on a pretrained VLM architecture. It incorporates a bird’s‑eye‑view perception head for 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert that generates future ego trajectories from shared representations. Experiments show strong 3D perception, driving scene understanding, and competitive motion‑planning performance while largely preserving general vision‑language capabilities.

By Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai