The paper introduces a new criticality metric specifically designed for vulnerable road users (VRUs) and a scenario‑independent prediction framework that applies to all traffic participants. The VRU‑centric metric improves pedestrian criticality classification by up to 50 %, while the prediction framework surpasses state‑of‑the‑art metrics by 275 %, achieving an F1‑score of 0.96 on the DeepAccident dataset. These advances enable more accurate, scenario‑agnostic safety assessments for autonomous driving systems.
By J\"org Gamerdinger, Victor Schwarzenberger, Philipp Schmid, Sven Teufel, Oliver Bringmann
LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.
By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
arXiv:2511. 14592v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns.
By Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao
arXiv:2609.05593v1 Announce Type: cross
Abstract: Autonomous robot navigation failures differ not only in categorical severity but also in the physical context in which they occur. A near-miss at low...
By Rifa Ferzana
arXiv:2609.37871v1 Announce Type: cross
Abstract: Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDr...
By Ziyi Luo, Zhe Sun, Yehao Lu, Lei Zhou, Lisheng Wu, Xuewei Li, Zequn Qin, Xi Li
DrivingBench is the first benchmark that tests general‑purpose vision‑language models on the task of driving a real Toyota Corolla around a parking‑lot cone course. The models receive live camera frames and issue steering and velocity commands, with inference latency counted as part of the challenge. In tests, only GPT‑6 Astra completed the course, while other models showed limited progress or failed to pass half the course.
By Aditya Ramabadran, Simon Mahns, Tobias Gessler
arXiv:2607. 11128v1 Announce Type: cross Abstract: Real-time driving risk assessment provides an essential basis for proactive safety by identifying and quantifying the danger of ongoing road interactions before adverse outcomes occur.
By Zhuoren Li, Yi Zhong, Weiqi Zhang, Xinrui Zhang, Lu Xiong, Chongfeng Wei, Bo Leng
arXiv:2603. 14841v3 Announce Type: replace-cross Abstract: Road crashes remain a leading cause of preventable fatalities.
By Joyjit Roy, Samaresh Kumar Singh, Sushanta Das
CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.
By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
arXiv:2606. 16313v1 Announce Type: cross Abstract: Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude.
By Qiao Sun, Weicheng Zheng, Yixin Huang, Hang Zhao
Reliable motion classification is critical for autonomous driving, as false dynamic predictions of static objects can cascade into unnecessary planner interventions. Unstable bounding box predictions can lead to spurious velocity estimates in tracking and falsely predicted trajectories.
FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with rubric-guided aggregation into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol revealed that safety compliance is the most discriminative dimension among 66 LLMs, with models showing up to 28-point differences in safety scores and recurrent failures such as safety violations under physically plausible predictions and instability in multi-step rollouts.
By Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang