arXiv Machine Learning

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

arXiv:2607. 07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety.

Hugging Face Trending Papers
Jul 8

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving data, and existing trajectory and language-augmented datasets seldom provide high-risk event labels, semantic annotations, and verifiable safety signals.

arXiv AI
Sep 21

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

HERMES is a holistic end‑to‑end multimodal driving framework that incorporates long‑tail semantic knowledge into trajectory planning for autonomous vehicles. It uses a foundation‑model‑assisted annotation pipeline to build Long‑Tail Scene Context and Long‑Tail Planning Context, capturing hazard‑centric scene information, maneuver intent, and risk‑aware guidance. A Tri‑Modal Driving Module then fuses multi‑view visual observations, historical ego‑motion, and long‑tail semantic instructions to generate intent‑ and risk‑aware trajectories, achieving consistent performance gains on a large‑scale real‑world long‑tail driving benchmark.

By Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran
arXiv Machine Learning
Sep 3

PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems

PRISM (Proactive Risk Intelligence and Safety Management) is an agentic multi-model architecture designed to shift autonomous transportation safety from reactive crash avoidance to proactive, continuous risk management. It uses inverse crash‑probability modeling to transform binary crash classifiers into dynamic safety scores, and runs three specialized models—trajectory kinematics, environmental risk, and VRU interaction—coordinated by a reinforcement‑learning reasoning layer. Across 1,296 naturalistic driving scenarios, PRISM achieved a mean safety score of 68/100, classified 77.6% of situations as advisory, and flagged 3.8% as near‑misses, with 11% requiring intervention or emergency response, highlighting trajectory risk and VRU proximity as key safety factors.

By Joyjit Roy, Samaresh Kumar Singh, Sushanta Das
arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
arXiv AI
Jul 10

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.

By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita