arXiv AI

AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports

AccidentSim is a framework that generates physically realistic vehicle collision videos by extracting physical clues from real-world accident reports. It uses a reliable physical simulator to replicate post-collision trajectories, builds a trajectory dataset, fine‑tunes a language model to predict consistent trajectories from user prompts, and finally renders high‑quality videos with Neural Radiance Fields. The resulting videos show strong visual and physical authenticity compared to existing methods.

arXiv AI
Jul 10

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.

By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita
arXiv AI
Sep 18

LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality

The paper introduces an automated pipeline that converts non‑critical driving scenes into safety‑critical scenarios by integrating computer vision, Large Language Models (LLMs), and Augmented Reality (AR). It detects and tracks road users, extracts safety features such as distance, velocity, motion direction, and Time‑to‑Collision (TTC), and evaluates scene criticality. Safe scenes are then modified by an LLM, which generates realistic collision‑inducing objects and behaviors that are overlaid onto the original scene using AR, achieving 97.52% safety classification accuracy on the nuScenes dataset and producing realistic scenarios like pedestrian crossings, rear overtaking vehicles, and sudden‑stop events.

By Noura Fady, Farah Khaled, Catherine M. Elias
arXiv AI
Sep 3

CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation

CrashDiffuser is a closed-loop VLM‑guided diffusion framework designed for fine‑grained safety‑critical traffic scenario generation. It separates semantic collision reasoning from trajectory synthesis via a hierarchical collision‑intent interface that specifies target contact regions (head, rear, or side). The system uses a vision‑language model to extract scene context and predict structured action tuples, which condition a diffusion model to produce executable adversarial trajectories, achieving high target‑collision and contact‑region control rates on WOMD‑derived scenarios.

By Shucheng Zhang, Yuang Zhang, Bingzhang Wang, Muhammad Monjurul Karim, Kehua Chen, Yinhai Wang
arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
Hugging Face Trending Papers
Jul 8

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving data, and existing trajectory and language-augmented datasets seldom provide high-risk event labels, semantic annotations, and verifiable safety signals.