arXiv Computer Vision

TrafficSignBench: Rule-Centric Closed-Loop Evaluation of Traffic-Sign Compliance in Autonomous Driving

arXiv AI
Jun 8

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

arXiv:2602. 16073v2 Announce Type: replace-cross Abstract: Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, obeying traffic rules, and making efficient progress.

By Kevin Kai-Chun Chang, Ekin Beyazit, Alberto Sangiovanni-Vincentelli, Tichakorn Wongpiromsarn, Sanjit A. Seshia
Hugging Face Trending Papers
Aug 18

Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

The paper introduces a plug‑and‑play method that injects traffic‑element signals—such as traffic lights and road signs—into end‑to‑end autonomous driving models with minimal architectural changes. By augmenting several public datasets with comprehensive traffic‑element annotations, the authors evaluate this integration across diverse driving paradigms, consistently improving performance on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. The approach achieves a new state‑of‑the‑art result on the challenging NAVSIM‑v2 benchmark, demonstrating the broad utility of traffic‑element awareness.

arXiv AI
Jun 16

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

arXiv:2606. 15749v1 Announce Type: cross Abstract: Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics.

By Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen
arXiv AI
Jun 17

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

arXiv:2606. 17362v1 Announce Type: cross Abstract: Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-dependent.

By Xinglong Sun, Kevin Xie, Jenny Schmalfuss, Despoina Paschalidou, Xiuming Zhang, Sanja Fidler, Kashyap Chitta, Jose M. Alvarez
arXiv AI
Jun 16

Driving, Fast or Slow? Neuro-Symbolic Guidance for Motion Prediction in Multi-Modal Ground Mobility

arXiv:2606. 15251v1 Announce Type: cross Abstract: Accurate and interpretable motion prediction for heterogeneous traffic spaces, including pedestrians, bicycles, cars, and trucks, is essential for safe autonomous navigation.

By Simon Kohaut, Felix Divo, Julius Hahnewald, Benedict Flade, Julian Eggert, Kristian Kersting, Devendra Singh Dhami
arXiv AI
Jul 21

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

arXiv:2511. 14592v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns.

By Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao
arXiv AI
Jun 4

MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation

arXiv:2606. 04513v1 Announce Type: new Abstract: Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized lane networks for hundreds of cities remains highly labor-intensive.

By Deguo Xia, Zihan Li, Haochen Zhao, Dong Xie, Yuyao Kong, Xiyan Liu, Jizhou Huang, Mengmeng Yang, Diange Yang
arXiv AI
Sep 3

CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation

CrashDiffuser is a closed-loop VLM‑guided diffusion framework designed for fine‑grained safety‑critical traffic scenario generation. It separates semantic collision reasoning from trajectory synthesis via a hierarchical collision‑intent interface that specifies target contact regions (head, rear, or side). The system uses a vision‑language model to extract scene context and predict structured action tuples, which condition a diffusion model to produce executable adversarial trajectories, achieving high target‑collision and contact‑region control rates on WOMD‑derived scenarios.

By Shucheng Zhang, Yuang Zhang, Bingzhang Wang, Muhammad Monjurul Karim, Kehua Chen, Yinhai Wang