arXiv AI

Is Your Trajectory Displacement Safe in Long-tail?

arXiv:2606. 16313v1 Announce Type: cross Abstract: Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude.

Hugging Face Trending Papers
Jul 8

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving data, and existing trajectory and language-augmented datasets seldom provide high-risk event labels, semantic annotations, and verifiable safety signals.

arXiv AI
6d ago

Auditing Latent-Space Monitors for Autonomous Driving

The paper audits runtime failure monitors that use a model’s internal representations to predict failures in autonomous driving tasks. Across two tasks—online vectorized map generation with LaneSegNet and end‑to‑end planning with VAD—the authors find that frame‑level errors can be predicted with high AUROC scores using supervised latent probes. However, adding latent features to baseline monitors that use only observable inputs and outputs does not yield statistically significant improvements, suggesting that internal representations may not provide additional predictive value beyond what is already observable.

By Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar
arXiv AI
Aug 21

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

arXiv:2604. 02022v4 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses.

By Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, Jing Shao, Xia Hu, Dongrui Liu
arXiv Computer Vision
Sep 3

Designing Versatile Samples for Learned Trajectory Scoring

The paper proposes a new training dataset that generates more informative positive and negative samples for trajectory scoring in autonomous driving. By perturbing logged human trajectories laterally toward the drivable boundary and longitudinally toward a leading vehicle, the dataset provides richer supervision than the planner’s default proposal pool. Using a transformer-based scorer trained on this dataset, the authors achieve improved EPDMS scores on two frozen planners, DiffusionDrive and MeanFuser, when evaluated on the NAVSIM navtrain dataset.

By Yaguang Li, Jiaru Zhang, Chuheng Wei, Can Cui, Ziran Wang
arXiv AI
Sep 21

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

HERMES is a holistic end‑to‑end multimodal driving framework that incorporates long‑tail semantic knowledge into trajectory planning for autonomous vehicles. It uses a foundation‑model‑assisted annotation pipeline to build Long‑Tail Scene Context and Long‑Tail Planning Context, capturing hazard‑centric scene information, maneuver intent, and risk‑aware guidance. A Tri‑Modal Driving Module then fuses multi‑view visual observations, historical ego‑motion, and long‑tail semantic instructions to generate intent‑ and risk‑aware trajectories, achieving consistent performance gains on a large‑scale real‑world long‑tail driving benchmark.

By Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran