arXiv AI

Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs

arXiv:2605. 21446v2 Announce Type: replace-cross Abstract: Interpretable autonomous driving planners depend not only on generating explanations, but also on those explanations remaining reliable under real-world sensor degradation.

arXiv AI
2d ago

Towards Reliable Vision-Language Models for Autonomous Driving

The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.

By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
arXiv Computer Vision
4d ago

CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving

CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.

By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
arXiv AI
Sep 4

LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving

LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.

By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
arXiv Computer Vision
Aug 26

Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation

The paper introduces Variance‑Guided Spatial Attention Fusion (VG‑SAF), a method for robust end‑to‑end driving that fuses camera and LiDAR data while handling asymmetric sensor degradation. VG‑SAF uses a physically grounded augmentor to generate dense reliability masks, modality‑specific experts to predict per‑pixel reliability scales, and a hybrid attention mechanism that gates unreliable cells and balances modalities. The approach also includes a Laplace uncertainty head to signal severe or combined sensor failures, and demonstrates improved closed‑loop robustness on the CARLA Longest6 benchmark across various degradation scenarios.

By Weizhi Tao, Zengwang Jin, Xiao Wang, Hailong Huang
arXiv Machine Learning
Sep 22

Correcting Learning-based Perception for Safety

The paper presents a two-step method to correct machine‑learning based perception for safety in autonomous systems. First, it uses offline computation to characterize uncertainties from the ML module via preimages of perception contracts. Then, at runtime, a risk heuristic selects specific states from these uncertain estimates to guide control decisions, reducing safety violations in adaptive cruise control scenarios while adding minimal delay.

By Yan Miao, Hussein Darir, Sayan Mitra
arXiv AI
Jul 28

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

arXiv:2607. 23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality.

By Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng