arXiv AI

Auditing Latent-Space Monitors for Autonomous Driving

The paper audits runtime failure monitors that use a model’s internal representations to predict failures in autonomous driving tasks. Across two tasks—online vectorized map generation with LaneSegNet and end‑to‑end planning with VAD—the authors find that frame‑level errors can be predicted with high AUROC scores using supervised latent probes. However, adding latent features to baseline monitors that use only observable inputs and outputs does not yield statistically significant improvements, suggesting that internal representations may not provide additional predictive value beyond what is already observable.

arXiv AI
6d ago

Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

The paper introduces Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that corrects intermediate waypoints of end-to-end driving policies while preserving the predicted endpoint. ECO does not require maps, privileged simulator state, or additional training, and can be applied to a wide range of waypoint-emitting policies. Experiments on two closed-loop simulators show that ECO significantly improves closed-loop performance, achieving top results in the HUGSIM Closed-Loop Driving Challenge and boosting scene scores on AlpaSim.

By Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta
arXiv AI
Aug 19

Physics-Grounded Causal Auditing of End-to-End Driving Planners

The paper introduces CADET, a training‑free framework for auditing, benchmarking, and repairing spurious reliance in pretrained end‑to‑end autonomous‑driving planners. It addresses the problem that such planners often learn statistical shortcuts—associating co‑occurring scene elements with driving decisions—rather than causal variables, which undermines reliability in rare scenarios. CADET can detect and correct these causal confusions without retraining the model or updating its parameters.

By Zikun Guo, Minglan Chen, Jinyou Zhai, Rongjin Zou
arXiv AI
Jun 15

CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners

arXiv:2606. 14438v1 Announce Type: cross Abstract: End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merely co-occur with expert actions (a roadside object, a building facade) with driving decisions, rather than the variables that causally determine them.

By Zikun Guo
arXiv AI
6d ago

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.

By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
arXiv Computer Vision
Sep 7

FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

The paper introduces FailureSpot, a label‑efficient method for detecting failures at the timestamp level in vision‑language‑action (VLA) policies. It first generates weak supervision from unlabeled VLA action chunks by identifying abnormal patterns, then employs active learning to annotate only the most uncertain trajectories. Experiments on multiple VLA policies demonstrate improved performance for both timestamp‑level and trajectory‑level failure detection.

By Jie Ma, Zongxi Liu, Yi Zhu