arXiv AI

Physics-Grounded Causal Auditing of End-to-End Driving Planners

The paper introduces CADET, a training‑free framework for auditing, benchmarking, and repairing spurious reliance in pretrained end‑to‑end autonomous‑driving planners. It addresses the problem that such planners often learn statistical shortcuts—associating co‑occurring scene elements with driving decisions—rather than causal variables, which undermines reliability in rare scenarios. CADET can detect and correct these causal confusions without retraining the model or updating its parameters.

arXiv AI
Jun 15

CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners

arXiv:2606. 14438v1 Announce Type: cross Abstract: End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merely co-occur with expert actions (a roadside object, a building facade) with driving decisions, rather than the variables that causally determine them.

By Zikun Guo
arXiv AI
Jul 1

What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning

arXiv:2606. 31106v1 Announce Type: cross Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics.

By Hyeonchang Jeon, Kyungbeom Kim, Eugene Vinitsky, Kyung-Joong Kim
arXiv AI
Sep 25

Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners

The paper argues that current embodied vision‑language planning benchmarks favor linguistic next‑token prediction over physically grounded next‑state reasoning, leading models to rely on language priors rather than true causal dependencies. To address this, the authors introduce Causal‑Plan‑Bench, a diagnostic suite covering four causal dimensions, and Causal‑Plan‑1M, a million‑scale corpus of explicit causal reasoning traces extracted from egocentric videos. Extensive experiments show that existing models perform poorly on these tasks, while a new model trained with a tailored recipe—Causal Planner based on Qwen3‑VL‑8B—achieves significant gains, demonstrating the feasibility of physically grounded causal reasoning.

By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Zirui Song, Lijie Wang, Chong Luo, Bei Liu, Yiming Li
arXiv AI
6d ago

Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization

The paper introduces Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that corrects intermediate waypoints of end-to-end driving policies while preserving the predicted endpoint. ECO does not require maps, privileged simulator state, or additional training, and can be applied to a wide range of waypoint-emitting policies. Experiments on two closed-loop simulators show that ECO significantly improves closed-loop performance, achieving top results in the HUGSIM Closed-Loop Driving Challenge and boosting scene scores on AlpaSim.

By Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta
arXiv AI
6d ago

Auditing Latent-Space Monitors for Autonomous Driving

The paper audits runtime failure monitors that use a model’s internal representations to predict failures in autonomous driving tasks. Across two tasks—online vectorized map generation with LaneSegNet and end‑to‑end planning with VAD—the authors find that frame‑level errors can be predicted with high AUROC scores using supervised latent probes. However, adding latent features to baseline monitors that use only observable inputs and outputs does not yield statistically significant improvements, suggesting that internal representations may not provide additional predictive value beyond what is already observable.

By Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar
arXiv Machine Learning
Sep 18

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

OPTED is a method for on‑policy fine‑tuning of end‑to‑end driving models that separates reinforcement learning from the policy update. A privileged teacher trained with RL on vectorized inputs (HD‑maps and bounding boxes) supervises the pre‑trained student during closed‑loop post‑training. Applied to the camera‑based models TransFuser and VaVAM in AlpaSim, OPTED boosts driving scores by 1.6× and 9.5×, respectively, while requiring roughly three orders of magnitude fewer simulator interactions than direct RL post‑training.

By Damiano Da Col, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, Christos Sakaridis
arXiv AI
Jul 2

Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces

arXiv:2603. 14354v3 Announce Type: replace-cross Abstract: End-to-End autonomous driving (E2E-AD) systems face challenges in lifelong learning, including catastrophic forgetting, difficulty in knowledge transfer across diverse scenarios, and spurious correlations between unobservable confounders and true driving intents.

By Jiayuan Du, Yuebing Song, Yiming Zhao, Xianghui Pan, Jiawei Lian, Yuchu Lu, Liuyi Wang, Chengju Liu, Qijun Chen
arXiv AI
Sep 24

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

arXiv:2609. 28366v1 Announce Type: cross Abstract: Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning.

By Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li