arXiv:2605. 21446v2 Announce Type: replace-cross Abstract: Interpretable autonomous driving planners depend not only on generating explanations, but also on those explanations remaining reliable under real-world sensor degradation.
By Abhinaw Priyadershi, Jelena Frtunikj
Intent misinterpretation during vehicle interactions causes recurring planning failures. We study a decision layer in which a language-guided intent module reads structured descriptors, computes a smo...
arXiv:2609.11592v2 Announce Type: replace
Abstract: Transportation agencies increasingly predict crash-injury severity with statistical and machine-learning models, but these models do not state how...
By Amir Rafe, Subasish Das
arXiv:2606. 11266v1 Announce Type: new Abstract: The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision has begun, and the Lagrange multiplier of PPO-Lagrangian grows only after the episode budget has been exceeded.
By Samuel Tetteh, Cody Fleming
arXiv:2606. 29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review.
By Mengdie Flora Wang, Haochen Xie, Guanghui Wang, Devin Zhang, Jae Oh Woo
arXiv:2609.22582v1 Announce Type: new
Abstract: End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a rank reports an outcome, not the behaviour behind...
By Ruolin Yang, Zilin Huang, Buoyue Wang, Zhengyang Wan, Yuhao Luo, Zihao Sheng, Sikai Chen
CS-WCP introduces confidence‑set weighted conformal prediction to provide robust prediction sets for large‑language‑model judges when deployment traffic shifts the prevalence of task or policy groups. By constructing simultaneous exact intervals for source and target group masses and taking the union over all compatible ratio vectors, CS‑WCP achieves high coverage (mean 0.973) with few failures across 336 constructed traffic shifts, outperforming standard source conformal prediction. The method offers an auditable coverage safeguard under uncertain mixture weights, focusing on conservative tail protection rather than tighter set sizes.
By Ibne Farabi Shihab, Fariya Afrin
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.
By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
arXiv:2608. 19376v1 Announce Type: cross Abstract: Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot vision-language models (VLMs).
By Jai Kumar Sharma, Amartya Dutta
PRISM (Proactive Risk Intelligence and Safety Management) is an agentic multi-model architecture designed to shift autonomous transportation safety from reactive crash avoidance to proactive, continuous risk management. It uses inverse crash‑probability modeling to transform binary crash classifiers into dynamic safety scores, and runs three specialized models—trajectory kinematics, environmental risk, and VRU interaction—coordinated by a reinforcement‑learning reasoning layer. Across 1,296 naturalistic driving scenarios, PRISM achieved a mean safety score of 68/100, classified 77.6% of situations as advisory, and flagged 3.8% as near‑misses, with 11% requiring intervention or emergency response, highlighting trajectory risk and VRU proximity as key safety factors.
By Joyjit Roy, Samaresh Kumar Singh, Sushanta Das
arXiv:2607. 20549v1 Announce Type: new Abstract: Trajectory datasets used in ADAS evaluation are heavily biased toward routine driving; genuine vehicle-to-vehicle conflict events are rare, and the rarer the event, the higher the cost when an ADAS system fails to handle it.
By Eni Solomon Laughter