arXiv AI

When and How Severely: Scenario-Specific Safety Envelopes for Driving VLAs

arXiv:2606. 14238v1 Announce Type: cross Abstract: Safety certification of Vision-Language-Action (VLA) driving planners under ISO 21448 (SOTIF) rests on an Operational Design Domain (ODD) specification that answers two complementary questions: when does the planner start to fail, and how severely does it fail once it does?

arXiv Machine Learning
Jun 11

Seeing Before Colliding: Anticipatory Safe RL with Frozen Vision-Language Models

arXiv:2606. 11266v1 Announce Type: new Abstract: The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision has begun, and the Lagrange multiplier of PPO-Lagrangian grows only after the episode budget has been exceeded.

By Samuel Tetteh, Cody Fleming
arXiv Machine Learning
Sep 24

CS-WCP: Robust Conformal Sets for LLM-Judge Traffic Shifts with Uncertain Group Proportions

CS-WCP introduces confidence‑set weighted conformal prediction to provide robust prediction sets for large‑language‑model judges when deployment traffic shifts the prevalence of task or policy groups. By constructing simultaneous exact intervals for source and target group masses and taking the union over all compatible ratio vectors, CS‑WCP achieves high coverage (mean 0.973) with few failures across 336 constructed traffic shifts, outperforming standard source conformal prediction. The method offers an auditable coverage safeguard under uncertain mixture weights, focusing on conservative tail protection rather than tighter set sizes.

By Ibne Farabi Shihab, Fariya Afrin
arXiv AI
2d ago

Towards Reliable Vision-Language Models for Autonomous Driving

The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.

By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
arXiv Machine Learning
Sep 3

PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems

PRISM (Proactive Risk Intelligence and Safety Management) is an agentic multi-model architecture designed to shift autonomous transportation safety from reactive crash avoidance to proactive, continuous risk management. It uses inverse crash‑probability modeling to transform binary crash classifiers into dynamic safety scores, and runs three specialized models—trajectory kinematics, environmental risk, and VRU interaction—coordinated by a reinforcement‑learning reasoning layer. Across 1,296 naturalistic driving scenarios, PRISM achieved a mean safety score of 68/100, classified 77.6% of situations as advisory, and flagged 3.8% as near‑misses, with 11% requiring intervention or emergency response, highlighting trajectory risk and VRU proximity as key safety factors.

By Joyjit Roy, Samaresh Kumar Singh, Sushanta Das