arXiv AI

From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving

arXiv:2608. 08941v1 Announce Type: cross Abstract: Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain.

arXiv AI
Sep 3

From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems

The paper proposes a structured method for verifying Operational Design Domain (ODD) coverage in safety‑critical AI systems, particularly for aviation. It combines parameter discretization, constraint‑based filtering, and criticality‑based dimension reduction to create a multi‑step verification process. Using simulation data from AI‑based mid‑air collision avoidance research, the authors demonstrate how this approach can meet EASA’s requirement for complete ODD coverage in high‑dimensional spaces.

By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv AI
Sep 10

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits. whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."

By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz
arXiv Machine Learning
Sep 3

PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems

PRISM (Proactive Risk Intelligence and Safety Management) is an agentic multi-model architecture designed to shift autonomous transportation safety from reactive crash avoidance to proactive, continuous risk management. It uses inverse crash‑probability modeling to transform binary crash classifiers into dynamic safety scores, and runs three specialized models—trajectory kinematics, environmental risk, and VRU interaction—coordinated by a reinforcement‑learning reasoning layer. Across 1,296 naturalistic driving scenarios, PRISM achieved a mean safety score of 68/100, classified 77.6% of situations as advisory, and flagged 3.8% as near‑misses, with 11% requiring intervention or emergency response, highlighting trajectory risk and VRU proximity as key safety factors.

By Joyjit Roy, Samaresh Kumar Singh, Sushanta Das
arXiv AI
Sep 24

Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation

Teach-to-Crash is a closed‑loop testing framework that uses a dual‑LLM architecture to generate collision‑inducing scenarios for autonomous driving systems. A high‑reasoning Teacher LLM controls the search when collision metrics stagnate, while a low‑reasoning Student LLM produces simulator‑executable scenarios in JSON. In a CARLA case study, Teach‑to‑Crash achieved the highest collision hit rate (90.79 %), the shortest mean time‑to‑collision (18.31 s), and superior diversity and avoidability metrics compared to other methods.

By Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim
arXiv AI
Jul 31

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

arXiv:2607. 26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction.

By Xinyu Yang, Tianxing Chen, Honghao Su, Minxuan Wang, Chenze Yu, Zhangzheng Tu, Yue Chen, Yuxiao Huo, Lingfeng Zhang, Yan Huang, Yan Qin, Shaolong Zhu, Qiwei Liang, Hekun Tian, Shujia Liu, Guangyu Chen, Junhao Gong, Zixuan Li, Wenwei Lin, Zijian Lin, Wenxuan Zhu, Eric J Chen, Yue Yuan, Qize Yu, Jiaqi Liang, Haowen Yan, Hengfei Zhao, Weijie Wan, Zikun Xiao, Junyuan Tang, Baijun Chen, Kai-Chong Lei, Kaixuan Wang, Kailun Su, Zanxin Chen, Yao Mu, Renjing Xu, Chuqiao Lyu, Qi Xiong, Ping Luo, Wenbo Ding
arXiv AI
Jul 21

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

arXiv:2511. 14592v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns.

By Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao