arXiv AI By Lev Sorokin, Matteo Biagiola, Andrea Stocco

Simulator Ensembles for Trustworthy Autonomous Driving Systems Testing

Read the original on arXiv AI →

arXiv:2503. 08936v3 Announce Type: replace-cross Abstract: Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (ADAS).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation

Teach-to-Crash is a closed‑loop testing framework that uses a dual‑LLM architecture to generate collision‑inducing scenarios for autonomous driving systems. A high‑reasoning Teacher LLM controls the search when collision metrics stagnate, while a low‑reasoning Student LLM produces simulator‑executable scenarios in JSON. In a CARLA case study, Teach‑to‑Crash achieved the highest collision hit rate (90.79 %), the shortest mean time‑to‑collision (18.31 s), and superior diversity and avoidability metrics compared to other methods.

By Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim
arXiv AI
Sep 10

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits. whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."

By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz