arXiv AI

Simulator Ensembles for Trustworthy Autonomous Driving Systems Testing

arXiv:2503. 08936v3 Announce Type: replace-cross Abstract: Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (ADAS).

arXiv AI
Sep 24

Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation

Teach-to-Crash is a closed‑loop testing framework that uses a dual‑LLM architecture to generate collision‑inducing scenarios for autonomous driving systems. A high‑reasoning Teacher LLM controls the search when collision metrics stagnate, while a low‑reasoning Student LLM produces simulator‑executable scenarios in JSON. In a CARLA case study, Teach‑to‑Crash achieved the highest collision hit rate (90.79 %), the shortest mean time‑to‑collision (18.31 s), and superior diversity and avoidability metrics compared to other methods.

By Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim
arXiv AI
Sep 10

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits. whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."

By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz
arXiv AI
Sep 18

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

The paper introduces SPAR, a closed‑loop simulation platform that couples real‑time AUV control software with a higher‑level orchestration layer for fault injection, prompting, and evaluation of large language models (LLMs) in diagnosing and recovering from anomalies. SPAR enables ensemble testing of LLMs, comparing a frontier model with three locally deployable LLMs on a mass‑shift fault scenario across 480 trials, revealing that model choice significantly affects diagnostic accuracy. The study demonstrates that while the frontier model consistently ranks the correct fault mechanism among its top hypotheses, local models succeed mainly when they follow the full diagnostic procedure, and overall diagnosis and operational decisions appear decoupled in this dataset.

By Khalid Halba, Kylie Cooper, James G. Bellingham
arXiv AI
Sep 7

ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems

ARIA is a multi‑agent large‑language‑model framework that autonomously runs end‑to‑end visual tests on Android infotainment systems. From simple scenario sentences, it executes interactions, generates reproducible scripts, and produces detailed reports with visual evidence. In evaluation on a manufacturer’s device, ARIA achieved a 93.3% completion rate, correctly identified all known defects, and demonstrated lower false‑positive rates compared to a single‑agent baseline.

By Ant\'onio Azevedo, Bruno Lima, Jo\~ao Pascoal Faria
Hugging Face Trending Papers
Aug 19

SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation

SCAPE is a scenario‑conditioned simulation‑augmented policy evaluation framework that predicts real‑world policy performance for specific scenarios using limited paired simulation‑and‑real samples and extensive simulation rollouts. It corrects sim‑to‑real bias in simulation labels before training the prediction model and calibrates prediction uncertainty via conformal prediction. Experiments on autonomous driving and quadruped velocity tracking show SCAPE reduces scenario‑level prediction error, improves testing sample efficiency, narrows calibrated prediction intervals, and generalizes better to out‑of‑distribution scenarios, enabling fine‑grained deployment strategies.

Hugging Face Trending Papers
Sep 17

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

The paper introduces SPAR, a closed‑loop simulation platform that couples real‑time AUV control software with a higher‑level orchestration layer to evaluate large language models (LLMs) for fault diagnosis and recovery. It demonstrates that a frontier LLM outperforms locally deployable models in identifying a mass‑shift fault, and shows that successful diagnosis depends on following a complete diagnostic procedure rather than premature conclusions. The study provides an architecture and ensemble evaluation methodology for LLM‑assisted mission management on low‑power autonomous underwater vehicles.