arXiv:2503. 08936v3 Announce Type: replace-cross Abstract: Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (ADAS).
By Lev Sorokin, Matteo Biagiola, Andrea Stocco
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
By Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
arXiv:2608. 19425v1 Announce Type: cross Abstract: Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions.
By Dijie Zhu, Seunghun Oh, Ruopeng Huang, Zhiyu Huang, Jiaqi Ma, Chen Tang
arXiv:2607. 14826v1 Announce Type: cross Abstract: Safe physical AI for robot actions are required not only likely to succeed but tested to be safe before execution.
By Naren Vasantakumaar, Tom Schierenbeck, Michael Beetz
SCAPE is a scenario‑conditioned simulation‑augmented policy evaluation framework that predicts real‑world policy performance for specific scenarios using limited paired simulation‑and‑real samples and extensive simulation rollouts. It corrects sim‑to‑real bias in simulation labels before training the prediction model and calibrates prediction uncertainty via conformal prediction. Experiments on autonomous driving and quadruped velocity tracking show SCAPE reduces scenario‑level prediction error, improves testing sample efficiency, narrows calibrated prediction intervals, and generalizes better to out‑of‑distribution scenarios, enabling fine‑grained deployment strategies.
The paper audits runtime failure monitors that use a model’s internal representations to predict failures in autonomous driving tasks. Across two tasks—online vectorized map generation with LaneSegNet and end‑to‑end planning with VAD—the authors find that frame‑level errors can be predicted with high AUROC scores using supervised latent probes. However, adding latent features to baseline monitors that use only observable inputs and outputs does not yield statistically significant improvements, suggesting that internal representations may not provide additional predictive value beyond what is already observable.
By Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar
arXiv:2607. 14439v1 Announce Type: new Abstract: Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks.
By Andrew Liao, Hanchen Cui, Karthik Desingh, Aryan Deshwal
arXiv:2606. 31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely.
By Huaze Tang, Bill Zeng, Chao Wang, Zhenpeng Shi, Qian Zhang, Wenbo Ding
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
By Tanmay Khandait, Preetom Biswas, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Giulia Pedrielli
The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.
By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
arXiv:2608.22421v1 Announce Type: new
Abstract: World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic pr...
By Zhanpeng Shi, Zi Liang, Rong Feng, Shiqin Tang, Xuyang Chen, Hongzong Li
arXiv:2607. 22697v1 Announce Type: new Abstract: Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution.
By Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez