arXiv:2503. 08936v3 Announce Type: replace-cross Abstract: Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (ADAS).
By Lev Sorokin, Matteo Biagiola, Andrea Stocco
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
By Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
arXiv:2607. 14826v1 Announce Type: cross Abstract: Safe physical AI for robot actions are required not only likely to succeed but tested to be safe before execution.
By Naren Vasantakumaar, Tom Schierenbeck, Michael Beetz
arXiv:2607. 14439v1 Announce Type: new Abstract: Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks.
By Andrew Liao, Hanchen Cui, Karthik Desingh, Aryan Deshwal
arXiv:2606. 31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely.
By Huaze Tang, Bill Zeng, Chao Wang, Zhenpeng Shi, Qian Zhang, Wenbo Ding
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
By Tanmay Khandait, Preetom Biswas, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Giulia Pedrielli
arXiv:2607. 22697v1 Announce Type: new Abstract: Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution.
By Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez
arXiv:2606. 31844v1 Announce Type: cross Abstract: A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments.
By Ziyan Wang, Tan Xiang, Peng Chen, Xintao Yan
arXiv:2606. 31131v1 Announce Type: new Abstract: To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial.
By Anjali Parashar, Chuchu Fan
arXiv:2606. 08508v1 Announce Type: cross Abstract: Generative robot policies fail unpredictably at deployment: they hesitate at critical moments, drift off-task, or commit to unrecoverable actions.
By Bingjia Huang, Xiangyu Li, Xiang Wang, Liang Mi, Zixu Hao, Weijun Wang, Hao Wu, Kun Li, Yunxin Liu, Ting Cao
arXiv:2606. 29898v1 Announce Type: cross Abstract: Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle.
By Haoxu Huang, Tongsam Zheng, Yifan Chen, Jiacheng You, Yang Gao
arXiv:2607. 01111v1 Announce Type: cross Abstract: Robot policies inevitably encounter failures when deployed in real environments.
By Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski, Jeff Schneider