The paper introduces prediction‑powered smoothing (PP‑S) and its taxonomy‑aware extension (PP‑TS) to improve point and interval estimates of domain‑specific AI performance when only a limited sample of labeled units is available. It also proposes a new design‑based cross‑validation score that is approximately unbiased for selecting between direct and smoothed estimators. Experiments on a curated benchmark and real‑world agent traffic show that the proposed methods outperform direct estimators in both accuracy and coverage, and that the new score matches the performance of an independent validation sample while providing more precise error estimates.
By Sho Kawano, Zehang Richard Li, Paul A. Parker
arXiv:2608. 13793v1 Announce Type: cross Abstract: Machine learning (ML) has become an indispensable part of modern engineering design workflows.
By Tyler R. Johnson, Kian Ben-Jacob, Christopher P. Muller, Ramin Bostanabad
arXiv:2608. 13719v1 Announce Type: new Abstract: Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets.
By Anjali Parashar, Rachel Luo, Apoorva Sharma, Sushant Veer, Edward Schmerling, Carson Sobolewski, Mingxin Yu, Chuchu Fan, Marco Pavone
arXiv:2606. 04314v1 Announce Type: new Abstract: As neural networks are increasingly deployed in safety-critical domains, testing is essential to evaluate and improve their reliability.
By Bin Duan, Meiru Che, Guowei Yang
The paper introduces a risk‑controlled framework for using large language models (LLMs) as judges in tasks without reference answers. By calibrating uncertainty thresholds on a held‑out set, the method ensures that the false discovery rate of accepted verdicts stays below a user‑specified level α with high probability, using finite‑sample Clopper–Pearson intervals. When the parametric judge lacks confidence, the instance is routed to a retrieval‑augmented mode with a second calibrated threshold, preserving the error guarantee while achieving higher coverage than single‑mode baselines.
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
By Tanmay Khandait, Preetom Biswas, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Giulia Pedrielli