TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% error while cutting inference counts by 60‑80% on image and language classification tasks.
arXiv:2608. 18900v2 Announce Type: replace Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.
By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
TestifAI is a deep learning testing framework that estimates model robustness against combinations of semantic input perturbations such as blur, brightness, and zoom. It allows users to define operational conditions as structured spaces with discrete severity levels and query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% estimation error while reducing inference counts by 60‑80% across five image and language classification tasks.
By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
arXiv:2606. 16883v1 Announce Type: cross Abstract: Generalization is a critical property of data-driven models, particularly deep learning models deployed in safety-critical applications.
By Abdul-Rauf Nuhu, Parham M. Kebria, Vahid Hemmati, Mahmoud N. Mahmoud, Edward Tunstel, Abdollah Homaifar
Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e. g.
arXiv:2608. 19882v1 Announce Type: new Abstract: Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.
By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
arXiv:2606. 04310v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) are increasingly being deployed in security-critical and safety-sensitive applications, which makes rigorous testing essential to identify and mitigate model weaknesses.
By Bin Duan, Matthew B. Dwyer, Guowei Yang
arXiv:2606. 04767v1 Announce Type: new Abstract: The robustness of deep neural networks is crucial for safety-critical deployments, yet existing evaluation methods are often attack-dependent and lack interpretability.
By Chong Zhang, Xiang Li, Jia Wang, Qiufeng Wang, Xiaobo Jin
arXiv:2602.01718v2 Announce Type: replace
Abstract: Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benc...
By Sora Nakai, Youssef Fadhloun, Kacem Mathlouthi, Kotaro Yoshida, Ganesh Talluri, Ioannis Mitliagkas, Hiroki Naganuma
arXiv:2606. 02267v1 Announce Type: new Abstract: The vulnerability of deep neural networks to adversarial examples poses a significant challenge for real-world deployment.
By Nicolas Stalder, Benjamin F. Grewe, Matteo Saponati, Pau Vilimelis Aceituno
arXiv:2606. 09746v1 Announce Type: cross Abstract: With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential.
By Sherwin Varghese, Matthew Wicker, Alessio Lomuscio
arXiv:2608. 10289v1 Announce Type: cross Abstract: Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-represented scenarios.
By Nusrat Jahan Mozumder, Divya Gopinath, Corina Pasareanu, Matthew Dwyer