arXiv AI

Upper Bounds on the Generalization Error of Deep Learning Models via Local Robustness and Stability

arXiv:2606. 16883v1 Announce Type: cross Abstract: Generalization is a critical property of data-driven models, particularly deep learning models deployed in safety-critical applications.

Hugging Face Trending Papers
Aug 19

TestifAI: Tomography-Based Testing for Deep Learning Systems

TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By using partial model tomography, TestifAI reconstructs higher-order perturbation effects from low-order tests, achieving less than 7% error while reducing inferences by 60–80%.

arXiv AI
Aug 20

\textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems

TestifAI is a deep learning testing framework that estimates model robustness against combinations of semantic input perturbations such as blur, brightness, and zoom. It allows users to define operational conditions as structured spaces with discrete severity levels and query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% estimation error while reducing inference counts by 60‑80% across five image and language classification tasks.

By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
Hugging Face Trending Papers
Aug 19

\textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems

TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% error while cutting inference counts by 60‑80% on image and language classification tasks.