arXiv AI

\textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems

TestifAI is a deep learning testing framework that estimates model robustness against combinations of semantic input perturbations such as blur, brightness, and zoom. It allows users to define operational conditions as structured spaces with discrete severity levels and query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% estimation error while reducing inference counts by 60‑80% across five image and language classification tasks.

Hugging Face Trending Papers
Aug 19

\textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems

TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% error while cutting inference counts by 60‑80% on image and language classification tasks.

Hugging Face Trending Papers
Aug 19

TestifAI: Tomography-Based Testing for Deep Learning Systems

TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By using partial model tomography, TestifAI reconstructs higher-order perturbation effects from low-order tests, achieving less than 7% error while reducing inferences by 60–80%.

arXiv Computer Vision
Aug 31

Image Augmentation as Test Generation for Deep Learning-Based Image Retrieval Systems

The paper reviews 50 image augmentation and generation techniques, categorizing them into ten groups, and conducts a large‑scale empirical study to assess their effectiveness as test generators for embedding‑based image retrieval systems. Using Amazon Titan and OpenCLIP embeddings, the authors evaluate the techniques across four dimensions—embedding‑space similarity, embedding uncertainty, semantic realism, and retrieval failure rate—on CIFAR‑10, ImageNet‑1K, and an industrial dataset. Results show that weather simulation and SaSPA yield the highest uncertainty and failure rates while maintaining realistic visuals, whereas GAN‑based methods produce low realism due to synthetic artifacts.

By Yehan De Silva, Anirudh Sridhar, Armin Lotfy, Nafiseh Kahani, Yvan Labiche, Ziyu Wang, Frank Ouyang, Clare Carty, Azalia Shamsaei