TestifAI: Tomography-Based Testing for Deep Learning Systems
arXiv:2608. 18900v2 Announce Type: replace Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.
TestifAI is a deep learning testing framework that estimates model robustness against combinations of semantic input perturbations such as blur, brightness, and zoom. It allows users to define operational conditions as structured spaces with discrete severity levels and query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% estimation error while reducing inference counts by 60‑80% across five image and language classification tasks.
arXiv:2608. 18900v2 Announce Type: replace Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.
TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% error while cutting inference counts by 60‑80% on image and language classification tasks.
TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By using partial model tomography, TestifAI reconstructs higher-order perturbation effects from low-order tests, achieving less than 7% error while reducing inferences by 60–80%.
arXiv:2606. 16883v1 Announce Type: cross Abstract: Generalization is a critical property of data-driven models, particularly deep learning models deployed in safety-critical applications.
Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e. g.
arXiv:2608. 19882v1 Announce Type: new Abstract: Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.
arXiv:2606. 04310v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) are increasingly being deployed in security-critical and safety-sensitive applications, which makes rigorous testing essential to identify and mitigate model weaknesses.
arXiv:2606. 04767v1 Announce Type: new Abstract: The robustness of deep neural networks is crucial for safety-critical deployments, yet existing evaluation methods are often attack-dependent and lack interpretability.
The paper reviews 50 image augmentation and generation techniques, categorizing them into ten groups, and conducts a large‑scale empirical study to assess their effectiveness as test generators for embedding‑based image retrieval systems. Using Amazon Titan and OpenCLIP embeddings, the authors evaluate the techniques across four dimensions—embedding‑space similarity, embedding uncertainty, semantic realism, and retrieval failure rate—on CIFAR‑10, ImageNet‑1K, and an industrial dataset. Results show that weather simulation and SaSPA yield the highest uncertainty and failure rates while maintaining realistic visuals, whereas GAN‑based methods produce low realism due to synthetic artifacts.
arXiv:2602.01718v2 Announce Type: replace Abstract: Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benc...
arXiv:2606. 02267v1 Announce Type: new Abstract: The vulnerability of deep neural networks to adversarial examples poses a significant challenge for real-world deployment.
arXiv:2606. 09746v1 Announce Type: cross Abstract: With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential.