arXiv:2606. 04310v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) are increasingly being deployed in security-critical and safety-sensitive applications, which makes rigorous testing essential to identify and mitigate model weaknesses.
By Bin Duan, Matthew B. Dwyer, Guowei Yang
arXiv:2607. 20046v1 Announce Type: cross Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important.
By Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
By Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
arXiv:2601. 15041v2 Announce Type: replace Abstract: The increasing deployment of deep learning systems requires systematic evaluation of their reliability in real-world scenarios.
By Oliver Wei{\ss}l, Vincenzo Riccio, Severin Kacianka, Andrea Stocco
TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By using partial model tomography, TestifAI reconstructs higher-order perturbation effects from low-order tests, achieving less than 7% error while reducing inferences by 60–80%.
arXiv:2608. 18900v2 Announce Type: replace Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.
By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% error while cutting inference counts by 60‑80% on image and language classification tasks.
arXiv:2608. 12144v1 Announce Type: cross Abstract: Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.
By Yidi Kao, Shawn Burnham, Tommi Rose Fahy, Ali Ghanbari
TestifAI is a deep learning testing framework that estimates model robustness against combinations of semantic input perturbations such as blur, brightness, and zoom. It allows users to define operational conditions as structured spaces with discrete severity levels and query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% estimation error while reducing inference counts by 60‑80% across five image and language classification tasks.
By Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
arXiv:2607. 05461v1 Announce Type: cross Abstract: Existing methods for testing deep neural networks (DNNs) primarily prioritize test inputs likely to reveal model faults under a fixed labeling budget.
By Bonan Shen, Wei-Jung Huang, Xin Liu, Jiazhou Gao, Tao Ning
arXiv:2607. 11228v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases.
By Anqi Li, Jie Zhang, Zhongqi Wang, Songkai Xue, Jiahao Wang, Shiguang Shan, Xilin Chen
While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities.