TESTNAV: Pareto-Guided Search for Compositional Robustness Testing
Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e. g.
arXiv:2608. 19882v1 Announce Type: new Abstract: Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.
Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e. g.
TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% error while cutting inference counts by 60‑80% on image and language classification tasks.
TestifAI is a deep learning testing framework that estimates model robustness against combinations of input perturbations efficiently. It allows users to define structured spaces of semantic perturbations and severity levels, then query robustness for any combination. By using partial model tomography, TestifAI reconstructs higher-order perturbation effects from low-order tests, achieving less than 7% error while reducing inferences by 60–80%.
TestifAI is a deep learning testing framework that estimates model robustness against combinations of semantic input perturbations such as blur, brightness, and zoom. It allows users to define operational conditions as structured spaces with discrete severity levels and query robustness for any combination. By employing partial model tomography, TestifAI reconstructs higher‑order perturbation effects from low‑order tests, achieving less than 7% estimation error while reducing inference counts by 60‑80% across five image and language classification tasks.
arXiv:2608. 18900v2 Announce Type: replace Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.
arXiv:2606. 02134v1 Announce Type: cross Abstract: Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations.
arXiv:2607. 11228v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases.
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e. g.
arXiv:2606. 06943v1 Announce Type: cross Abstract: Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations.
While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities.
arXiv:2602.01718v2 Announce Type: replace Abstract: Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benc...
arXiv:2607. 24516v1 Announce Type: cross Abstract: While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed.