arXiv:2403. 06013v2 Announce Type: replace Abstract: This paper delves into the critical area of deep learning robustness, challenging the conventional belief that classification robustness and explanation robustness in image classification systems are inherently correlated.
By Tiejin Chen, Wenwang Huang, Linsey Pang, Dongsheng Luo, Hua Wei
arXiv:2606. 14965v1 Announce Type: new Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit.
By Shadman Islam, Agustinus Kristiadi, Mostafa Milani
The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.
By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez
arXiv:2606. 01746v1 Announce Type: cross Abstract: Modern neural networks are highly susceptible to adversarial perturbations.
By Kai Wang
arXiv:2603. 23318v2 Announce Type: replace Abstract: Among the different possible strategies for evaluating the reliability of individual predictions of classifiers, robustness quantification stands out as a method that evaluates how much uncertainty a classifier could cope with before changing its prediction.
By Rodrigo F. L. Lassance, Jasper De Bock
arXiv:2509. 07605v2 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge to supervised classification, particularly in critical domains like medical diagnostics and anomaly detection where minority class instances are rare.
By Ali Nawaz, Amir Ahmad, Shehroz S. Khan
The paper introduces a diagnostic test for class sensitivity in saliency methods, revealing that many popular techniques produce nearly identical explanations regardless of the predicted class. This limitation appears across different architectures and datasets, indicating a structural issue. To address this, the authors propose CASE, a contrastive explanation method that isolates features uniquely discriminative for the predicted class, and demonstrate its improved fidelity and class specificity through experiments.
By Dane Williamson, Yangfeng Ji, Matthew Dwyer
arXiv:2609.13914v1 Announce Type: new
Abstract: Machine-learning models are commonly developed under an assumption that training and test data are sufficiently complete, balanced, labelled, and drawn...
By Masoumeh Zareapoor
Counterfactual (CF) explanations for time-series classifiers are usually evaluated one example at a time: what minimal edit flips this single window's prediction? We argue that the more informative qu...
arXiv:2608. 13190v1 Announce Type: new Abstract: Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable.
By Qianqian Wang, Yunshan Li, Dawei Huang, Wenwu Gong, Lili Yang
The paper introduces techniques for measuring the robustness of predictions made by two generative classifiers—naive Bayes classifiers and generative forests—whose underlying models are probabilistic graphical models. Robustness is defined as the degree to which the classifier’s distribution can be perturbed without altering its prediction, with perturbations explored via epsilon‑contamination, total variation distance, and chi‑squared divergence neighborhoods. Experiments on benchmark datasets show that the computed robustness values can serve as indicators of prediction trustworthiness and are compared against other existing indicators.
By Adri\'an Detavernier, Jasper De Bock
arXiv:2608.23164v1 Announce Type: cross
Abstract: Counterfactual (CF) explanations for time-series classifiers are usually evaluated one example at a time: what minimal edit flips this single window'...
By Syed Muhammad Hamza Zaidi, Szymon Bobek, Grzegorz J. Nalepa, Myra Spiliopoulou