arXiv:2608.23164v1 Announce Type: cross
Abstract: Counterfactual (CF) explanations for time-series classifiers are usually evaluated one example at a time: what minimal edit flips this single window'...
By Syed Muhammad Hamza Zaidi, Szymon Bobek, Grzegorz J. Nalepa, Myra Spiliopoulou
arXiv:2610.00895v1 Announce Type: cross
Abstract: Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such st...
By Sidney Bender, Benedikt Kunz, Ahmed Zeid, Shinichi Nakajima, Klaus-Robert M\"uller, Marco Morik
The paper introduces a unified evaluation protocol for robust counterfactual explanations (CFE), testing six robust methods and two baselines across four tabular datasets under eight types of model change. It shows that robustness scores vary by change type and that methods designed for one change family may not transfer to others, with RobX performing most consistently. The study emphasizes the need for a common protocol that defines model changes, measures their impact, and separates generation performance from robustness.
By Marcin Kostrzewa, Maciej Zi\k{e}ba
arXiv:2607. 06637v1 Announce Type: new Abstract: In this work, we propose a unified approach for diagnosing misclassification and assessing the robustness of black-box classifiers.
By Evgenii Kuriabov, David Miller, Jia Li
arXiv:2606. 04209v1 Announce Type: new Abstract: Counterfactual explanations seek small, semantically meaningful changes to an input that alter a model's prediction, and are widely used to interpret and audit machine learning systems.
By Ioanna Gemou, Matteo Gamba, Randall Balestriero, Ritambhara Singh
arXiv:2603. 22016v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verification, repeated attempts, or unnecessary exploration that wastes computation and can even overturn the correct answer.
By Xinyan Wang, Xiaogeng Liu, Ming Pei, Chaowei Xiao
arXiv:2603. 16436v2 Announce Type: replace Abstract: Counterfactual explanations (CE) explain model decisions by identifying input modifications that lead to different predictions.
By Yikai Gu, Lele Cao, Bo Zhao, Lei Lei, Lei You
The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.
By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez
This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase efficiency and safety.
The paper presents a comparative analysis of six state‑of‑the‑art counterfactual explainers for graph neural networks, focusing on methods that can both add and remove edges to alter model predictions. It evaluates these explainers across diverse real‑world and synthetic datasets, covering binary and multi‑class graph and node classification tasks, using a range of quantitative and qualitative metrics. The study highlights the trade‑offs between explanation size, coverage, and quality, aiming to pinpoint each method’s strengths and weaknesses to inform future research.
By Maria Myrto Villia, Filippos Gouidis, Theodore Patkos, Panos Trahanias
arXiv:2607. 22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?
By Yassine Oueslati, Daniil Kirilenko, Martin Gjoreski, Marc Langheinrich
arXiv:2608. 04777v1 Announce Type: cross Abstract: Oscillatory signals, such as vibration, carry class-discriminative information in specific frequency bands; perturbing them in raw feature space for counterfactual analysis easily destroys their temporal structure and produces physically implausible results.
By Udo Schlegel, Julian Rakuschek, Thomas Seidl, Andreas Holzinger, Tobias Schreck, Javier Del Ser