arXiv:2602. 07453v2 Announce Type: replace Abstract: Decision tree ensembles are widely used in critical domains, making robustness and sensitivity analysis essential to their trustworthiness.
By Namrita Varshney, Ashutosh Gupta, Arhaan Ahmad, Tanay V. Tayal, S. Akshay
arXiv:2606. 01746v1 Announce Type: cross Abstract: Modern neural networks are highly susceptible to adversarial perturbations.
By Kai Wang
arXiv:2512. 13003v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is essential for determining when a supervised model encounters inputs that differ meaningfully from its training distribution.
By Min Lu, Hemant Ishwaran
arXiv:2606. 20208v1 Announce Type: new Abstract: Machine learning models are predominantly evaluated through predictive performance metrics such as ranking quality, prediction error, or classification accuracy.
By Guillaume Olivier Delplanque (LIG), Pierre Genev\`es (LIG), Nabil Laya\"ida (LIG,TYREX), Zephirin Faure
arXiv:2603. 23318v2 Announce Type: replace Abstract: Among the different possible strategies for evaluating the reliability of individual predictions of classifiers, robustness quantification stands out as a method that evaluates how much uncertainty a classifier could cope with before changing its prediction.
By Rodrigo F. L. Lassance, Jasper De Bock
arXiv:2605. 27618v2 Announce Type: replace Abstract: Despite the wide use of explainability techniques to attempt to understand the behavior of Artificial Intelligence (AI), the generated explanations may not always be reliable.
By Tom\'as Pereira, Jo\~ao Vitorino, Eva Maia, Isabel Pra\c{c}a
arXiv:2607. 06637v1 Announce Type: new Abstract: In this work, we propose a unified approach for diagnosing misclassification and assessing the robustness of black-box classifiers.
By Evgenii Kuriabov, David Miller, Jia Li
arXiv:2607. 19386v1 Announce Type: new Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores.
By Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady
arXiv:2602. 14161v2 Announce Type: replace Abstract: Detecting prompt injection, jailbreak attacks, and harmful requests is critical for deploying LLM-based agents safely, yet current evaluation practices in this literature overestimate generalization.
By Max Fomin
arXiv:2606. 25004v1 Announce Type: new Abstract: In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality.
By Gefei Tan, Adria Gascon, Sarah Meiklejohn, Mariana Raykova
arXiv:2604. 25077v2 Announce Type: replace Abstract: Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots.
By Hamid Osooli, Kareema Batool, Rick Gentry, Tiasa Singha Roy, Ashwin Gupta, Anirudha Ramesh
arXiv:2607. 03075v1 Announce Type: new Abstract: Safety-critical applications require classifiers that are both robust and reliable.
By Nicolas Sournac, Ahmed Baha Ben Jmaa, Bertrand Braeckeveldt