arXiv:2603. 23318v2 Announce Type: replace Abstract: Among the different possible strategies for evaluating the reliability of individual predictions of classifiers, robustness quantification stands out as a method that evaluates how much uncertainty a classifier could cope with before changing its prediction.
By Rodrigo F. L. Lassance, Jasper De Bock
arXiv:2607. 03075v1 Announce Type: new Abstract: Safety-critical applications require classifiers that are both robust and reliable.
By Nicolas Sournac, Ahmed Baha Ben Jmaa, Bertrand Braeckeveldt
arXiv:2604. 05324v2 Announce Type: replace Abstract: Statistical evaluation aims to estimate the generalization performance of a model using held-out i.
By Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao
arXiv:2501. 18897v4 Announce Type: replace-cross Abstract: Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty quantification.
By Zijun Gao, Yan Sun, Han Su
arXiv:2510. 05709v2 Announce Type: replace-cross Abstract: LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations are available for classical inference, and (ii) test prompts are independent.
By Mary Llewellyn, Isobel Thornton, James Bishop, Annie Gray
arXiv:2606. 25004v1 Announce Type: new Abstract: In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality.
By Gefei Tan, Adria Gascon, Sarah Meiklejohn, Mariana Raykova
The paper demonstrates that common binary classification metrics—Matthews' correlation coefficient, Cohen's κ, the F-score, and the Jaccard similarity—are not robust to extreme class imbalance, as the Bayes classifier’s true positive rate tends to zero when the minority class proportion vanishes. To address this, the authors propose robustified versions of these metrics that include a tuning parameter, ensuring that the Bayes-optimal classifier’s threshold remains bounded and its true positive rate stays above zero even in highly imbalanced scenarios. The study provides theoretical bounds, simulation results, and practical guidance on applying these robust metrics to real data, such as a credit‑default dataset, and discusses their relationship to ROC and precision‑recall curves.
By Hajo Holzmann, Bernhard Klar
arXiv:2603. 27270v2 Announce Type: replace Abstract: Credal sets, i.
By Xabier Gonzalez-Garcia, Siu Lun Chau, Julian Rodemann, Michele Caprio, Krikamol Muandet, Humberto Bustince, S\'ebastien Destercke, Eyke H\"ullermeier, Yusuf Sale
arXiv:2607. 06637v1 Announce Type: new Abstract: In this work, we propose a unified approach for diagnosing misclassification and assessing the robustness of black-box classifiers.
By Evgenii Kuriabov, David Miller, Jia Li
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv:2608. 08489v1 Announce Type: new Abstract: Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination.
By Subhabrata Majumdar, Anand Deo, Partha Pratim Saha, Abhik Ghosh
arXiv:2605. 13830v2 Announce Type: replace-cross Abstract: Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence verifying properties on these models has been an active topic of study over the last decade.
By Ajinkya Naik, Chaitanya Garg, S. Akshay, Ashutosh Gupta, Kuldeep S. Meel