The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2608.27704v1 Announce Type: new
Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version,...
By Madhusudan Srinivasan, Namith Nishal Raphae
The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.
By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv:2608. 09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken.
By Ange-Cl\'ement Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal
Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions. Consequently, users lack a reliable signal for deciding when a prediction can be trusted.
arXiv:2607. 20046v1 Announce Type: cross Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important.
By Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin
arXiv:2608. 20326v1 Announce Type: cross Abstract: Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions.
By Parampreet Singh, Anushka Singh, Sumit Kumar, Vipul Arora
arXiv:2606. 29484v1 Announce Type: cross Abstract: Modern deepfake detectors are rarely consumed as bare classifiers.
By Md Anas Biswas
The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.
By Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao
arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
The paper introduces Alignment Forecasting, a method for predicting whether fine‑tuning a language model on a given dataset will increase specific alignment failures such as deception or sycophancy. It presents ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions across many models, datasets, and failure modes, and shows that a simple forecasting scaffold using an LLM’s assessment of dataset bias can outperform baseline forecasters. The authors demonstrate that filtering out high‑risk training examples identified by the forecaster can improve alignment in multiple‑choice evaluations, though benefits in open‑ended conversations remain uncertain.
By Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
The paper introduces the Temperature Scaling Attack (TSA), a training‑time method that degrades model confidence calibration while keeping predictive accuracy largely intact. TSA injects temperature scaling with a learning‑rate coupling during local federated training, shifting confidence scores and causing significant calibration errors (e.g., a 145% increase on CIFAR‑100) with less than a 2% drop in accuracy. The authors provide a convergence analysis for non‑IID settings and demonstrate TSA’s effectiveness across three benchmarks, robust aggregation, and post‑hoc calibration defenses, highlighting its impact on mission‑critical systems such as healthcare verification and autonomous driving.
By Kichang Lee, Jaeho Jin, JaeYeon Park, Songkuk Kim, JeongGil Ko