arXiv:2608.21262v1 Announce Type: cross
Abstract: Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cut...
By Adam Noonan
arXiv:2608. 01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability.
By Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov
arXiv:2607. 18162v1 Announce Type: new Abstract: The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al.
By Shreyas Pradeepkumar Khandale
arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.
By Mame Diarra Toure, David A. Stephens
Signal‑Routed Temperature Scaling (SRTS‑BCE) is a 10‑parameter, argmax‑preserving calibration method that separates calibration objectives from adaptive capacity. It cross‑fits a correctness‑risk score over six logit statistics and assigns a top‑label BCE temperature to each of three risk groups, generalizing TvA‑TS when K=1. Experiments on fine‑tuned CIFAR‑100 and ViT‑B/16 show that SRTS‑BCE reduces ECE from 1.65 to 0.96 with a small calibration budget, outperforming higher‑capacity SMART+BCE when only 250 examples are available, and revealing a budget‑dependent ranking reversal on Swin‑T.
By Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv:2608. 05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human.
By Jianru Shen
arXiv:2609.00691v1 Announce Type: new
Abstract: Post-hoc out-of-distribution detectors are fitted on a finite reference set, so every score they produce is an estimate. If we had chosen a different s...
By Donghoon Lee, Shinjin Kang
arXiv:2609.36721v1 Announce Type: new
Abstract: Fixed-budget adaptation from heterogeneous data sources requires deciding not only how much data to use, but how much exposure each source receives. Si...
By Yupeng Chang, Yuan Wu
arXiv:2606. 15910v2 Announce Type: replace Abstract: A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors.
By Reza Khanmohammadi, Kundan Thind, Mohammad M. Ghassemi
arXiv:2607. 13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning.
By Wisdom Dogah
arXiv:2608. 08424v1 Announce Type: cross Abstract: Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage.
By Chenchen Peng, Mixia Wu, Qijing Yan, Zhiqi Shen, Jie Zhang
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang