arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.
By Mame Diarra Toure, David A. Stephens
arXiv:2605. 08446v3 Announce Type: replace Abstract: Bayesian neural networks are typically trained against the evidence lower bound (ELBO), whose Jensen gap closes only when the variational posterior is exact.
By Pavel Prochazka
The paper introduces CalSAM, a lightweight adaptation framework that fine‑tunes only the mask decoder of the Segment Anything Model (SAM) while keeping its encoders frozen. CalSAM employs a Feature Fisher Information Penalty (FIP) to reduce encoder sensitivity to domain shift and a Confidence Misalignment Penalty (CMP) to curb overconfident voxel‑wise errors. Experiments on cross‑center, scanner‑shift, and motion‑corrupted brain MRI datasets show significant gains in Dice similarity coefficient, Hausdorff distance, and expected calibration error, with only a modest training‑time overhead.
By Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari
arXiv:2609.01072v1 Announce Type: new
Abstract: Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only...
By Daehwan Kim, Haejun Chung, Ikbeom Jang
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2607. 18162v1 Announce Type: new Abstract: The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al.
By Shreyas Pradeepkumar Khandale
arXiv:2608. 19376v1 Announce Type: cross Abstract: Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot vision-language models (VLMs).
By Jai Kumar Sharma, Amartya Dutta
arXiv:2607. 13682v2 Announce Type: cross Abstract: Radiative Gaussian splatting reconstructs sparse-view CT fast and accurately, and recent work attaches per-Gaussian posteriors to yield per-voxel uncertainty maps.
By Chulin Zhao, Yiran Xu, Shu Liu
MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.
By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan
arXiv:2606. 28654v1 Announce Type: cross Abstract: Deep Neural Network (DNN) classifiers suffer from poor calibration when their softmax outputs (predictive confidence) deviate from the empirical likelihoods.
By Thiru Thillai Nadarasar Bahavan, Sachith Seneviratne, Saman Halgamuge
arXiv:2608. 09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken.
By Ange-Cl\'ement Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal
arXiv:2512.04305v3 Announce Type: replace
Abstract: Vision-language models (VLMs) such as CLIP are increasingly adapted across decentralized data silos, yet the reliability of their predictions under...
By Mainak Singha, Masih Aminbeidokhti, Paolo Casari, Gianni Franchi, Elisa Ricci, Subhankar Roy