arXiv Machine Learning

Focal Calibration Loss: Controlling Posterior Distortion in Deep Neural Classifiers

arXiv Machine Learning
Jun 18

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.

By Mame Diarra Toure, David A. Stephens
arXiv Computer Vision
Sep 11

Confidence-Calibrating Regularization for Robust Brain MRI Segmentation Under Domain Shift

The paper introduces CalSAM, a lightweight adaptation framework that fine‑tunes only the mask decoder of the Segment Anything Model (SAM) while keeping its encoders frozen. CalSAM employs a Feature Fisher Information Penalty (FIP) to reduce encoder sensitivity to domain shift and a Confidence Misalignment Penalty (CMP) to curb overconfident voxel‑wise errors. Experiments on cross‑center, scanner‑shift, and motion‑corrupted brain MRI datasets show significant gains in Dice similarity coefficient, Hausdorff distance, and expected calibration error, with only a modest training‑time overhead.

By Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari
arXiv Computer Vision
Aug 28

MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models

MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.

By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan