arXiv AI

How Much Can Reliability Drift Under a Fixed Confidence Distribution?

arXiv Machine Learning
Aug 4

Conformalized Large Language Models under Configuration Shift

arXiv:2608. 01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability.

By Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov
arXiv Machine Learning
Jun 18

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.

By Mame Diarra Toure, David A. Stephens
arXiv AI
1d ago

Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets

Signal‑Routed Temperature Scaling (SRTS‑BCE) is a 10‑parameter, argmax‑preserving calibration method that separates calibration objectives from adaptive capacity. It cross‑fits a correctness‑risk score over six logit statistics and assigns a top‑label BCE temperature to each of three risk groups, generalizing TvA‑TS when K=1. Experiments on fine‑tuned CIFAR‑100 and ViT‑B/16 show that SRTS‑BCE reduces ECE from 1.65 to 0.96 with a small calibration budget, outperforming higher‑capacity SMART+BCE when only 250 examples are available, and revealing a budget‑dependent ranking reversal on Swin‑T.

By Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv AI
Sep 2

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.

By Filippo Cenacchi, Longbing Cao, Runze Yang