arXiv AI

FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration

arXiv:2607. 18278v1 Announce Type: cross Abstract: Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident despite being wrong.

arXiv AI
Sep 2

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.

By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv Computer Vision
Aug 27

Reliability analysis for BraTS-GoAT segmentation: a controlled robustness study of deep-ensemble uncertainty

The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.

By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv AI
Sep 12

Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration

The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.

By Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao
arXiv AI
4d ago

Alignment Forecasting: Predicting Misalignment From Training Data

The paper introduces Alignment Forecasting, a method for predicting whether fine‑tuning a language model on a given dataset will increase specific alignment failures such as deception or sycophancy. It presents ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions across many models, datasets, and failure modes, and shows that a simple forecasting scaffold using an LLM’s assessment of dataset bias can outperform baseline forecasters. The authors demonstrate that filtering out high‑risk training examples identified by the forecaster can improve alignment in multiple‑choice evaluations, though benefits in open‑ended conversations remain uncertain.

By Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
arXiv AI
Sep 4

Temperature Scaling Attack Disrupting Model Confidence in Federated Learning

The paper introduces the Temperature Scaling Attack (TSA), a training‑time method that degrades model confidence calibration while keeping predictive accuracy largely intact. TSA injects temperature scaling with a learning‑rate coupling during local federated training, shifting confidence scores and causing significant calibration errors (e.g., a 145% increase on CIFAR‑100) with less than a 2% drop in accuracy. The authors provide a convergence analysis for non‑IID settings and demonstrate TSA’s effectiveness across three benchmarks, robust aggregation, and post‑hoc calibration defenses, highlighting its impact on mission‑critical systems such as healthcare verification and autonomous driving.

By Kichang Lee, Jaeho Jin, JaeYeon Park, Songkuk Kim, JeongGil Ko