arXiv Computer Vision
Sep 17

Beyond Performance Metrics: Uncertainty Mapping of Label Ambiguity in Fazekas Score Prediction

The paper introduces an uncertainty‑mapping framework for periventricular Fazekas score prediction, moving beyond standard performance metrics. By linking uncertainty to positions in the learned feature space, it identifies class‑boundary regions where predictions are ambiguous and highlights potential label disagreements. The study shows that different loss functions affect the uncertainty profile, with some models achieving clearer class separation and more localized uncertainty.

By Susanne Schmid, Johanna Ospel, Richard Frayne, Roberto Souza
arXiv Machine Learning
Aug 27

Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals

The study investigates how confidence intervals (CIs) behave in medical imaging AI by analyzing 24 segmentation and classification tasks with 19 models per task, various metrics, aggregation strategies, and CI methods. It finds that required sample sizes for reliable CIs vary widely, CI behavior depends on performance metrics, aggregation strategy, and problem type, and that different CI methods differ in reliability and precision. The authors provide a decision tree to guide researchers in selecting appropriate CI methods, aiming to support future consensus guidelines on reporting performance uncertainty.

By Pascaline Andr\'e (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Charles Heitz (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Evangelia Christodoulou (German Cancer Research Center), Annika Reinke (German Cancer Research Center), Carole H. Sudre (Unit for Lifelong Health and Ageing at UCL, Department of Population Science and Experimental Medicine and Hawkes InstituteCentre for Medical Image Computing, Department of Computer Science, University College London, UK), Michela Antonelli (School of Biomedical Engineering and Imaging Science, King's College London, UK), Patrick Godau (German Cancer Research Center), M. Jorge Cardoso (School of Biomedical Engineering and Imaging Science, King's College London, UK), Antoine Gilson (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Sophie Tezenas du Montcel (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Ga\"el Varoquaux (SODA project team, Inria Saclay-\^Ile-de-France, France), Lena Maier-Hein (German Cancer Research Center), Olivier Colliot (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France)
arXiv Machine Learning
Jul 30

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.

By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
arXiv Machine Learning
Sep 1

Uncertainty of Vision Medical Foundation Models

The paper investigates how pre‑training strategy, dataset size, and domain affect uncertainty estimation in vision medical foundation models. It compares point‑prediction calibration with conformal (region) prediction across retinal, histopathological, and chest X‑ray models, finding that domain‑specific, self‑supervised pre‑training yields better calibration and more efficient conformal sets. The study shows that standard recalibration alone cannot fully reconcile uncertainty differences between models trained on different data sources.

By Haoxu Huang, Narges Razavian
Hugging Face Trending Papers
4d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.