arXiv:2607. 19999v1 Announce Type: new Abstract: This Good Practice Guide presents work done in the QUMPHY project (Uncertainty quantification for machine learning models applied to photoplethysmography signals) that considered both machine learning and uncertainty quantification for problems which used photoplethysmography (PPG) signals from wearable devices as input.
By P. Harris, C. Bench, M. Rinkevi\v{c}ius, V. Marozas, L. Coquelin, A. Thompson, M. Nandi, U. Hackstein, P. J. Aston
The paper introduces a benchmark for out‑of‑distribution (OOD) detection in electroencephalography (EEG) machine learning, evaluates a wide range of OOD methods, and assesses their impact on two clinical downstream prediction tasks. It distinguishes between OOD detection and model uncertainty estimation, which are often conflated, and shows how combining complementary methods can create a robust safety net for deploying EEG‑based models in high‑risk settings.
By Philipp Bomatter, Henry Gouk
arXiv:2504. 18433v3 Announce Type: replace Abstract: Uncertainty quantification is crucial in machine learning, yet most (axiomatic) studies of uncertainty measures focus on classification, leaving a gap in regression settings with limited formal justification and evaluations.
By Christopher B\"ulte, Yusuf Sale, Timo L\"ohr, Paul Hofman, Gitta Kutyniok, Eyke H\"ullermeier
arXiv:2608. 28063v1 Announce Type: new Abstract: Multi-organ ultrasound classifiers increasingly combine attention, mixture-of-experts routing, uncertainty gating, and evidential deep learning (EDL) objectives to address heterogeneous anatomy and acquisition.
By Yang Song, Pengbo Sun, Shichang Feng, Ye Zhu, Xin Xu, Ziran Wang
arXiv:2607. 20582v1 Announce Type: cross Abstract: Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases.
By Frederik Hauke, Patrick Wienholt, Christiane Kuhl, Dyke Ferber, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn
The study investigates how confidence intervals (CIs) behave in medical imaging AI by analyzing 24 segmentation and classification tasks with 19 models per task, various metrics, aggregation strategies, and CI methods. It finds that required sample sizes for reliable CIs vary widely, CI behavior depends on performance metrics, aggregation strategy, and problem type, and that different CI methods differ in reliability and precision. The authors provide a decision tree to guide researchers in selecting appropriate CI methods, aiming to support future consensus guidelines on reporting performance uncertainty.
By Pascaline Andr\'e (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Charles Heitz (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Evangelia Christodoulou (German Cancer Research Center), Annika Reinke (German Cancer Research Center), Carole H. Sudre (Unit for Lifelong Health and Ageing at UCL, Department of Population Science and Experimental Medicine and Hawkes InstituteCentre for Medical Image Computing, Department of Computer Science, University College London, UK), Michela Antonelli (School of Biomedical Engineering and Imaging Science, King's College London, UK), Patrick Godau (German Cancer Research Center), M. Jorge Cardoso (School of Biomedical Engineering and Imaging Science, King's College London, UK), Antoine Gilson (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Sophie Tezenas du Montcel (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Ga\"el Varoquaux (SODA project team, Inria Saclay-\^Ile-de-France, France), Lena Maier-Hein (German Cancer Research Center), Olivier Colliot (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France)
arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.
By Mame Diarra Toure, David A. Stephens
The paper introduces egRUE, an explainable uncertainty estimation method that merges uncertainty quantification with feature‑level explanations for medical AI predictions. egRUE incorporates prediction explanations into its uncertainty calculation and decomposes uncertainty into contributions from individual features. Experiments and a user study with medical experts show that egRUE improves reliability, interpretability, and calibrated trust compared to existing methods.
By Li Rong Wang, Jamie Duell, Xinran Xu, Thomas C. Henderson, Yu Yue Hew, Pik Wan Erica Chiang, Xiao Wei Alstar Ang, Bingwen Eugene Fan, Xiuyi Fan
arXiv:2607. 24035v1 Announce Type: cross Abstract: Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for individual predictions may systematically misrepresent model behavior.
By Nils Gumpfer, Michael Guckert, Samuel Sossalla, Birgit A{\ss}mus, Jennifer Hannig
The paper investigates how pre‑training strategy, dataset size, and domain affect uncertainty estimation in vision medical foundation models. It compares point‑prediction calibration with conformal (region) prediction across retinal, histopathological, and chest X‑ray models, finding that domain‑specific, self‑supervised pre‑training yields better calibration and more efficient conformal sets. The study shows that standard recalibration alone cannot fully reconcile uncertainty differences between models trained on different data sources.
By Haoxu Huang, Narges Razavian
arXiv:2606. 11988v1 Announce Type: new Abstract: The distinction between aleatoric and epistemic uncertainty has received considerable attention in machine learning research, mainly in the context of supervised learning but also in other settings such as generative modeling.
By Yusuf Sale, Christopher B\"ulte, Felix Czaja, Joshua Stiller, Eyke H\"ullermeier
arXiv:2609.15180v1 Announce Type: new
Abstract: Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliabl...
By Mingcheng Zhu, Jinning Liang, Tingting Zhu