The paper evaluates uncertainty estimation (UE) methods for clinical vision‑language models (VLMs) on visual question answering (VQA). Across 8 UE techniques and 12 VLMs, UE quality tracks model accuracy, degrading where performance is weakest, and fails to signal uncertainty when models are stressed by hiding the correct answer (NOTA perturbations). However, UE on unperturbed inputs reliably predicts which predictions will collapse under NOTA, suggesting UE can diagnose model fragility.
By Arnisa Fazla, Alberto Testoni, Ameen Abu-Hanna, Barbara Plank, Iacer Calixto
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
By Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
arXiv:2606. 15910v2 Announce Type: replace Abstract: A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors.
By Reza Khanmohammadi, Kundan Thind, Mohammad M. Ghassemi
arXiv:2608. 16643v1 Announce Type: cross Abstract: Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation.
By Yifan Zhang, Rahmatollah Beheshti
The paper investigates how pre‑training strategy, dataset size, and domain affect uncertainty estimation in vision medical foundation models. It compares point‑prediction calibration with conformal (region) prediction across retinal, histopathological, and chest X‑ray models, finding that domain‑specific, self‑supervised pre‑training yields better calibration and more efficient conformal sets. The study shows that standard recalibration alone cannot fully reconcile uncertainty differences between models trained on different data sources.
By Haoxu Huang, Narges Razavian
arXiv:2205. 04599v2 Announce Type: replace-cross Abstract: Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explanations that are difficult for clinicians and patients to interpret.
By Mohammad Eslami, Solale Tabarestani, Saber Kazeminasab, Ehsan Adeli, Glyn Elwyn, Tobias Elze, Mengyu Wang, Nazlee Zebardast, Lucia Sobrin, Nassir Navab, Daniel Shu Wei Ting, Malek Adjouadi
arXiv:2608.22059v1 Announce Type: cross
Abstract: Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited....
By Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen
arXiv:2609.17753v1 Announce Type: new
Abstract: Reference labels used to train medical image classification models are not always as certain as they may appear, and this uncertainty has implications...
By Susanne Schmid, Johanna Ospel, Richard Frayne, Roberto Souza
arXiv:2509.19375v2 Announce Type: replace-cross
Abstract: Large language models are increasingly used for clinical text classification, where overconfident misclassifications can directly affect pati...
By Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil
arXiv:2608. 19807v1 Announce Type: new Abstract: Vision-language models (VLMs) can estimate physical quantities such as duration, speed, and acceleration from visual observations, but existing benchmarks primarily assess overall model performance against annotated ground truth.
By Rongyu Yu, Ke Niu, Fengxiang He
arXiv:2607. 26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment.
By Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Sch\"onlieb, Anna Korhonen, Anna Breger
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart.