Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA
arXiv:2603. 24481v2 Announce Type: replace Abstract: Miscalibrated confidence scores are a practical obstacle to deploying AI in clinical settings.
arXiv:2607. 20582v1 Announce Type: cross Abstract: Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases.
arXiv:2603. 24481v2 Announce Type: replace Abstract: Miscalibrated confidence scores are a practical obstacle to deploying AI in clinical settings.
arXiv:2606. 15910v2 Announce Type: replace Abstract: A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors.
arXiv:2607. 16317v1 Announce Type: cross Abstract: Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic.
The paper introduces egRUE, an explainable uncertainty estimation method that merges uncertainty quantification with feature‑level explanations for medical AI predictions. egRUE incorporates prediction explanations into its uncertainty calculation and decomposes uncertainty into contributions from individual features. Experiments and a user study with medical experts show that egRUE improves reliability, interpretability, and calibrated trust compared to existing methods.
The study investigates how confidence intervals (CIs) behave in medical imaging AI by analyzing 24 segmentation and classification tasks with 19 models per task, various metrics, aggregation strategies, and CI methods. It finds that required sample sizes for reliable CIs vary widely, CI behavior depends on performance metrics, aggregation strategy, and problem type, and that different CI methods differ in reliability and precision. The authors provide a decision tree to guide researchers in selecting appropriate CI methods, aiming to support future consensus guidelines on reporting performance uncertainty.
The paper introduces MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations. MedDream outperforms existing diagnostic and generative AI models across eight clinical datasets, improving diagnostic reasoning, resident concordance, and evidence generation. Targeted synthetic augmentation guided by subgroup performance gaps further enhances model performance, particularly for Asian patients.
The paper investigates how pre‑training strategy, dataset size, and domain affect uncertainty estimation in vision medical foundation models. It compares point‑prediction calibration with conformal (region) prediction across retinal, histopathological, and chest X‑ray models, finding that domain‑specific, self‑supervised pre‑training yields better calibration and more efficient conformal sets. The study shows that standard recalibration alone cannot fully reconcile uncertainty differences between models trained on different data sources.
arXiv:2603.18792v3 Announce Type: replace Abstract: Uncertainty quantification (UQ) is crucial in safety-critical applications such as medical image segmentation. Total uncertainty is typically decom...
arXiv:2609.37848v1 Announce Type: cross Abstract: Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even...
arXiv:2609.15180v1 Announce Type: new Abstract: Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliabl...
The paper introduces ConRad, a reinforcement learning framework that fine‑tunes large vision‑language models to generate calibrated verbalized confidence estimates for radiology reports. ConRad offers both a single report‑level confidence score and a sentence‑level variant, trained with the GRPO algorithm and logarithmic scoring rewards to encourage truthful self‑assessment. Experiments show significant calibration improvements over existing methods, and clinical evaluation indicates that report‑level scores align well with clinicians’ judgments, enabling targeted review of low‑confidence statements.
arXiv:2508. 07617v2 Announce Type: replace-cross Abstract: AI has the potential to augment human decision making.