Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
arXiv:2603. 21693v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have shown strong potential for medical Visual Question Answering (VQA), yet they remain prone to hallucinations, defined as generating responses that contradict the input image, posing serious risks in clinical settings.
The paper introduces ConRad, a reinforcement learning framework that fine‑tunes large vision‑language models to generate calibrated verbalized confidence estimates for radiology reports. ConRad offers both a single report‑level confidence score and a sentence‑level variant, trained with the GRPO algorithm and logarithmic scoring rewards to encourage truthful self‑assessment. Experiments show significant calibration improvements over existing methods, and clinical evaluation indicates that report‑level scores align well with clinicians’ judgments, enabling targeted review of low‑confidence statements.
arXiv:2609.06419v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-base...
PROVE is a black‑box hallucination detector for medical visual question answering that tailors its verification strategy to each question’s evidential structure. It classifies questions into three regimes, activates a subset of five operators per regime, and calibrates operator importance using deterministic question‑answer features to produce a risk score. On 8048 test samples across three medical VQA benchmarks and four state‑of‑the‑art vision‑language models, PROVE achieves an AUROC of 0.821, surpassing the best baseline by 0.159 with consistent improvements across all models and datasets.
The paper evaluates uncertainty estimation (UE) methods for clinical vision‑language models (VLMs) on visual question answering (VQA). Across 8 UE techniques and 12 VLMs, UE quality tracks model accuracy, degrading where performance is weakest, and fails to signal uncertainty when models are stressed by hiding the correct answer (NOTA perturbations). However, UE on unperturbed inputs reliably predicts which predictions will collapse under NOTA, suggesting UE can diagnose model fragility.