arXiv Machine Learning

Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models

arXiv:2604. 08941v2 Announce Type: replace Abstract: Medical Vision-Language Models (VLMs) answering binary presence questions on chest radiographs can fail in two linked ways: they are confidently wrong, and they change answers when a clinically equivalent question is rephrased.

arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv AI
Sep 15

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

arXiv:2609.15635v1 Announce Type: cross Abstract: A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a pai...

By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
arXiv Machine Learning
Sep 18

Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction

This study independently reproduces the dissociation reported by Zhao (2026) regarding chain-of-thought entropy in large language models. It confirms that the shape of the entropy trajectory predicts answer correctness, while the total entropy drop magnitude does not, across four open-weight models and two benchmarks (GSM8K and MATH‑500). The reproduction also maps settings where the magnitude signal holds or fails and documents protocol differences not reported in the original work.

By Theodore O. Cochran
arXiv AI
Jul 21

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain

arXiv:2603. 21693v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have shown strong potential for medical Visual Question Answering (VQA), yet they remain prone to hallucinations, defined as generating responses that contradict the input image, posing serious risks in clinical settings.

By Mohammad Asadi, Tahoura Nedaee, Jack W. O'Sullivan, Euan Ashley, Ehsan Adeli