Hugging Face Trending Papers

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

arXiv AI
Sep 15

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

arXiv:2609.15635v1 Announce Type: cross Abstract: A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a pai...

By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
arXiv Computation and Language
Sep 4

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

The paper introduces MedQA-MM, a benchmark that exposes shortcut reasoning in medical multimodal multiple-choice questions. By auditing prompts, images, and modalities, the authors show that models often rely on textual cues rather than visual evidence, with full-input accuracy at 62.63% but only 5.21% when restricted to text. The study highlights the need for route-level evidence to validate true medical image reasoning.

By Benlu Wang, Yifan Zhang, Jiaqing Yu, Chin Siang Ong, Juncheng Huang, Zhuohao Li, Zhenyu Zhang, Arman Cohan, Hong Yu, Zonghai Yao