arXiv:2603. 28387v2 Announce Type: replace Abstract: Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts.
By Doan Nam Long Vu, Simone Balloccu
arXiv:2609.00593v1 Announce Type: new
Abstract: Neuroradiologists rarely read a brain MRI in isolation, yet automated brain-MRI report generation has been built almost entirely for single studies. Te...
By Krish Patel, Peirong Liu
The paper investigates how medical vision‑language models (VLMs) behave when faced with distribution shifts such as changes in acquisition domain, supervision, or evaluation protocol. Using datasets like NIH ChestXray14, CheXpert, PadChest, and OpenI, the authors isolate cross‑dataset visual transfer, evaluate multimodal alignment, and quantify source‑proxy leakage in frozen embeddings. They find that self‑supervised visual initialization improves transfer, adversarial adaptation is only marginally helpful, and that multimodal retrieval performance drops under external stress tests while source‑proxy information remains recoverable, highlighting hidden failure modes in medical VLMs.
By Ayoub Louaye Bouaziz, Lokmane Chebouba, Yassine Himeur
arXiv:2608.28714v1 Announce Type: cross
Abstract: Objective: Deep learning accelerates brain MRI four- to tenfold, but models can erase lesions or synthesize false tissue - failures pixel-averaged me...
By Dat Tat Mai, Thai Viet Pham, Thu Nguyen Thi Dang, James Jin Kang
arXiv:2606. 10066v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) are evaluated on public benchmarks whose images and question-answer pairs have been freely downloadable for years, yet reported accuracy assumes these examples were absent from pretraining.
By Bruce Changlong Xu, Lan Wu, Alexander Ryu
arXiv:2608. 06429v1 Announce Type: cross Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior.
By Yong Yang, Roger Newman-Norlund, Xiang Guan, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Sophie Arheix-Parras, Srihari Nelakuditi, Leonardo Bonilha, Christopher Rorden, Rutvik H. Desai, Julius Fridriksson
arXiv:2608. 12689v1 Announce Type: cross Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically.
By Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
arXiv:2606. 00123v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks.
By Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su, Taiping Qu, Jingwei Guo, Nan Zhang, Hui Wang, Zhen Zhou, Kairui Bo, Yan Chen, Yue Ren, Shuai Li, Lei Xu, Henggui Zhang
arXiv:2607. 15047v1 Announce Type: cross Abstract: Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated detection from neuropsychological drawing tests remains fundamentally constrained by data scarcity, class imbalance, and diagnostic ambiguity near clinical boundaries.
By Javad Khoramdel, Farhad Hoseyni, Amirhossein Nikoofard
arXiv:2605. 18419v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology.
By Franciskus Xaverius Erick, Johanna Paula M\"uller, Bernhard Kainz
MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.
By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan
The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.
By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson