arXiv:2608. 04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis.
By Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia
arXiv:2609.36557v1 Announce Type: cross
Abstract: Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reaso...
By Janet Wang, Yunbei Zhang, Xiao Wang, Jihun Hamm
arXiv:2608. 05960v1 Announce Type: cross Abstract: Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume.
By Maulik Chevli, Johannes Brandt, Rickmer Braren, Daniel Rueckert, Philip M\"uller
arXiv:2509.22404v2 Announce Type: replace
Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
By Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun
arXiv:2606. 12590v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved strong performance across medical imaging tasks, yet they remain prone to factual inconsistencies, poor visual grounding, and misalignment with clinically meaningful feedback.
By Shayan Mohammadizadehsamakosh, Pritam Sarkar, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv:2605. 18419v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology.
By Franciskus Xaverius Erick, Johanna Paula M\"uller, Bernhard Kainz
The paper introduces Spatial‑FAD, a few‑shot medical anomaly detection framework that fuses Vision‑Language Model (CLIP) semantics with spatial priors from Vision Foundation Models (DINO). A VFM‑enhanced adapter injects structural affinity into CLIP features, while a sliding‑window aggregation produces high‑resolution embeddings for finer lesion localization. Prototype‑enhanced support memory further improves efficiency and performance, yielding significant gains on Liver CT, Retinal OCT, and Brain MRI datasets, notably an 11.4% Dice improvement in 4‑shot scenarios.
By Juzheng Miao, Yuchen Yuan, Cheng Chen, Pheng-Ann Heng
arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.
By Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Klaus H. Maier-Hein, Panagiotis Korfiatis, Valentina Salvatelli, Javier Alvarez-Valle, Fernando P\'erez-Garc\'ia
arXiv:2608. 19825v1 Announce Type: cross Abstract: Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems.
By Yunseo Lee, Hyun Jun Kim, Heeseung Shin, Changwon Lim
arXiv:2511. 18676v2 Announce Type: replace-cross Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.
By Yongcheng Yao, Yongshuo Zong, Raman Dutt, Yongxin Yang, Sotirios A Tsaftaris, Timothy Hospedales
arXiv:2603. 07131v4 Announce Type: replace-cross Abstract: Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis.
By Shuai Lu, Meng Wang, Jia Guo, Jiawei Du, Bo Liu, Shengzhu Yang, Weihang Zhang, Huazhu Fu, Huiqi Li
The paper investigates how medical vision‑language models (VLMs) behave when faced with distribution shifts such as changes in acquisition domain, supervision, or evaluation protocol. Using datasets like NIH ChestXray14, CheXpert, PadChest, and OpenI, the authors isolate cross‑dataset visual transfer, evaluate multimodal alignment, and quantify source‑proxy leakage in frozen embeddings. They find that self‑supervised visual initialization improves transfer, adversarial adaptation is only marginally helpful, and that multimodal retrieval performance drops under external stress tests while source‑proxy information remains recoverable, highlighting hidden failure modes in medical VLMs.
By Ayoub Louaye Bouaziz, Lokmane Chebouba, Yassine Himeur