The paper introduces a biology-informed heterogeneous graph representation that models retinal vessel segments, intercapillary areas, and the foveal avascular zone to predict diabetic retinopathy stages from OCTA images. This graph-based approach reframes staging as a graph-level classification task solved with a graph neural network, achieving AUC-ROC values up to 84% and outperforming biomarker-based classifiers, CNNs, and vision transformers. The method also provides detailed, interpretable explanations by precisely localizing abnormal vessels and non-perfusion areas.
By Laurin Lux, Alexander H. Berger, Maria Romeo Tricas, Richard Rosen, Alaa E. Fayed, Sobha Sivaprasada, Linus Kreitner, Jonas Weidner, Martin J. Menten, Daniel Rueckert, Johannes C. Paetzold
arXiv:2512.13742v3 Announce Type: replace-cross
Abstract: Medical image classifiers detect gastrointestinal diseases well, but they do not explain their decisions. Large language models can generate...
By Md. Najib Hasan (Wichita State University, USA), Imran Ahmad (Wichita State University, USA), Sourav Basak Shuvo (Khulna University of Engineering and Technology, Bangladesh), Md. Mahadi Hasan Ankon (Khulna University of Engineering and Technology, Bangladesh), Nazmul Siddique (Ulster University, UK), Hui Wang (Queen's University Belfast, UK)
arXiv:2607. 04673v1 Announce Type: cross Abstract: Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer little clinical interpretability.
By Cheng Huang, Jia Zhang, Yi Jiang, Yang Liu, Karanjit Kooner, Yadi Liu, Tsengdar Lee, Yang Xie, Wenqi Shi, Guanghua Xiao
arXiv:2605. 22547v3 Announce Type: replace-cross Abstract: Medical image diagnosis has achieved significant progress with deep learning, yet existing methods often rely on isolated visual evidence and lack the ability to effectively leverage similar cases and external knowledge.
By Yiming Xu, Yixuan Liu, Yuhang Zhang, Ling Zheng, Yihan Wang, Qi Song
arXiv:2511. 18676v2 Announce Type: replace-cross Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.
By Yongcheng Yao, Yongshuo Zong, Raman Dutt, Yongxin Yang, Sotirios A Tsaftaris, Timothy Hospedales
arXiv:2605. 23995v4 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data.
By Chathura Wimalasiri, Kishor Nandakishor, Marimuthu Palaniswami
The paper introduces the Cross‑Modal Triage Network (CMTN), a multimodal deep‑learning model that fuses a Swin Transformer V2 visual encoder with a PubMedBERT text encoder to perform severity‑based triage, pathology detection, and generate visual explanations for chest radiographs. Trained on 34,639 image‑text pairs from MIMIC‑CXR‑JPG, the CMTN achieves high ordinal agreement with reference labels (QWK = 0.9341) and excellent pathology detection (macro‑AUROC = 0.9970) while operating with 34 ms latency. However, a blinded clinical audit revealed low agreement with expert radiologists (QWK = 0.1399) and only modest spatial‑semantic concordance in heatmaps, underscoring the gap between algorithmic performance and clinical judgment.
By Zinah Ghulam, Richa Mittal, Eranga Ukwatta
NV-Reason-CT is a generative vision‑language model designed for chest and abdominal CT analysis that preserves native 3D visual encoding and incorporates radiologist‑guided reasoning. The system couples a 3D vision transformer with a language model, feeding all visual tokens and their 3D coordinates directly into language decoding to maintain volumetric spatial information. Trained on a curated corpus of about 550,000 multimodal instruction examples, the model supports abnormality classification, report generation, and interactive reasoning, achieving strong performance on CT benchmarks and reducing expert interpretation time by 50%.
By Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
arXiv:2607.19261v4 Announce Type: replace-cross
Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating...
By Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin
arXiv:2608.22323v1 Announce Type: new
Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus...
By Lai Wei, Yuchao Chen, Zhenbiao Cao, Xiaojin Zhang, Zhongyu Wei, Bangting Wang, Wei Chen, Xiang Bai
OncoVision is a privileged‑information training framework that learns from mammography images and clinical data during training but performs inference using only mammographic images. It employs an attention‑based encoder‑decoder to jointly segment masses, calcifications, axillary findings, and breast tissue, and predicts ten structured clinical features such as BI‑RADS. Two late‑fusion strategies (Independent and Dependent) integrate imaging, radiomic, and clinical information to improve diagnostic precision, and a retrospective multi‑reader study showed higher diagnostic confidence, reduced reading time, and segmentation accuracy comparable to or better than radiologists.
By Istiak Ahmed, Galib Ahmed, K. Shahriar Sanjid, Md. Tanzim Hossain, Md. Nishan Khan, Md. Misbah Khan, Md. Arifur Rahman, Sheikh Anisul Haque, Sharmin Akhtar Rupa, Mohammed Mejbahuddin Mia, Mahmud Hasan Mostofa Kamal, Md. Mostafa Kamal Sarker, M. Monir Uddin