The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.
By Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
This study presents four deep‑learning pipelines—two‑dimensional and three‑dimensional—for segmenting age‑related macular degeneration (AMD) and diabetic macular edema (DME) lesions in optical coherence tomography (OCT) images. The models achieve Dice scores between 0.76 and 0.82 and demonstrate strong volumetric and surface calibration (r_vol, r_surf ≥ 0.97) on an in‑domain validation set. Generalization was assessed on the OLIVES clinical cohort using proxy metrics such as biomarker AUROC, central subfield thickness correlation, and longitudinal concordance, showing that the predictions still track clinical biomarkers outside the training distribution, albeit with reduced strength.
By Lucia Sundberg, Zhihao Zhao, M. Ali Nasseri
arXiv:2607. 03959v1 Announce Type: cross Abstract: Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, highlighting the need for accurate and accessible screening tools.
By Rashadul Hasan Badhon, Atalie Carina Thompson, Jennifer I. Lim, Theodore Leng, Minhaj Nur Alam
arXiv:2607. 05825v1 Announce Type: cross Abstract: Background.
By Fred Mutisya, Oscar Onyango, Sarah Sitati, Syokau Ilovi, Aeesha NJ Malik, Brenda W'mosi, Brian Makini, Jalemba Aluuvala, Josiah Onyango, Rachael Kanguha Mmene, Steven Wanyee
arXiv:2509.25549v3 Announce Type: replace-cross
Abstract: Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical t...
By Mohammadmahdi Eshragh, Emad A. Mohammed, Behrouz Far, Ezekiel Weis, Carol L Shields, Sandor R Ferenczy, Trafford Crump
The study evaluates fundus-specific foundation models (FM) for detecting diabetic macular edema (DME) in retinal images. It compares two popular FM—RETFound and FLAIR—against a lightweight EfficientNet-B0 backbone across multiple datasets (IDRiD, MESSIDOR-2, and OCT-and-Eye-FundusImages). Results indicate that FM do not consistently outperform fine‑tuned CNNs; EfficientNet-B0 often matches or exceeds FM performance, with FLAIR being the most competitive FM.
By Franco Javier Arellano, Jos\'e Ignacio Orlando
The paper introduces GlobeReady, a clinician-friendly platform that leverages the RetiGlobe foundation model for ophthalmic image diagnostics without requiring fine-tuning. RetiGlobe was pretrained in two stages: first with self-supervised learning on 38 million synthetic images, then with contrastive learning on 475,845 real image‑text pairs from diverse ethnicities, devices, and regions. GlobeReady was evaluated on 488,448 images from multiple international centers and tested prospectively with 31 ophthalmologists, also exploring domain generalisability, uncertainty quantification, OOD detection, and feature-based case retrieval.
By Meng Wang, Tian Lin, Qingshan Hou, Aidi Lin, Lianyu Wang, Jingcheng Wang, Qingsheng Peng, Truong X. Nguyen, Zhi Da Soh, Xiayin Zhang, Jingyan Yang, Danqi Fang, Ke Zou, Ting Xu, Can Can Xue, Ten Cheer Quek, Qinkai Yu, Minxin Liu, Hui Zhou, Zixuan Xiao, Guiqin He, Huiyu Liang, Tingkun Shi, Man Chen, Zhuangling Lin, Linna Liu, Yuanyuan Peng, Li Jia Chen, Chi Ming Chan, Xiaohong Li, Junren He, Zhirong Xu, Tingbing Fang, Yanli Wang, Qingzhi Wang, Wenyi Hu, Yujie Wang, Li Li, Jiaying Ye, Tonghui Ye, Liang Lyu, Yongjian Lu, Ruoshi Cai, Yiwen Tang, Qiuming Hu, Junhong Chen, Zhenhua Zhang, Cheng Chen, Yitian Zhao, Dianbo Liu, Jianhua Wu, Xinjian Chen, Changqing Zhang, Xiaojun Wu, Triet Thanh Nguyen, Yanda Meng, Yalin Zheng, Daoqiang Zhang, Xiaochun Cao, Yih Chung Tham, Ye Zhang, Ying Han, Alvin L Young, Mary Ho, Carmen K M Chan, Clement C Tham, Zhuoting Zhu, Carol Y. Cheung, Tien Yin Wong, Huazhu Fu, Haoyu Chen, Ching-Yu Cheng
arXiv:2608. 09752v1 Announce Type: cross Abstract: Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation to every image regardless of the underlying disease distribution.
By Nagur Shareef Shaik, Jeongwoo Park, Yeong-Jin Kim, Jaeuk Jung, Hyunjung Oh, Dong Hye Ye
MVC-Bench is a new benchmark designed to evaluate the calibration of vision‑language models (VLMs) and medical VLMs (Medical‑VLMs) for medical image classification. It tests calibration across robustness to modality, backbone, and domain shift; effectiveness of calibration strategies and prompt‑tuning methods; and stability under prompt‑template and random‑seed variations. The benchmark includes eight backbones, three medical modalities (fundus imaging, histopathology, chest X‑ray), and compares post‑hoc, train‑time, and zero‑shot calibration approaches, reporting accuracy, Expected Calibration Error (ECE), Maximum Calibration Error (MCE), and Adaptive Calibration Error (ACE) over 1,638 experiments, while also proposing a Multi‑Class Margin (MCM) regularization technique that improves ECE in most settings.
By Ashshak Sharifdeen, Shihab Aaqil Ahamed, Ufaq Khan, Muhammad Akhtar Munir Sujair Ibrahim, Mohamed Rafeek Mareer Ahamed, Yutong Xie, Imran Razzak, Muhammad Haris Khan
Automated diabetic retinopathy (DR) grading from colour fundus photographs can achieve strong predictive performance, but clinical interpretation requires more than an image-level label. It requires understanding how lesion evidence is distributed around retinal vessels and how this evidence relates to quantitative vascular biomarkers.
arXiv:2603. 18846v3 Announce Type: replace-cross Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL).
By Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi, Philipp Berens
arXiv:2608. 07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.
By Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher