arXiv AI

Clinician-Friendly Foundation Models for Ophthalmic Image Diagnostics without Fine-Tuning or Technical Barriers

The paper introduces GlobeReady, a clinician-friendly platform that leverages the RetiGlobe foundation model for ophthalmic image diagnostics without requiring fine-tuning. RetiGlobe was pretrained in two stages: first with self-supervised learning on 38 million synthetic images, then with contrastive learning on 475,845 real image‑text pairs from diverse ethnicities, devices, and regions. GlobeReady was evaluated on 488,448 images from multiple international centers and tested prospectively with 31 ophthalmologists, also exploring domain generalisability, uncertainty quantification, OOD detection, and feature-based case retrieval.

arXiv Machine Learning
Aug 31

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.

By Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
arXiv AI
Jun 16

EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining

arXiv:2606. 15129v1 Announce Type: cross Abstract: Color fundus photography (CFP) is the mainstay for large-scale retinal screening, yet its diagnostic capacity is constrained by the lack of depth-resolved structural information.

By Zhuo Deng, Ruiheng Zhang, Ziheng Zhang, Weihao Gao, Yitong Li, Qian Wang, Lei Shao, Jiaoyue Dong, Zhixi Zeng, Lijian Fang, Haibo Wang, Xiaobin Lin, Tao Liu, Zhicheng Du, Zhengwei Zhang, Lin Yang, Zheng Gong, Xinyu Zhao, Zhenquan Wu, Fang Li, Zhiguang Zhou, Guoming Zhang, Sun Jing, Han Lv, Wenbin We, Lan Ma
arXiv AI
Sep 1

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

Co-Annotator is a clinical AI system that distills expert gaze and dictation into two guidance components: a gaze‑aligned Vision Transformer that highlights fixation‑aligned areas of interest (AOIs) and an ontology‑bounded vision‑language model that pre‑fills editable biomarker summaries for retinal OCT. In controlled studies, each modality independently improved diagnostic accuracy and biomarker generation, and when combined across two academic institutions, the system increased correct diagnoses per minute by 40% and reduced comment editing time by 67% without compromising accuracy.

By Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor
arXiv Computer Vision
Sep 3

Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection

The study evaluates fundus-specific foundation models (FM) for detecting diabetic macular edema (DME) in retinal images. It compares two popular FM—RETFound and FLAIR—against a lightweight EfficientNet-B0 backbone across multiple datasets (IDRiD, MESSIDOR-2, and OCT-and-Eye-FundusImages). Results indicate that FM do not consistently outperform fine‑tuned CNNs; EfficientNet-B0 often matches or exceeds FM performance, with FLAIR being the most competitive FM.

By Franco Javier Arellano, Jos\'e Ignacio Orlando
arXiv AI
Sep 25

Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study

The study developed a Vision Transformer-based deep learning model with uncertainty estimation to detect glaucoma from colour fundus photographs across multi‑ethnic populations, including those with high myopia. Using 56,483 images for training, the model achieved an internal AUROC of 98.7% and maintained high performance (AUROC 86.4–99.6%) on 16 external datasets from eight countries. In high‑myopia eyes, the model outperformed ophthalmologists and matched specialists when full clinical data were available.

By Raghavan Lavanya, Yangqin Feng, Ten Cheer Quek, Quan V. Hoang, Linda Yi-Chieh Poon, Jost B. Jonas, Ya Xing Wang, Vinay Nangia, Jin Wook Jeoung, Sehie Park, SoYeon Kim, Benjamin Y Xu, Sreenidhi Iyengar Munimadugu, Paul Mitchell, Gerald Liew, Yanin Suwan, Jirayu Hong-amata, Sahil Thakur, Monisha E Nongipur, Tina Wong, Rahat Husain, Ng Si Rui, Yamon Syn, Phey Feng Lo, Nicholas Tan Yi Qiang, Shaista Hussain, Xiaofeng Lei, Zhi Da Soh, Marco Yu, Haslina Hamzah, Zizhou Wang, Yan Wang, Liangli Zhen, Xinxing Xu, Tien-Yin Wong, Tin Aung, Rachel S Chong, Yong Liu, Ching-Yu Cheng
arXiv Computer Vision
Sep 23

Interpretable AI plus Handheld, Portable Retinal Photographs: A Low-Cost Glaucoma Screening Solution for West Africa

The study presents an interpretable AI framework for glaucoma screening using low-cost handheld retinal cameras in a West African population. Trained on 681 participants, the system achieved high performance across vessel segmentation, cup/disc segmentation, and optic nerve head feature detection, with classification AUCs of 0.85 for the handheld device and 0.93 for a tabletop camera. The model provides confidence scores and visual explanations to aid clinical interpretation.

By Charis Y. N. Chiang, Tarela Sarimiye, Adeyinka Ashaye, Martin Buist, Michael A. Hauser, Olusola Olawoye, Micha\"el J. A. Girard