arXiv Computer Vision

Robust retinal biometrics for patient identity verification and retrieval across age and imaging devices

arXiv Computer Vision
Aug 25

When the Edit Changes the Patient: Measuring Identity Preservation in Counterfactual Retinal Images

arXiv:2608.23024v1 Announce Type: new Abstract: Counterfactual medical image generation aims to modify an existing image to reflect a hypothetical scenario in which certain characteristics of the ima...

By Andrea Posada, Wenke Karbole, Bach Ngoc Doan, Alexander Weers, Solmaz Abdolrahimzadeh, Maria Patsiamanidi, Kahkashan Haider, Vaishali Khare, Daniel Rueckert, Andrew Lotery, Sobha Sivaprasad, Martin J. Menten
arXiv AI
Sep 1

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

Co-Annotator is a clinical AI system that distills expert gaze and dictation into two guidance components: a gaze‑aligned Vision Transformer that highlights fixation‑aligned areas of interest (AOIs) and an ontology‑bounded vision‑language model that pre‑fills editable biomarker summaries for retinal OCT. In controlled studies, each modality independently improved diagnostic accuracy and biomarker generation, and when combined across two academic institutions, the system increased correct diagnoses per minute by 40% and reduced comment editing time by 67% without compromising accuracy.

By Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor
arXiv AI
Jun 16

EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining

arXiv:2606. 15129v1 Announce Type: cross Abstract: Color fundus photography (CFP) is the mainstay for large-scale retinal screening, yet its diagnostic capacity is constrained by the lack of depth-resolved structural information.

By Zhuo Deng, Ruiheng Zhang, Ziheng Zhang, Weihao Gao, Yitong Li, Qian Wang, Lei Shao, Jiaoyue Dong, Zhixi Zeng, Lijian Fang, Haibo Wang, Xiaobin Lin, Tao Liu, Zhicheng Du, Zhengwei Zhang, Lin Yang, Zheng Gong, Xinyu Zhao, Zhenquan Wu, Fang Li, Zhiguang Zhou, Guoming Zhang, Sun Jing, Han Lv, Wenbin We, Lan Ma
arXiv Machine Learning
Jul 24

Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

arXiv:2607. 21068v1 Announce Type: new Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reducing dependency on specialist-only assessment, for which neural network-based deep learning (DL) models have been widely utilized.

By Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee
arXiv Computation and Language
Sep 7

Retinal OCTA Phenotyping with LLM Reporting for Alzheimer's Disease

The study introduces an explainable OCTA pipeline for Alzheimer’s disease phenotyping that combines vessel segmentation, layer‑specific biomarker extraction, and label‑free phenotyping with large language model (LLM) reporting. Using 117 images from 39 subjects, the segmentation models achieved high ROC‑AUC (0.916–0.970) and Dice scores (0.695–0.781), and six vascular biomarkers were used to create subject‑level profiles for exploratory clustering. LLMs (GPT, Gemini, Llama) produced measurement‑grounded reports evaluated for citation faithfulness and diagnostic caution, offering a transparent, non‑diagnostic link between retinal vascular data and Alzheimer’s research.

By Progga Paromita Dutta, Jeba Maliha, Md Rafiul Kabir
arXiv Machine Learning
Aug 31

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.

By Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz