arXiv Machine Learning

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.

arXiv Machine Learning
Jul 7

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

arXiv:2607. 02569v1 Announce Type: cross Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using RETFound, a self-supervised vision-transformer retinal foundation model used here as a frozen feature encoder, and the public APTOS 2019 and DDR diabetic retinopathy fundus image datasets.

By Karim Mardhani
arXiv Machine Learning
Jul 24

Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

arXiv:2607. 21068v1 Announce Type: new Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reducing dependency on specialist-only assessment, for which neural network-based deep learning (DL) models have been widely utilized.

By Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee
arXiv Machine Learning
Aug 19

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

This study investigates whether vision‑based models for surgical skill assessment learn representations that transfer across different scoring rubrics (GOALS and OSATS) using the LASANA and JIGSAWS datasets. By evaluating end‑to‑end training, Adaptive Sharpness‑Aware Minimization, and self‑supervised/contrastive pretraining, the authors find that models pretrained on JIGSAWS can transfer reasonably well to LASANA, but transfer to JIGSAWS fails, likely due to annotation inconsistencies. Control experiments with a Kinetics‑pretrained backbone show that task‑specific heads carry most of the skill prediction load, while the backbone provides general spatiotemporal features.

By Hanna Hoffmann, Felix von Bechtolsheim, Stefanie Speidel, Rebecca Hisey
arXiv Machine Learning
Aug 11

Disentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert Routing

arXiv:2608. 09752v1 Announce Type: cross Abstract: Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation to every image regardless of the underlying disease distribution.

By Nagur Shareef Shaik, Jeongwoo Park, Yeong-Jin Kim, Jaeuk Jung, Hyunjung Oh, Dong Hye Ye