arXiv Machine Learning

Towards Interpretable Foundation Models for Retinal Fundus Images

arXiv:2603. 18846v3 Announce Type: replace-cross Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL).

Hugging Face Trending Papers
Aug 27

Domain-Specific Self-Supervised Representation Learning for Retinal Fundus Classification

The paper explores contrastive self‑supervised learning (SSL) for retinal fundus image classification, comparing SimSiam and SimCLR under limited data and computational resources. It investigates how retinal‑specific augmentation strategies and training parameters affect representation quality, evaluated through linear probing and fine‑tuning on multi‑disease classification and diabetic retinopathy grading tasks. The results demonstrate that tailored augmentations enable lightweight SSL models to learn transferable representations, reducing reliance on large annotated datasets while achieving competitive performance.

arXiv Machine Learning
Aug 28

Domain-Specific Self-Supervised Representation Learning for Retinal Fundus Classification

The paper explores contrastive self‑supervised learning (SSL) for retinal fundus image classification, comparing SimSiam and SimCLR under limited data and computational resources. It investigates how retinal‑specific augmentation strategies and training parameters affect representation quality, evaluated through linear probing and fine‑tuning on multi‑disease classification and diabetic retinopathy grading tasks. Results indicate that tailored augmentations enable lightweight SSL models to learn transferable representations, reducing reliance on large annotated datasets while achieving competitive performance.

By Bekzat Nurlanbekova, Fung Fung Ting
arXiv Computer Vision
Sep 2

Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications

Pix2Rep-v2 is a self‑supervised learning framework that learns pixel‑ and voxel‑level representations for dense medical imaging tasks, using a redundancy‑reduction objective and equivariance principles to scale to 3D and wide field‑of‑view data. The method is evaluated on four datasets across multiple modalities, tasks, and backbones, demonstrating higher data‑efficiency in few‑shot scenarios and competitive performance, such as a +9.3 Dice point improvement in one‑shot segmentation on the M&Ms‑2 dataset. An in‑context dense prototype approach is also proposed, eliminating the need for downstream training.

By S. Sifaoui, E. Angelini, S. Toupin, T. Pezel, L. Le Folgoc
arXiv Computer Vision
Sep 7

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.

By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv Computer Vision
Sep 3

Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection

The study evaluates fundus-specific foundation models (FM) for detecting diabetic macular edema (DME) in retinal images. It compares two popular FM—RETFound and FLAIR—against a lightweight EfficientNet-B0 backbone across multiple datasets (IDRiD, MESSIDOR-2, and OCT-and-Eye-FundusImages). Results indicate that FM do not consistently outperform fine‑tuned CNNs; EfficientNet-B0 often matches or exceeds FM performance, with FLAIR being the most competitive FM.

By Franco Javier Arellano, Jos\'e Ignacio Orlando