Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence.
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
The study evaluates fundus-specific foundation models (FM) for detecting diabetic macular edema (DME) in retinal images. It compares two popular FM—RETFound and FLAIR—against a lightweight EfficientNet-B0 backbone across multiple datasets (IDRiD, MESSIDOR-2, and OCT-and-Eye-FundusImages). Results indicate that FM do not consistently outperform fine‑tuned CNNs; EfficientNet-B0 often matches or exceeds FM performance, with FLAIR being the most competitive FM.
By Franco Javier Arellano, Jos\'e Ignacio Orlando
arXiv:2606. 16234v1 Announce Type: cross Abstract: Fundus fluorescein angiography (FFA) is critical for assessing retinal vascular abnormalities, but its acquisition is invasive and not always feasible.
By Tengfei Ma, Ruiqi Wu, Chenran Zhang, Ye Geng, Na Su, Xiangyuan Duanmu, Tao Zhou, Yi Zhou, Wen Fan
arXiv:2608. 09752v1 Announce Type: cross Abstract: Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation to every image regardless of the underlying disease distribution.
By Nagur Shareef Shaik, Jeongwoo Park, Yeong-Jin Kim, Jaeuk Jung, Hyunjung Oh, Dong Hye Ye
arXiv:2607. 03959v1 Announce Type: cross Abstract: Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, highlighting the need for accurate and accessible screening tools.
By Rashadul Hasan Badhon, Atalie Carina Thompson, Jennifer I. Lim, Theodore Leng, Minhaj Nur Alam
arXiv:2608. 10316v1 Announce Type: cross Abstract: Multi-modal learning combining medical images and clinical text is promising for disease diagnosis.
By Zijian Gu, Weikai Lin, Shuang Zhou, Zihan Chen, Song Wang
arXiv:2603. 18846v3 Announce Type: replace-cross Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL).
By Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi, Philipp Berens
arXiv:2609.06699v1 Announce Type: cross
Abstract: Glaucoma is a progressive optic neuropathy characterized by irreversible damage to the optic nerve, making timely diagnosis critical to prevent perma...
By Abdullah Al Shafi, Nishat Sadaf Lira, Abrar Hasan, Kazi Saeed Alam, Swapnil Kundu Argha
arXiv:2608.30824v1 Announce Type: new
Abstract: Combining whole-body magnetic resonance imaging (WB-MRI) with clinical variables has the potential to improve systemic disease diagnosis by leveraging...
By Laura Daza, Marta Hasny, Cristina Gonz\'alez, Julia A. Schnabel
OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.
By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
arXiv:2609.22271v1 Announce Type: new
Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes...
By Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert