arXiv AI

An automated method of identifying incorrectly labelled images based on the sequences of loss functions of deep learning networks

arXiv:2607. 02594v1 Announce Type: cross Abstract: Deep learning is widely applied in medical image analysis, but up to 10% of manually labelled images may be incorrect, degrading model performance.

arXiv Computer Vision
Sep 3

Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection

The study evaluates fundus-specific foundation models (FM) for detecting diabetic macular edema (DME) in retinal images. It compares two popular FM—RETFound and FLAIR—against a lightweight EfficientNet-B0 backbone across multiple datasets (IDRiD, MESSIDOR-2, and OCT-and-Eye-FundusImages). Results indicate that FM do not consistently outperform fine‑tuned CNNs; EfficientNet-B0 often matches or exceeds FM performance, with FLAIR being the most competitive FM.

By Franco Javier Arellano, Jos\'e Ignacio Orlando
arXiv Machine Learning
Jul 28

Long-Tailed Medical Image Classification

arXiv:2607. 23883v1 Announce Type: cross Abstract: In this paper, we examine the difficulties of using standard techniques for medical image classification due to long-tailed distributions (wherein rarer conditions have very few samples) resulting in bias towards diagnosing common diseases and away from rarer diseases.

By Nathanael Ren, Saagar Arya
arXiv Machine Learning
Jul 8

Reliable Mislabel Detection for Video Capsule Endoscopy Data

arXiv:2602. 06938v2 Announce Type: replace-cross Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets.

By Julia Werner, Julius Oexle, Oliver Bause, Maxime Le Floch, Franz Brinkmann, Hannah Tolle, Jochen Hampe, Oliver Bringmann
arXiv Machine Learning
Aug 31

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.

By Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
arXiv AI
Jul 7

Comparison of Loss Functions for Robust Deep Learning-based Echocardiography Segmentation when Learning with Partially Labelled Data from Multiple Domains

arXiv:2607. 05008v1 Announce Type: cross Abstract: Echocardiography is the first imaging modality used for assessing cardiac function, and accurate segmentation of cardiac structures is essential for deriving biomarkers.

By Iman Islam, Esther Puyol-Ant\'on, Bram Ruijsink, Andrew J. Reader, Andrew P. King
arXiv AI
Jun 16

Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection

arXiv:2606. 16868v1 Announce Type: cross Abstract: While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections such as contour disagreement, missing or additional structures, and confused labels.

By Markus Bujotzek, Dimitrios Bounias, Stefan Denner, Ralf Floca, Maximilian Fischer, Peter Neher, Klaus Maier-Hein
arXiv Computer Vision
Sep 18

Performance of Machine Learning Classification in Sonomammogram Images using BI-RADS

This study evaluates the classification accuracy of six modern deep‑learning architectures—VGG19, ResNet50, GoogleNet, ConvNeXt, EfficientNet, and Vision Transformers—on breast ultrasound images categorized by BI‑RADS. Using 2,945 training images and 936 validation images from 1,540 patients, the models were tested in full fine‑tuning, linear evaluation, and training‑from‑scratch settings. The best performance was achieved with full fine‑tuning, yielding 76.39 % accuracy and a 67.94 % F1 score.

By Malitha Gunawardhana, Norbert Zolek
arXiv Machine Learning
Sep 17

Interpretable Retinal Disease Prediction Using Biology-Informed Heterogeneous Graph Representations

The paper introduces a biology-informed heterogeneous graph representation that models retinal vessel segments, intercapillary areas, and the foveal avascular zone to predict diabetic retinopathy stages from OCTA images. This graph-based approach reframes staging as a graph-level classification task solved with a graph neural network, achieving AUC-ROC values up to 84% and outperforming biomarker-based classifiers, CNNs, and vision transformers. The method also provides detailed, interpretable explanations by precisely localizing abnormal vessels and non-perfusion areas.

By Laurin Lux, Alexander H. Berger, Maria Romeo Tricas, Richard Rosen, Alaa E. Fayed, Sobha Sivaprasada, Linus Kreitner, Jonas Weidner, Martin J. Menten, Daniel Rueckert, Johannes C. Paetzold