The paper presents a two‑pipeline framework for retinal fundus analysis that combines four‑class disease classification with vessel segmentation. It fine‑tunes eight ImageNet‑pretrained CNNs on the FIVES dataset, applies five gradient‑based explanation methods to assess model interpretability, and benchmarks ten U‑Net variants—including transformer‑based and attention‑enhanced architectures—on the FIVES and DRIVE datasets. The best classification results come from ResNet101 (94.17% accuracy), while the strongest segmentation performance is achieved by Attention U‑Net with a ResNet101V2 backbone, improving DRIVE IoU from 60.80% to 64.83%.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah
arXiv:2606. 16153v1 Announce Type: cross Abstract: Medical image segmentation plays a critical role in clinical diagnostics, treatment planning, disease monitoring, and neurological disorder identification.
By Pengyu Zhu, Xiaojing Zhang, Kunbo Zhang, Chunyan Zhang, Zhenyu Wang
The study evaluates four deep‑learning segmentation architectures—Unet, PSPNet, Linknet, and FPN—paired with six pre‑trained encoders to predict COVID‑19 lesions in CT images. Experiments on three COVID‑19 CT datasets show high accuracy, achieving a maximum binary F1‑score of 98% and multi‑class F1‑scores of 75% and 77%. The work aims to provide a standardized performance benchmark for medical image segmentation and a reference for other imaging scenarios.
By Sarmad Khan, Basim Azam, Arslan Shaukat
arXiv:2608.23745v1 Announce Type: cross
Abstract: Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, surgical guidance, and diseas...
By Mohammad Mahdi Danesh Pajouh, Sara Saeedi
OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.
By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
The paper introduces Report Supervision (R‑Super), a framework that uses radiology reports to supervise tumor segmentation models. By incorporating loss functions that align segmentation outputs with report‑derived tumor counts, sizes, and locations, R‑Super improves detection and segmentation performance. Experiments on kidney and pancreatic tumors show up to a 15% increase in F1‑Score and DSC compared to mask‑only training, outperforming methods like CLIP and multi‑task learning.
By Pedro R. A. S. Bassia, Wenxuan Li, Jakob Wasserthal, Jieneng Chen, Xinze Zhou, Zheren Zhu, Chuntung Zhuanga, Sergio Decherchi, Andrea Cavalli, Kang Wang, Yang Yang, Alan Yuille, Zongwei Zhou
The paper introduces a biology-informed heterogeneous graph representation that models retinal vessel segments, intercapillary areas, and the foveal avascular zone to predict diabetic retinopathy stages from OCTA images. This graph-based approach reframes staging as a graph-level classification task solved with a graph neural network, achieving AUC-ROC values up to 84% and outperforming biomarker-based classifiers, CNNs, and vision transformers. The method also provides detailed, interpretable explanations by precisely localizing abnormal vessels and non-perfusion areas.
By Laurin Lux, Alexander H. Berger, Maria Romeo Tricas, Richard Rosen, Alaa E. Fayed, Sobha Sivaprasada, Linus Kreitner, Jonas Weidner, Martin J. Menten, Daniel Rueckert, Johannes C. Paetzold
This study presents a clinically relevant framework for evaluating deep neural networks that segment lymphoma lesions in PET/CT images, addressing gaps such as out‑of‑distribution testing and comparison with expert annotators. Using 611 multi‑institutional cases, the authors assess four networks (ResUNet, SegResNet, DynUNet, SwinUNETR) with lesion‑specific metrics, detection criteria, and metabolic‑characteristic‑based thresholds, finding that models perform best on large, intense lesions. The work also demonstrates that network errors mirror those of physicians, highlighting shared challenges with small, faint lesions.
By Shadab Ahamed, Yixi Xu, Sara Kurkowska, Claire Gowdy, Joo H. O, Ingrid Bloise, Don Wilson, Patrick Martineau, Fran\c{c}ois B\'enard, Fereshteh Yousefirizi, Rahul Dodhia, Juan M. Lavista, William B. Weeks, Carlos F. Uribe, Arman Rahmim
The paper introduces a method for generalizable brain tumor segmentation in the BraTS 2026 Challenge. It builds on the nnU-Net framework with a large residual encoder, adding semi‑supervised learning via pseudo‑labels and a tumor‑aware deformable augmentation that locally deforms lesions while preserving surrounding anatomy. The approach improves Dice and NSD scores across all tumor regions compared to labeled‑only baselines, demonstrating the complementary benefits of self‑training and the proposed augmentation.
By Henrique Zan Grande, Jeovane Honorio Alves, Rayson Laroca, Andre Gustavo Hochuli
InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.
By Ziquan Liu, Zhewei Zhu, Xuyang Shi
The paper introduces FreNet, a feature reconfiguration framework that incorporates visual priors for medical lesion segmentation. FreNet performs pixel‑level reconfiguration before encoding using an Implicit Prior Neural Network (IPNN) that leverages SAM, and feature‑level reconfiguration during encoding via a Dual‑domain Feature Reconfiguration (DFR) module, which includes a Frequency Decoupling Module (FDM) and a Spatial Localization Module (SLM). Experiments on nine benchmarks across three imaging modalities show that FreNet outperforms state‑of‑the‑art methods, achieving a 5.0% Dice improvement over the best baseline on the ETIS dataset and a 7.2% improvement over SAM.
By Yinan Liu, Jiankang Hong, Zhen Gao, Ye Lu
The study evaluates fundus-specific foundation models (FM) for detecting diabetic macular edema (DME) in retinal images. It compares two popular FM—RETFound and FLAIR—against a lightweight EfficientNet-B0 backbone across multiple datasets (IDRiD, MESSIDOR-2, and OCT-and-Eye-FundusImages). Results indicate that FM do not consistently outperform fine‑tuned CNNs; EfficientNet-B0 often matches or exceeds FM performance, with FLAIR being the most competitive FM.
By Franco Javier Arellano, Jos\'e Ignacio Orlando