InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.
By Ziquan Liu, Zhewei Zhu, Xuyang Shi
arXiv:2608. 15537v1 Announce Type: cross Abstract: Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts.
By Wang Jiangtao, Nur Intan Raihana Ruhaiyem, Fu Panpan, Yang Yu, Huang Yan
The paper introduces MedSegLatDiff, a diffusion-based framework that combines a variational autoencoder (VAE) with a latent diffusion model for medical image segmentation. By compressing images into a low-dimensional latent space, the method reduces noise and speeds up training, while a weighted cross‑entropy loss preserves tiny structures such as small nodules. Evaluated on ISIC‑2018, CVC‑Clinic, and LIDC‑IDRI datasets, MedSegLatDiff achieves state‑of‑the‑art Dice and IoU scores, generates diverse segmentation hypotheses, and produces confidence maps that enhance interpretability and reliability for clinical deployment.
By Ngoc Huynh Trinh, Hai Toan Nguyen, Son Ba Luong, Quoc Long Tran
arXiv:1812.00877v2 Announce Type: replace-cross
Abstract: Segmentation of skin lesion boundaries in dermoscopic imaging is an important prerequisite step for computer-aided diagnosis of malignant mel...
By Glib Kechyn
MDSkin-Net is a multi‑task skin lesion analysis framework that integrates Pattern Analysis priors into a hybrid CNN‑Transformer architecture. It introduces a Pattern Analysis‑Guided Attention Module (PAGAM) with improved Efficient Channel Attention, Multi‑Scale Spatial Attention, and Biased Asymmetry Attention, along with a multi‑scale spatial alignment regularization that uses segmentation masks as soft supervision. Trained only on the ISIC 2017 training split, the model achieves high segmentation and classification performance on multiple datasets, demonstrating strong zero‑shot generalization across different cohorts.
By Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas, Nikolaos Papanikolopoulos
arXiv:2609.24116v1 Announce Type: new
Abstract: Although deep learning has advanced Whole Slide Image (WSI) Analysis, tissue artifacts like bubbles and folds often cause silent failures by concealing...
By Hyeseong Lee, Eunsu Kim, D M Bappy, Ho Heon Kim, Youngsuk Lee, Se Young Chun, Jang-Hwan Choi, Sung Hak Lee, Sangjeong Ahn
The paper introduces an unsupervised approach to medical image segmentation by training a Denoising Diffusion Probabilistic Model (DDPM) on 21 unlabeled abdominal CT scans to learn anatomical features. The encoder weights from the DDPM are transferred to a U‑Net for downstream segmentation on the BTCV multi‑organ dataset, resulting in a significant Dice score improvement for liver segmentation from 0.75 to 0.93. In low‑data regimes, diffusion‑pretrained models retain robust performance, achieving high Dice scores even with only 10% of labeled data.
By Akshat G, Divyansh Gupta, Shaleen Bhatnagar, Shilpa Ankalaki, Tusar Kanti Mishra
The paper introduces MIFR, a modality‑invariant and fair representation framework for skin disease classification that jointly processes clinical photographs and dermoscopic images using ViT‑based encoders. It employs a five‑component multi‑objective loss to balance classification accuracy, fairness across skin tones, class alignment, and modality invariance. Experiments on paired and external datasets demonstrate competitive predictive performance and fairness, with t‑SNE visualizations confirming alignment of embeddings from different modalities.
By Asonyu Senge Njih, Yvan Guifo Fodjo, Vianney Kengne Tchendji, Jerry Lacmou Zeutouo, Kerol Djoumessi
MoSSGate is a plug‑and‑play module for U‑Net that improves skin lesion segmentation by combining boundary‑aware spatial gating, an external memory modulator, and parallel 2D state‑space modeling for efficient global context aggregation. The design limits long‑range propagation to informative regions, adapts dynamically to each sample, and preserves sharp lesion boundaries while keeping computational cost low. Experiments on ISIC 2017 and 2018 show state‑of‑the‑art accuracy (86.3%/85.9% mIoU, 92.6%/90.6% Dice) with fewer FLOPs than most CNN baselines.
By Anum Awan, Mahnoor Buriro, Muhammad Younas Khan, Md Imam Ahasan
Acquiring pixel-level annotations for medical image segmentation is a severe bottleneck. Traditional U-Net architectures, while effective, learn local texture patterns and lack awareness of global ana...
Skin diseases represent a major global public health burden, yet machine learning tools developed to assist in their diagnosis suffer from two critical limitations: reliance on only one modality for d...
BruNet is a new cross‑domain transfer framework for automatic bruise segmentation that combines a ViT‑based visual encoder (either self‑supervised DINOv3 or pretrained LingBot‑Vision) with a SAM‑based mask decoder. The model is trained on the HAM10000 skin lesion dataset and evaluated on a separate bruise dataset without any fine‑tuning, achieving superior performance over CNN‑based models, state‑of‑the‑art segmentation models, ChatGPT‑4o/5‑assisted SAM2 zero‑shot baselines, and the medical‑oriented MedSAM. This work represents the first study to address pixel‑level localisation of bruises, demonstrating strong cross‑domain generalisation.
By Qiming Wang, Richard J. Motley, Ebube E. Obi, Xianfang Sun, Paul L. Rosin