arXiv Computer Vision
Sep 23

Decoupling Disease, Covariates, and Individual Variability: A Unified Disentanglement Framework for Medical Image Classification

The paper introduces MedIDL, a Medical Imaging Disentanglement Learning framework that separates disease-related features from confounding covariates and individual variability in medical images. It achieves this by projecting image features into three orthogonal latent spaces—disease classification, covariate alignment, and a Gaussian head for individual variation—using specialized disentanglement heads. Across seven diverse imaging datasets, MedIDL surpasses state‑of‑the‑art supervised and self‑supervised methods in classification accuracy, and its latent representations and gradient‑based visualizations align with known clinical patterns.

By Shengjie Zhang, Jinglin Zhang, Zhuangzhuang Jiang, Ziqi Yu, Yipin Zhang, Qi Zhang, Xiang Chen, Haibo Yang, Fei Gao, Longbiao Cui, Yuan Zhou, Xiao-Yong Zhang, Alzheimer's Disease Neuroimaging Initiative
arXiv AI
Sep 3

InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation

InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.

By Ziquan Liu, Zhewei Zhu, Xuyang Shi