SCDM: Spatial-Contextual Disentanglement Mamba via Differential Inference for Efficient Image Classification
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 14355v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal representations capturing shared morpho-transcriptomic structure.
The paper introduces MedIDL, a Medical Imaging Disentanglement Learning framework that separates disease-related features from confounding covariates and individual variability in medical images. It achieves this by projecting image features into three orthogonal latent spaces—disease classification, covariate alignment, and a Gaussian head for individual variation—using specialized disentanglement heads. Across seven diverse imaging datasets, MedIDL surpasses state‑of‑the‑art supervised and self‑supervised methods in classification accuracy, and its latent representations and gradient‑based visualizations align with known clinical patterns.
arXiv:2509.22404v2 Announce Type: replace Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
arXiv:2607. 03103v1 Announce Type: cross Abstract: Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and precluding knowledge sharing across anatomically related tasks.
arXiv:2603.13044v2 Announce Type: replace-cross Abstract: Medical image segmentation (MIS) is a fundamental component of computer-assisted diagnosis and clinical decision support. Over the past decad...
InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.