When Fusion Fails: Corruption-Aware Rebalanced Fusion for Multi-Modal Medical Image Segmentation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 00491v1 Announce Type: cross Abstract: Deep learning-based CT segmentation systems often achieve high accuracy on clean benchmark images, but their performance may degrade under heterogeneous clinical imaging conditions such as noise, resolution loss, contrast variation, intensity shift, and artifacts.
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
CoMLP introduces a cooperatively-gated MLP module that fuses multimodal medical data—such as imaging modalities and clinical reports—without relying on computationally heavy cross-attention. The module uses regional and dilated MLP interactions to capture both local and global cross-modal dependencies, enabling fine-grained fusion at high spatial resolutions. Experiments on five segmentation benchmarks, covering 2D/3D images and diverse anatomical regions, show consistent improvements over state-of-the-art multi-modal and language-guided methods, highlighting the effectiveness of MLP-based interaction for medical image segmentation.
arXiv:2606.08906v2 Announce Type: replace Abstract: In many binary segmentation tasks, most multimodal methods rely on fixed feature concatenation for cross-modal interaction and straightforward deco...
arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.
arXiv:2608. 19788v1 Announce Type: cross Abstract: Trustworthy multimodal fusion in clinical settings requires handling incomplete and heterogeneous modality subsets across institutions, where privacy constraints prohibit centralized data sharing.