Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
arXiv:2607.12896v3 Announce Type: replace
Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fr...
By Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye, Jianfeng Cao, Yixuan Yuan, Ting Ma
CoMLP introduces a cooperatively-gated MLP module that fuses multimodal medical data—such as imaging modalities and clinical reports—without relying on computationally heavy cross-attention. The module uses regional and dilated MLP interactions to capture both local and global cross-modal dependencies, enabling fine-grained fusion at high spatial resolutions. Experiments on five segmentation benchmarks, covering 2D/3D images and diverse anatomical regions, show consistent improvements over state-of-the-art multi-modal and language-guided methods, highlighting the effectiveness of MLP-based interaction for medical image segmentation.
By Mingyuan Meng, Shuchang Ye, Mingjian Li, Zhenyu Zhao, Jinman Kim, Lei Bi
InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.
By Ziquan Liu, Zhewei Zhu, Xuyang Shi
arXiv:2609.00272v1 Announce Type: new
Abstract: Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Mult...
By Paul Schneider, Nazim Haouchine
arXiv:2608.22679v1 Announce Type: new
Abstract: Semantic segmentation has rapidly advanced with deep learning; however, challenges remain in effectively capturing local and global contexts as well as...
By Changki Sung, Hyungtae Lim, Wanhee Kim, Youngwoo Seo, Hyun Myung
The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significant modal disparities caused by distinct imaging mechanisms. These intrinsic inconsistencies are prone to introducing pseudo-changes, thereby constraining detection accuracy.
arXiv:2608.30371v1 Announce Type: new
Abstract: Automatic cardiac image segmentation is pivotal for diagnosing and treating cardiac diseases. In this work, we introduce MCSeg, a volumetric transforme...
By Zhiyu Ye, Hairong Zheng, Tong Zhang
arXiv:2609.10261v1 Announce Type: new
Abstract: Multi-modal medical image segmentation leverages complementary diagnostic information, yet fusion can underperform single-modality baselines when spati...
By Yuchen Pei, Xiaoyu Hu, Yixiong Zou, Dingwen Hu, Hui Chu, Yutao Ma, Shijun Qiu, Gang Li
The paper introduces SARTM, a framework that adapts the Segment Anything Model (SAM) for RGB‑thermal (RGB‑T) semantic segmentation. It fine‑tunes SAM with LoRA layers, incorporates language guidance, and employs a Cross‑Modal Knowledge Distillation module to bridge modality gaps. The approach also modifies the segmentation head and adds an auxiliary semantic head, achieving superior performance on MFNET, PST900, and FMB benchmarks.
By Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
arXiv:2608.20944v1 Announce Type: new
Abstract: Multimodal object detection in remote sensing faces challenges due to semantic heterogeneity and modality-specific noise interference. To this end, we...
By Xin Wu, Zhenyu Gao, Qiankun Zhang, Shaoyong Guo
The paper introduces MSM‑Seg, a dual‑memory segmentation framework for 3D multi‑modal brain tumor segmentation. It combines a modality‑and‑slice memory attention module to capture cross‑modal and spatial‑slice dependencies, a multi‑scale category‑agnostic prompt encoder for whole‑tumor guidance, and a modality‑adaptive fusion decoder to integrate complementary decoding information. Experiments on various MRI datasets show that MSM‑Seg surpasses state‑of‑the‑art methods for metastases and glioma tumor segmentation.
By Yuxiang Luo, Qing Xu, Hai Huang, Yuqi Ouyang, Xiangjian He, Zhen Chen, Wenting Duan, Jiebo Luo