The paper introduces LoG, a localization‑infused vision‑language fusion framework for text‑guided medical image segmentation. LoG jointly performs multi‑scale target localization to explicitly capture target‑oriented semantics and employs three levels of localization‑infused fusion—feature, attention, and loss—to integrate spatial information into segmentation. Experiments on three benchmark datasets across three imaging modalities show that LoG consistently outperforms state‑of‑the‑art methods.
By Songyue Han, Mingye Zou, Shuchang Ye, Lei Bi, Mingyuan Meng
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
The paper introduces MSM‑Seg, a dual‑memory segmentation framework for 3D multi‑modal brain tumor segmentation. It combines a modality‑and‑slice memory attention module to capture cross‑modal and spatial‑slice dependencies, a multi‑scale category‑agnostic prompt encoder for whole‑tumor guidance, and a modality‑adaptive fusion decoder to integrate complementary decoding information. Experiments on various MRI datasets show that MSM‑Seg surpasses state‑of‑the‑art methods for metastases and glioma tumor segmentation.
By Yuxiang Luo, Qing Xu, Hai Huang, Yuqi Ouyang, Xiangjian He, Zhen Chen, Wenting Duan, Jiebo Luo
arXiv:2607. 09481v1 Announce Type: cross Abstract: Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture.
By Yungeng Liu, Xuanzi Fang, Haijin Zeng, Qi Dai, Yongyong Chen
BiCLIP is a bidirectional multimodal framework that enhances medical image segmentation by allowing visual features to iteratively refine textual representations, improving semantic alignment. It incorporates an augmentation consistency objective to stabilize learning against perturbed inputs. Experiments on QaTa-COV19 and MosMedData+ show that BiCLIP outperforms state‑of‑the‑art image‑only and multimodal baselines, achieving strong performance even with only 1% labeled data and resisting common clinical artifacts such as motion blur and low‑dose CT noise.
By Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mourad Oussalah
arXiv:2601.09879v2 Announce Type: replace-cross
Abstract: Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report gen...
By Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.
By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv:2608.21786v2 Announce Type: replace
Abstract: General image fusion aims to integrate complementary information from multiple source images, but existing methods often rely on task-specific mode...
By Xingxin Xu, Siqi Zhao, Xin Li, Xinjie Yao, Yiming Sun, Pengfei Zhu
arXiv:2610.00279v1 Announce Type: new
Abstract: The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease p...
By Eirini Cholopoulou, Dimitrios E. Diamantis, Dimitris K. Iakovidis
Medical image segmentation relies on the ability of encoder-decoder architectures to translate rich feature representations into accurate pixel-level predictions under challenging conditions such as low contrast, structural ambiguity, and scale variability. While recent advances in large-scale pretraining and transformer-based encoders have substantially improved feature extraction, segmentation accuracy remains constrained by decoder design, particularly in terms of cross-scale alignment, contextual integration, and boundary preservation.
arXiv:2608. 20229v1 Announce Type: cross Abstract: Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts.
By Mosharof Hossain, Md Rabiul Islam, Limon Halder, Erchin Serpedin, Md Kamrul Hasan
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.