The paper introduces LoG, a localization‑infused vision‑language fusion framework for text‑guided medical image segmentation. LoG jointly performs multi‑scale target localization to explicitly capture target‑oriented semantics and employs three levels of localization‑infused fusion—feature, attention, and loss—to integrate spatial information into segmentation. Experiments on three benchmark datasets across three imaging modalities show that LoG consistently outperforms state‑of‑the‑art methods.
By Songyue Han, Mingye Zou, Shuchang Ye, Lei Bi, Mingyuan Meng
arXiv:2608. 13690v1 Announce Type: cross Abstract: Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context.
By Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu
arXiv:2607. 09481v1 Announce Type: cross Abstract: Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture.
By Yungeng Liu, Xuanzi Fang, Haijin Zeng, Qi Dai, Yongyong Chen
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.
By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv:2505.03380v2 Announce Type: replace
Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains...
By Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li
arXiv:2601.06847v2 Announce Type: replace-cross
Abstract: Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We po...
By Mengmeng Zhang, Xiaoping Wu, Hao Luo, Fan Wang, Yisheng Lv
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
The paper investigates how much clinical text influences pixel‑level predictions in multimodal medical image segmentation. It shows that segmentation performance is largely insensitive to the choice of fusion module, but that the impact of text varies across datasets: removing text severely degrades performance on BUSI and BTMRI, while it has only a marginal effect on ISIC and Kvasir‑SEG. Using an Evidence Decoupling Decoder, the authors reveal that text mainly modulates global semantic context rather than spatial localization, and that the specific semantic components driving sensitivity differ by dataset.
By Ziquan Liu, Zhewei Zhu, Xuyang Shi
arXiv:2605. 15720v2 Announce Type: replace-cross Abstract: Medical referring image segmentation (MRIS) predicts lesion masks from medical images and natural-language referring expressions, but acquiring paired pixel-level annotations and referring texts is costly.
By Yuchen Li, Ziru Wei, Zhen Zhao, Yi Liu, Luping Zhou
arXiv:2608. 05683v1 Announce Type: cross Abstract: Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions.
By Jiaxuan Li, Qing Xu, Xiangjian He, Yue Li, Daokun Zhang, Fiseha B. Tesema, Rong Qu
arXiv:2608. 10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging.
By Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu