arXiv AI

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

arXiv:2608. 13690v1 Announce Type: cross Abstract: Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context.

arXiv Computer Vision
Sep 10

Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks

arXiv:2505.03380v2 Announce Type: replace Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains...

By Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li
arXiv Computer Vision
Aug 25

Localization-Infused Vision-Language Semantic Fusion for Text-Guided Medical Image Segmentation

The paper introduces LoG, a localization‑infused vision‑language fusion framework for text‑guided medical image segmentation. LoG jointly performs multi‑scale target localization to explicitly capture target‑oriented semantics and employs three levels of localization‑infused fusion—feature, attention, and loss—to integrate spatial information into segmentation. Experiments on three benchmark datasets across three imaging modalities show that LoG consistently outperforms state‑of‑the‑art methods.

By Songyue Han, Mingye Zou, Shuchang Ye, Lei Bi, Mingyuan Meng
arXiv Computer Vision
Sep 16

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

BiCLIP is a bidirectional multimodal framework that enhances medical image segmentation by allowing visual features to iteratively refine textual representations, improving semantic alignment. It incorporates an augmentation consistency objective to stabilize learning against perturbed inputs. Experiments on QaTa-COV19 and MosMedData+ show that BiCLIP outperforms state‑of‑the‑art image‑only and multimodal baselines, achieving strong performance even with only 1% labeled data and resisting common clinical artifacts such as motion blur and low‑dose CT noise.

By Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mourad Oussalah
arXiv Machine Learning
Sep 25

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.

By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond
arXiv Computer Vision
Sep 2

Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT

Instance-Guided Report Anchoring (IGRA) is a model-agnostic module that links each abnormality instance in a chest CT to the corresponding finding in a radiology report during training, while discarding text components at inference. By reformulating free-text grounding as multi-label volumetric segmentation, IGRA allows all abnormality categories to be predicted in a single image-only forward pass. The method improves Dice scores by 22.5% over the strongest image-only baseline and matches state‑of‑the‑art performance on single-finding subsets, with consistent gains across multiple 3D segmentation backbones and datasets.

By Zhenyu Bu, Haoyan Ding, Chushu Shen, Xinyuan Zheng, Peiyu Duan, Xueqi Guo, Sepehr Farhand, Yoshihisa Shinagawa, Gerardo Hermosillo, Chaowei Wu
arXiv Machine Learning
Aug 4

Semi-MedRef: Semi-Supervised Medical Referring Image Segmentation with Cross-Modal Alignment

arXiv:2605. 15720v2 Announce Type: replace-cross Abstract: Medical referring image segmentation (MRIS) predicts lesion masks from medical images and natural-language referring expressions, but acquiring paired pixel-level annotations and referring texts is costly.

By Yuchen Li, Ziru Wei, Zhen Zhao, Yi Liu, Luping Zhou
arXiv AI
Jul 28

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.

By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang