arXiv AI By Yungeng Liu, Xuanzi Fang, Haijin Zeng, Qi Dai, Yongyong Chen

Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation

Read the original on arXiv AI →

arXiv:2607. 09481v1 Announce Type: cross Abstract: Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 25

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.

By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond
arXiv Computer Vision
Sep 16

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

BiCLIP is a bidirectional multimodal framework that enhances medical image segmentation by allowing visual features to iteratively refine textual representations, improving semantic alignment. It incorporates an augmentation consistency objective to stabilize learning against perturbed inputs. Experiments on QaTa-COV19 and MosMedData+ show that BiCLIP outperforms state‑of‑the‑art image‑only and multimodal baselines, achieving strong performance even with only 1% labeled data and resisting common clinical artifacts such as motion blur and low‑dose CT noise.

By Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mourad Oussalah
arXiv Computer Vision
Aug 25

Localization-Infused Vision-Language Semantic Fusion for Text-Guided Medical Image Segmentation

The paper introduces LoG, a localization‑infused vision‑language fusion framework for text‑guided medical image segmentation. LoG jointly performs multi‑scale target localization to explicitly capture target‑oriented semantics and employs three levels of localization‑infused fusion—feature, attention, and loss—to integrate spatial information into segmentation. Experiments on three benchmark datasets across three imaging modalities show that LoG consistently outperforms state‑of‑the‑art methods.

By Songyue Han, Mingye Zou, Shuchang Ye, Lei Bi, Mingyuan Meng