arXiv:2608. 10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging.
By Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu
arXiv:2609.23139v1 Announce Type: new
Abstract: Most medical vision-language models (VLMs) excel at open-ended report generation and VQA but provide limited support for structured, fine-grained clini...
By Le Thien Phuc Nguyen, Thien Nguyen, Thanh-Huy Nguyen, Gia Minh Hoang, Anh Mai Vu, Ulas Bagci
arXiv:2608. 11335v1 Announce Type: cross Abstract: Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries.
By Md Maklachur Rahman, Tracy Hammond
Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.
By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond
arXiv:2608. 04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis.
By Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia
Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language fe...
SEG-SAM is a unified medical image segmentation model that builds on the Segment Anything Model (SAM) by integrating semantic medical knowledge. It introduces a semantic‑aware decoder separate from SAM’s original decoder to handle both semantic segmentation of prompted objects and classification of unprompted objects. The model also incorporates key medical category characteristics from large language models via a text‑to‑vision semantic module and uses a cross‑mask spatial alignment strategy to improve overlap between predictions, achieving superior performance over existing SAM‑based and task‑specific methods.
By Shuangping Huang, Hao Liang, Qingfeng Wang, Chulong Zhong, Zijian Zhou, Miaojing Shi
arXiv:2608. 13690v1 Announce Type: cross Abstract: Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context.
By Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy.
BiCLIP is a bidirectional multimodal framework that enhances medical image segmentation by allowing visual features to iteratively refine textual representations, improving semantic alignment. It incorporates an augmentation consistency objective to stabilize learning against perturbed inputs. Experiments on QaTa-COV19 and MosMedData+ show that BiCLIP outperforms state‑of‑the‑art image‑only and multimodal baselines, achieving strong performance even with only 1% labeled data and resisting common clinical artifacts such as motion blur and low‑dose CT noise.
By Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mourad Oussalah
arXiv:2608. 19825v1 Announce Type: cross Abstract: Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems.
By Yunseo Lee, Hyun Jun Kim, Heeseung Shin, Changwon Lim
arXiv:2608.23853v1 Announce Type: new
Abstract: The interpretation of endoscopic imagery in ulcerative colitis is complex and subjective, with variability in human assessment and subtle mucosal infla...
By Alexis Ivan Escamilla-Lopez, Gilberto Ochoa-Ruiz, Salvador Hinojosa, Sharib Ali