arXiv AI

MedSAM3: Delving into Segment Anything with Medical Concepts

MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.

arXiv AI
Aug 20

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

The paper introduces MedUAG, a unified medical multimodal model that supports both understanding and generation tasks. It presents MedUAGCorpus, the largest dataset of over 6 million instances across 14 imaging modalities, and MedUAGBench, a benchmark covering 12 diverse generation tasks with standardized protocols. Experiments show that MedUAG performs strongly across many medical understanding and generation tasks, setting a competitive baseline for future medical multimodal systems.

By Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu
arXiv Computer Vision
Aug 27

SEG-SAM: Semantic-Guided SAM for Unified Medical Image Segmentation

SEG-SAM is a unified medical image segmentation model that builds on the Segment Anything Model (SAM) by integrating semantic medical knowledge. It introduces a semantic‑aware decoder separate from SAM’s original decoder to handle both semantic segmentation of prompted objects and classification of unprompted objects. The model also incorporates key medical category characteristics from large language models via a text‑to‑vision semantic module and uses a cross‑mask spatial alignment strategy to improve overlap between predictions, achieving superior performance over existing SAM‑based and task‑specific methods.

By Shuangping Huang, Hao Liang, Qingfeng Wang, Chulong Zhong, Zijian Zhou, Miaojing Shi
arXiv AI
Sep 15

Concept-Grounded Reasoning with Prompt-Driven Localization for Interpretable Structured Report Generation

The paper introduces CORAL, a multimodal framework that combines spatial grounding and concept-level supervision for medical report generation. CORAL uses a prompt-driven segmentation model to localize lesions and a Concept Bottleneck module to predict multi-class clinical attributes, feeding these textual concept tokens and mask-modulated visual features into a multimodal large language model. Experiments on BUS-CoT and IU X-ray datasets show that CORAL improves diagnostic accuracy, concept consistency, and report quality compared to existing general-purpose and medical MLLMs.

By Xinyue Xu, Hongbin Lin, Juangui Xu, Hualiang Wang, Lehan Wang, Lijie Hu, Weiyang Liu, Adrian Weller, Xiaomeng Li