arXiv AI

Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation

arXiv:2606. 04705v1 Announce Type: cross Abstract: Semantic segmentation in medical imaging is a critical yet challenging task due to data scarcity and high variability across modalities.

arXiv AI
Sep 15

MedSAM3: Delving into Segment Anything with Medical Concepts

MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.

By Anglin Liu, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen
arXiv Computer Vision
Aug 31

Prompt-Guided Interactive Segmentation of Interstitial Lung Disease in Thoracic CT

The paper introduces the first adaptation of the MedSAM2 foundation model for interactive 3D segmentation of interstitial lung disease (ILD) on thoracic CT scans. It evaluates three fine‑tuning strategies and four prompt types—bounding‑boxes, points, lassos, and scribbles—finding that full model fine‑tuning yields the best performance, improving Dice scores by 4.7 percentage points over the baseline. A proof‑of‑concept workflow is presented where MedSAM2 is first initialized with an automatic prior and then refined by radiologist prompts, with all resources released on GitHub.

By Vasilis Dedousis, Lubnaa Abdur Rahman, Lorenzo Brigat{\omicron}, Ethan Dack, Andreas Christe, Christoph Frank, Manuela Funke-Chambour, Justus Roos, Adrian Huber, Lukas Ebner, Stavroula Mougiakakou
arXiv AI
Aug 25

SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging

The paper introduces Segment Anything Small (SAS), a data‑augmentation method that improves deep‑learning segmentation of small anatomical structures in ultrasound images. SAS uses two transformations: resizing and embedding organ thumbnails into a black background to vary organ scale, and adding noise to regions of interest to mimic tissue texture variability. Experiments on one internal and five external datasets show Dice score gains up to 0.35, with an average improvement of 0.16, and demonstrate that SAS enhances model robustness and generalizability without adding hallucinations or artifacts.

By Danielle L. Ferreira, Ahana Gangopadhyay, Hsi-Ming Chang, Ravi Soni, Gopal Avinash
arXiv AI
Sep 2

GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation

GazeRefine is a training‑free framework that uses eye‑gaze data as an inference‑time prompt for zero‑shot medical image segmentation. It converts sparse, duration‑weighted fixations into foreground and background priors that initialize semantic prototypes in a frozen DINOv3 feature space, then iteratively refines these prototypes through discrimination, affinity propagation, and anchoring to the gaze guidance. The method achieves strong results on colonoscopy polyp segmentation and competitive performance on prostate MRI, demonstrating that gaze‑guided prototype refinement can enable segmentation without dense expert annotations or model fine‑tuning.

By Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri, Taifour Yousra, Bin Wang, Max Bengtsson, Gorkem Durak, Elif Keles, Zuheng Ming, Marek Penhaker, Azeddine Beghdadi, Ulas Bagci, Aladine Chetouani
arXiv AI
Aug 20

A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3

The study investigates how few expert-annotated cases are needed to fine‑tune MedSAM3 for abdominal organ segmentation using Low‑Rank Adaptation (LoRA). With only 10 annotated CT or MRI cases, the LoRA‑adapted models achieve performance comparable to specialist systems that require orders of magnitude more data, including reliable gallbladder segmentation and near‑state‑of‑the‑art results for liver, kidneys, and spleen. The approach also generalizes to cardiac segmentation on the Whole Heart dataset, and training takes only 3–5 hours per organ on a single GPU, roughly twice as fast as nnU-Net.

By Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot
arXiv Computer Vision
Aug 27

SEG-SAM: Semantic-Guided SAM for Unified Medical Image Segmentation

SEG-SAM is a unified medical image segmentation model that builds on the Segment Anything Model (SAM) by integrating semantic medical knowledge. It introduces a semantic‑aware decoder separate from SAM’s original decoder to handle both semantic segmentation of prompted objects and classification of unprompted objects. The model also incorporates key medical category characteristics from large language models via a text‑to‑vision semantic module and uses a cross‑mask spatial alignment strategy to improve overlap between predictions, achieving superior performance over existing SAM‑based and task‑specific methods.

By Shuangping Huang, Hao Liang, Qingfeng Wang, Chulong Zhong, Zijian Zhou, Miaojing Shi