Pixel-level annotation remains a major bottleneck in medical image segmentation, making weak supervision an attractive yet under-constrained alternative. We propose OBBSeg, an intermediate supervision paradigm guided by Oriented Bounding Boxes (OBBs) that bridges the gap between full and weak supervision.
MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.
By Anglin Liu, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen
The paper introduces the first adaptation of the MedSAM2 foundation model for interactive 3D segmentation of interstitial lung disease (ILD) on thoracic CT scans. It evaluates three fine‑tuning strategies and four prompt types—bounding‑boxes, points, lassos, and scribbles—finding that full model fine‑tuning yields the best performance, improving Dice scores by 4.7 percentage points over the baseline. A proof‑of‑concept workflow is presented where MedSAM2 is first initialized with an automatic prior and then refined by radiologist prompts, with all resources released on GitHub.
By Vasilis Dedousis, Lubnaa Abdur Rahman, Lorenzo Brigat{\omicron}, Ethan Dack, Andreas Christe, Christoph Frank, Manuela Funke-Chambour, Justus Roos, Adrian Huber, Lukas Ebner, Stavroula Mougiakakou
arXiv:2509. 25594v2 Announce Type: replace-cross Abstract: Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented.
By Bangwei Guo, Yunhe Gao, Meng Ye, Difei Gu, Yang Zhou, Leon Axel, Dimitris Metaxas
arXiv:2608.22619v1 Announce Type: cross
Abstract: Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-t...
By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Mahmudul Hasan, Tracy Hammond
The paper introduces Segment Anything Small (SAS), a data‑augmentation method that improves deep‑learning segmentation of small anatomical structures in ultrasound images. SAS uses two transformations: resizing and embedding organ thumbnails into a black background to vary organ scale, and adding noise to regions of interest to mimic tissue texture variability. Experiments on one internal and five external datasets show Dice score gains up to 0.35, with an average improvement of 0.16, and demonstrate that SAS enhances model robustness and generalizability without adding hallucinations or artifacts.
By Danielle L. Ferreira, Ahana Gangopadhyay, Hsi-Ming Chang, Ravi Soni, Gopal Avinash
arXiv:2608.23745v1 Announce Type: cross
Abstract: Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, surgical guidance, and diseas...
By Mohammad Mahdi Danesh Pajouh, Sara Saeedi
arXiv:2607.12896v3 Announce Type: replace
Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fr...
By Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye, Jianfeng Cao, Yixuan Yuan, Ting Ma
arXiv:2509.22404v2 Announce Type: replace
Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
By Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun
GazeRefine is a training‑free framework that uses eye‑gaze data as an inference‑time prompt for zero‑shot medical image segmentation. It converts sparse, duration‑weighted fixations into foreground and background priors that initialize semantic prototypes in a frozen DINOv3 feature space, then iteratively refines these prototypes through discrimination, affinity propagation, and anchoring to the gaze guidance. The method achieves strong results on colonoscopy polyp segmentation and competitive performance on prostate MRI, demonstrating that gaze‑guided prototype refinement can enable segmentation without dense expert annotations or model fine‑tuning.
By Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri, Taifour Yousra, Bin Wang, Max Bengtsson, Gorkem Durak, Elif Keles, Zuheng Ming, Marek Penhaker, Azeddine Beghdadi, Ulas Bagci, Aladine Chetouani
The study investigates how few expert-annotated cases are needed to fine‑tune MedSAM3 for abdominal organ segmentation using Low‑Rank Adaptation (LoRA). With only 10 annotated CT or MRI cases, the LoRA‑adapted models achieve performance comparable to specialist systems that require orders of magnitude more data, including reliable gallbladder segmentation and near‑state‑of‑the‑art results for liver, kidneys, and spleen. The approach also generalizes to cardiac segmentation on the Whole Heart dataset, and training takes only 3–5 hours per organ on a single GPU, roughly twice as fast as nnU-Net.
By Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot
SEG-SAM is a unified medical image segmentation model that builds on the Segment Anything Model (SAM) by integrating semantic medical knowledge. It introduces a semantic‑aware decoder separate from SAM’s original decoder to handle both semantic segmentation of prompted objects and classification of unprompted objects. The model also incorporates key medical category characteristics from large language models via a text‑to‑vision semantic module and uses a cross‑mask spatial alignment strategy to improve overlap between predictions, achieving superior performance over existing SAM‑based and task‑specific methods.
By Shuangping Huang, Hao Liang, Qingfeng Wang, Chulong Zhong, Zijian Zhou, Miaojing Shi