arXiv Computer Vision

One-Shot Adaptive Segmentation For Scientific Images

arXiv Machine Learning
Sep 14

Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images

The paper introduces Spatial‑FAD, a few‑shot medical anomaly detection framework that fuses Vision‑Language Model (CLIP) semantics with spatial priors from Vision Foundation Models (DINO). A VFM‑enhanced adapter injects structural affinity into CLIP features, while a sliding‑window aggregation produces high‑resolution embeddings for finer lesion localization. Prototype‑enhanced support memory further improves efficiency and performance, yielding significant gains on Liver CT, Retinal OCT, and Brain MRI datasets, notably an 11.4% Dice improvement in 4‑shot scenarios.

By Juzheng Miao, Yuchen Yuan, Cheng Chen, Pheng-Ann Heng
arXiv Computer Vision
Aug 26

Example-based Robust Abnormality Detection with Minimal Annotations using Exemplar Med-DETR

arXiv:2608.24281v1 Announce Type: new Abstract: Reducing annotation requirements remains a key challenge in developing robust medical object detectors. To address this, Vision-Language (VL) object de...

By Sheethal Bhat, Bogdan Georgescu, Awais Mansoor, Mathias Zinnen, Pranjal Sahu, Florin C. Ghesu, Sasa Grbic, Andreas Maier
arXiv Computer Vision
Sep 4

Exemplar: Classical Priors Complement Frozen Features for Few-Shot Microscopy Segmentation at Native Resolution

Exemplar is a few‑shot segmentation method that combines a frozen DINOv3 backbone with a fixed bank of classical native‑resolution filter responses in a single lightweight head. Trained only from support masks, it achieves a mean foreground intersection‑over‑union of 0.782 across eleven biomedical imaging datasets, outperforming either component alone and surpassing five other few‑shot methods in 54 of 55 comparisons. With a single annotated mask, Exemplar reaches 0.703, higher than a from‑scratch nnU‑Net trained on the same mask, and while nnU‑Net eventually overtakes it with eight masks, it requires 16–77× longer to fit.

By Michal Pr\r{u}\v{s}ek, Adam Novoz\'amsk\'y, Filip \v{S}roubek
arXiv Computer Vision
Sep 25

Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

The paper introduces a dual‑input, multi‑task learning framework that jointly segments and classifies bone tumors by applying bidirectional cross‑modal attention between a lesion crop and the full radiograph. Using a YOLO‑based detector and a dual‑stream DenseNet121 architecture, the model fuses fine‑grained lesion detail with global anatomical context through a novel cross‑modal attention fusion strategy and hierarchical multi‑scale feature fusion. On the multi‑institutional Bone Tumor X‑ray Radiograph Dataset, the approach outperforms single‑input baselines, achieving a Dice coefficient of 0.896 and a macro‑averaged F1‑score of 0.928, with an AUC of 0.999 for malignant osteosarcoma.

By S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul
arXiv Computer Vision
Sep 23

GAD-MambaUNet: Direction-Group Mamba with Gradient-Adaptive DINOv3 Distillation for Lightweight Medical Image Segmentation

The paper introduces GAD-MambaUNet, a lightweight medical image segmentation network that integrates efficient local modeling, Direction-Group Graph Selective Scan (DG‑GSS) for structured information exchange, and training‑time supervision from a frozen DINOv3 teacher with Gradient‑Adaptive Distillation. GAD‑MambaUNet demonstrates a strong accuracy‑efficiency trade‑off compared to other lightweight and general segmentation methods, and ablation studies confirm the benefits of DG‑GSS and DINOv3‑GAD supervision. Future work aims to refine teacher‑student alignment and apply the framework to multi‑class and multi‑modal medical segmentation tasks.

By Fang Wang, Huitao Li, Wenhan Chao, Zheng Zhuo, Xinxin Yang
arXiv Computer Vision
Aug 28

Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation

The paper introduces an unsupervised domain adaptation framework that aligns redundancy-reducing features to enable accurate 3D segmentation of cone-beam CT (CBCT) without target-domain annotations or inference-time adaptation. The method is architecture-agnostic, working with both CNN-based and ViT-based foundation models, and is evaluated on two liver segmentation benchmarks for interventional vascular procedures and radiation therapy. Results show that even large pretrained segmentation networks need explicit feature-space bridging to generalize across diagnostic CT and CBCT, and the proposed approach consistently outperforms existing pretrained foundation models and UDA strategies.

By Gauthier Miralles, Loic Le Folgoc, Vincent Jugnon, Pietro Gori
arXiv AI
Sep 2

GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation

GazeRefine is a training‑free framework that uses eye‑gaze data as an inference‑time prompt for zero‑shot medical image segmentation. It converts sparse, duration‑weighted fixations into foreground and background priors that initialize semantic prototypes in a frozen DINOv3 feature space, then iteratively refines these prototypes through discrimination, affinity propagation, and anchoring to the gaze guidance. The method achieves strong results on colonoscopy polyp segmentation and competitive performance on prostate MRI, demonstrating that gaze‑guided prototype refinement can enable segmentation without dense expert annotations or model fine‑tuning.

By Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri, Taifour Yousra, Bin Wang, Max Bengtsson, Gorkem Durak, Elif Keles, Zuheng Ming, Marek Penhaker, Azeddine Beghdadi, Ulas Bagci, Aladine Chetouani
arXiv Computer Vision
Sep 11

DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging

DINO-Med introduces a patch‑based framework that adapts natural‑image foundation models, specifically DINOv3, to multi‑modal medical imaging. The method uses training‑free registration, automated localization, and mask‑filtered patch extraction to aggregate patch‑level features into subject‑level diagnostics. In liver fibrosis staging, DINOv3 outperforms handcrafted radiomics, ResNet, and SAM‑Med2D features, achieving 78.4% accuracy for mild fibrosis (S1) and 75.8% for cirrhosis (S4) on the CARE 2025 cohort.

By Boya Wang, Ruizhe Li, Chao Chen, Xin Chen