One-Shot Adaptive Segmentation For Scientific Images
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 24453v1 Announce Type: cross Abstract: Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly.
arXiv:2606. 17972v1 Announce Type: cross Abstract: Self-supervised DINO models provide strong transferable visual representations, yet applying them directly to image segmentation remains challenging.
The paper introduces Spatial‑FAD, a few‑shot medical anomaly detection framework that fuses Vision‑Language Model (CLIP) semantics with spatial priors from Vision Foundation Models (DINO). A VFM‑enhanced adapter injects structural affinity into CLIP features, while a sliding‑window aggregation produces high‑resolution embeddings for finer lesion localization. Prototype‑enhanced support memory further improves efficiency and performance, yielding significant gains on Liver CT, Retinal OCT, and Brain MRI datasets, notably an 11.4% Dice improvement in 4‑shot scenarios.
arXiv:2609.18578v1 Announce Type: new Abstract: Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transf...
arXiv:2608.24281v1 Announce Type: new Abstract: Reducing annotation requirements remains a key challenge in developing robust medical object detectors. To address this, Vision-Language (VL) object de...
Exemplar is a few‑shot segmentation method that combines a frozen DINOv3 backbone with a fixed bank of classical native‑resolution filter responses in a single lightweight head. Trained only from support masks, it achieves a mean foreground intersection‑over‑union of 0.782 across eleven biomedical imaging datasets, outperforming either component alone and surpassing five other few‑shot methods in 54 of 55 comparisons. With a single annotated mask, Exemplar reaches 0.703, higher than a from‑scratch nnU‑Net trained on the same mask, and while nnU‑Net eventually overtakes it with eight masks, it requires 16–77× longer to fit.