Crane is a CLIP‑based framework for zero‑shot anomaly detection that enhances dense localization by adapting the vision encoder with a correlation‑based attention module and conditioning learnable prompts on global image context. It further fuses anomaly‑relevant patch features into the global representation for more sensitive image‑level detection, and a variant called Crane+ leverages DINOv2 spatial correlations for stronger pixel‑level performance. Across seven industrial benchmarks, Crane raises mean image‑level AP by 4.5% and Crane+ boosts mean pixel‑level AUPRO by 9.0%.
By Alireza Salehi, Mohammadreza Salehi, Reshad Hosseini, Cees G. M. Snoek, Makoto Yamada, Mohammad Sabokrou
arXiv:2608.24281v1 Announce Type: new
Abstract: Reducing annotation requirements remains a key challenge in developing robust medical object detectors. To address this, Vision-Language (VL) object de...
By Sheethal Bhat, Bogdan Georgescu, Awais Mansoor, Mathias Zinnen, Pranjal Sahu, Florin C. Ghesu, Sasa Grbic, Andreas Maier
The paper introduces Spatial‑FAD, a few‑shot medical anomaly detection framework that fuses Vision‑Language Model (CLIP) semantics with spatial priors from Vision Foundation Models (DINO). A VFM‑enhanced adapter injects structural affinity into CLIP features, while a sliding‑window aggregation produces high‑resolution embeddings for finer lesion localization. Prototype‑enhanced support memory further improves efficiency and performance, yielding significant gains on Liver CT, Retinal OCT, and Brain MRI datasets, notably an 11.4% Dice improvement in 4‑shot scenarios.
By Juzheng Miao, Yuchen Yuan, Cheng Chen, Pheng-Ann Heng
Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on independent local patch features, leaving the global contextual information encoded by Vision Transformers (ViTs) underexploited.
arXiv:2607. 23924v1 Announce Type: cross Abstract: Vision foundation models have enabled strong training-free anomaly detection (AD).
By Jyun-Ze Tang, Po-Han Huang, Ming-Ching Chang, Chih-Fan Hsu, Jeng-Lin Li
arXiv:2609.16785v1 Announce Type: new
Abstract: Zero-shot anomaly detection aims to localize anomalies without target-domain samples. Existing CLIP-based methods suffer from coarse anomaly maps and l...
By Xuezhi Xiang, Guanghao Wu, Heqi Xiang, Jiayao Liu, Xiaoheng Li, Yiming Chen, Shanjun Zhang
Zero-shot anomaly detection aims to localize anomalies without target-domain samples. Existing CLIP-based methods suffer from coarse anomaly maps and limited semantic prompts. We propose PSMP-CLIP, in...
BiCLIP is a bidirectional multimodal framework that enhances medical image segmentation by allowing visual features to iteratively refine textual representations, improving semantic alignment. It incorporates an augmentation consistency objective to stabilize learning against perturbed inputs. Experiments on QaTa-COV19 and MosMedData+ show that BiCLIP outperforms state‑of‑the‑art image‑only and multimodal baselines, achieving strong performance even with only 1% labeled data and resisting common clinical artifacts such as motion blur and low‑dose CT noise.
By Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mourad Oussalah
Unified visual anomaly detection seeks to train a single detector that can be deployed across categories, domains, and application scenarios. In the few-shot transfer regime, the key challenge is to estimate an episode-specific boundary for an unseen target category from a small support set.
arXiv:2608.23723v1 Announce Type: new
Abstract: Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual f...
By Wenyang Liu, Tianyi Liu, Dongshuo Zhang, Kejun Wu, Adams Wai-Kin Kong
The paper introduces PL‑SCEA, a method that reconfigures the attention mechanism of frozen Vision Foundation Models to better detect and localize anomalies in industrial images with few training examples. PL‑SCEA preserves the semantic context of pretrained query‑key attention while adding token‑adaptive self‑correlations over contextualized value features, then applies positive‑correlation filtering and power‑law reweighting to highlight task‑relevant relationships. The resulting features are fed into a lightweight variational autoencoder to produce reconstruction‑based anomaly scores, achieving competitive image‑level detection and strong pixel‑level localization on MVTec AD and VisA datasets.
By Xiaoyu Yang, Qixing Wu, Huixian Zhao, Changlong Jin
Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language fe...