arXiv Computer Vision

Beyond Normal References: Discriminative Few-Shot Anomaly Detection

This paper introduces IDEAL, a framework for few-shot anomaly detection that uses both normal and anomalous reference examples. IDEAL suppresses irrelevant normal variations and encodes intrinsic deviation vectors that capture discriminative anomaly directions. Experiments on eight real-world datasets show that IDEAL generalizes to unseen anomalies and outperforms existing methods.

arXiv Computer Vision
Sep 4

PL-SCEA: Reconfiguring Pretrained Attention for Few-Shot Industrial Anomaly Detection

The paper introduces PL‑SCEA, a method that reconfigures the attention mechanism of frozen Vision Foundation Models to better detect and localize anomalies in industrial images with few training examples. PL‑SCEA preserves the semantic context of pretrained query‑key attention while adding token‑adaptive self‑correlations over contextualized value features, then applies positive‑correlation filtering and power‑law reweighting to highlight task‑relevant relationships. The resulting features are fed into a lightweight variational autoencoder to produce reconstruction‑based anomaly scores, achieving competitive image‑level detection and strong pixel‑level localization on MVTec AD and VisA datasets.

By Xiaoyu Yang, Qixing Wu, Huixian Zhao, Changlong Jin
arXiv Computer Vision
Sep 3

DPA: Decoupling Product-Agnostic Anomaly Representations for Zero-shot Anomaly Generation

The paper introduces DPA, a diffusion-based framework that decouples product-agnostic anomaly representations to enable zero-shot anomaly generation. By reusing real anomalies from existing source products and filtering them for plausibility, DPA learns product-irrelevant anomaly embeddings that can be transferred across products. An adaptive mask-guided pipeline and a training-free labeling module further refine the realism and localization of generated anomalies, leading to improved performance on MVTec-AD, VisA, and a new anomaly-transfer benchmark.

By Hang Yao, Yansheng Fu, Ming Liu, Zifei Yan, Yanli Ji, Hongzhi Zhang, Wangmeng Zuo
arXiv Computer Vision
Sep 4

Neural-Collapse-guided Task-Free Continual Anomaly Detection

The paper introduces NC‑TFAD, a task‑free continual anomaly detection framework that leverages neural‑collapse geometry to learn from non‑stationary data streams without task boundaries. It freezes a pretrained backbone, aligns streaming features to a simplex Equiangular Tight Frame prototype space, and uses synthetic anomaly anchors, inter‑ and intra‑class regularization, and a Focal Neural Collapse Contrastive loss to stabilize representations and enhance normal‑anomaly separability. A normal‑patch‑prototype‑guided localization branch generates calibrated anomaly heatmaps, and extensive experiments on MVTec AD and VisA demonstrate that NC‑TFAD outperforms existing task‑free continual learning and unified anomaly detection baselines in both image‑level detection and pixel‑level localization.

By Xiaotong Kong, Chaoyang Song, Ziai Zhou, Jinxia Zhang, Kanjian Zhang, Haikun Wei
arXiv Machine Learning
Sep 23

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

The paper investigates whether the performance of anomaly detection systems can be predicted without labeled anomalies. For kNN-based detectors, it derives a lower bound on AUC that links detection performance to the separation and variance of inlier and outlier scores, and uses this to analyze how density variation, intrinsic dimensionality, and domain mismatch affect score variability. The authors introduce pseudo‑anomaly probes that provide a reference for estimating relative score separation, and demonstrate through experiments on DCASE benchmarks that these probes enable anomaly‑free model selection to outperform conventional development‑set selection, especially under domain shift.

By Kevin Wilkinghoff, Zheng-Hua Tan
Hugging Face Trending Papers
Aug 5

VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-based VAD, which focuses on motion dynamics rather than raw video data.