arXiv AI

Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

arXiv:2606. 09670v1 Announce Type: cross Abstract: Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec.

Hugging Face Trending Papers
Jun 8

Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these methods face challenges when foundational assumptions - such as consistent object scale, viewpoint, background, illumination, and centered placement - are violated.

arXiv Computer Vision
Aug 27

See More, Detect Less? Taming Information Leakage in Multi-View Anomaly Detection

The paper identifies a problem in multi‑view anomaly detection called cross‑view information leakage, where fusing multiple inspection views can cause normal features to mask anomalies during reconstruction. To address this, the authors propose GLAD, a framework that uses a Global‑Local Attention Driven approach, combining vision foundation model features with two fusion modules: Multi‑view Merging Attention for local, weighted fusion and Object‑Guided Attention for global context aggregation. Experiments on Real‑IAD and MANTA‑Tiny demonstrate that GLAD outperforms existing methods across various metrics, underscoring the importance of restricting information flow to preserve the reconstruction gap.

By Shang-Fu Chen, Kuan-Chuan Peng, Jhih-Ciang Wu, Wen-Huang Cheng, Kai-Lung Hua
arXiv Machine Learning
Sep 10

Contrastive Knowledge Distillation for Anomaly Detection in Multi-Illumination/Focus Display Images

The paper introduces a contrastive learning approach for anomaly detection in multi-illumination and multi-focus display images. It builds on Multiresolution Knowledge Distillation (MKD) and proposes Multiresolution Contrastive Distillation (MCD), which eliminates the need for explicit positive/negative pairs by adjusting distances between teacher and student features. A blending module aggregates multi-channel data into a three‑channel input, and the method achieves superior AUROC and accuracy on the MMdAD dataset compared to state‑of‑the‑art baselines.

By Jihyun Lee, Hangil Park, Yongmin Seo, Taewon Min, Joodong Yun, Jaewon Kim, Tae-Kyun Kim
arXiv Computer Vision
Sep 7

Training-Free Logical and Structural Anomaly Detection via Calibrated Fusion

The paper introduces a training‑free anomaly detector that simultaneously handles structural and logical defects by calibrating heterogeneous anomaly cues with statistics from normal images. This calibration aligns frozen representations, allowing their fusion without extra training or part‑level supervision. The resulting method achieves state‑of‑the‑art AUROC scores on MVTec‑LOCO and remains competitive on MVTec‑AD.

By Changyi Li, Miao Yu, Kai Dong, Yu Xiao
arXiv Computer Vision
Sep 16

Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

Probe‑VAD introduces an ordinal binary‑probing framework that leverages frozen vision‑language models for training‑free video anomaly detection. By querying ten ordered severity thresholds and extracting YES/NO continuation likelihoods, it builds a cumulative severity profile that is converted into a continuous anomaly score with isotonic projection for ordinal consistency. Experiments on public benchmarks show that this simple interface yields superior performance at low computational cost, avoiding the limitations of caption‑based compression or restricted numerical scoring.

By Jiawei Gu, Qilin Zhao, Tengkuo Guo, Zhiming Zhong, Shuangqing Zhang, Fan Lyu, Fang Zhao, Guo-Sen Xie, Caifeng Shan