arXiv AI

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

arXiv:2607. 18850v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect reasoning.

arXiv Computer Vision
Sep 25

Industrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space

The paper introduces Anomaly‑LR, a defect‑grounded latent reasoning framework for industrial anomaly detection that builds a global understanding of an image and then refines anomaly‑relevant representations directly in visual latent space. It also presents IAD‑LR‑22K, a new instruction dataset with 22,228 image‑question pairs and detailed annotations. Experiments demonstrate that Anomaly‑LR outperforms comparable‑scale methods on multiple IAD benchmarks without needing external references or tools.

By Jaron Yeh, Yen-Wei Chang, Jiang Liu, Shao-Yuan Lo
arXiv Computer Vision
Sep 16

Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

Probe‑VAD introduces an ordinal binary‑probing framework that leverages frozen vision‑language models for training‑free video anomaly detection. By querying ten ordered severity thresholds and extracting YES/NO continuation likelihoods, it builds a cumulative severity profile that is converted into a continuous anomaly score with isotonic projection for ordinal consistency. Experiments on public benchmarks show that this simple interface yields superior performance at low computational cost, avoiding the limitations of caption‑based compression or restricted numerical scoring.

By Jiawei Gu, Qilin Zhao, Tengkuo Guo, Zhiming Zhong, Shuangqing Zhang, Fan Lyu, Fang Zhao, Guo-Sen Xie, Caifeng Shan
arXiv AI
3d ago

TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization

The paper introduces TED (Text-Axis Evidence Decomposition), a post‑hoc scoring method that improves anomaly localization in CLIP‑based detectors without altering the backbone or prompts. TED evaluates whether ambiguous responses are better supported by defect patches or normal patches, thereby distinguishing true defects from visually complex normal regions. Experiments show that TED significantly enhances pixel‑level localization across frozen VLM backbones and adapted hosts, especially under hard‑false‑positive competition.

By JinYoung Kim, Geonho Kim, GiJeong Park, Geonu Lee, YoungJoon Yoo
arXiv Computer Vision
Sep 3

DPA: Decoupling Product-Agnostic Anomaly Representations for Zero-shot Anomaly Generation

The paper introduces DPA, a diffusion-based framework that decouples product-agnostic anomaly representations to enable zero-shot anomaly generation. By reusing real anomalies from existing source products and filtering them for plausibility, DPA learns product-irrelevant anomaly embeddings that can be transferred across products. An adaptive mask-guided pipeline and a training-free labeling module further refine the realism and localization of generated anomalies, leading to improved performance on MVTec-AD, VisA, and a new anomaly-transfer benchmark.

By Hang Yao, Yansheng Fu, Ming Liu, Zifei Yan, Yanli Ji, Hongzhi Zhang, Wangmeng Zuo
arXiv AI
Sep 3

Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework

The paper introduces a new task called Quality Anomaly Perception for UGC Image Enhancement (UEAP) and presents the first benchmark dataset, UEAP-4k, featuring fine‑grained annotations of anomaly categories, locations, and severity levels in real‑world user‑generated content. It proposes the Difference‑Fusion Anomaly Perception Method (DFAP‑UGC), which fuses explicit differences between enhanced images and their references using dense spatial querying, regional verification, and quality‑aware ranking to robustly identify localized anomalies. A Locality‑Aware Dynamic Task Prioritization (LADTP) training strategy is also introduced to enable efficient end‑to‑end learning without multi‑stage overhead, and experiments demonstrate that DFAP‑UGC outperforms adapted classical baselines.

By Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi, Tingting Jiang