The paper introduces Anomaly‑LR, a defect‑grounded latent reasoning framework for industrial anomaly detection that builds a global understanding of an image and then refines anomaly‑relevant representations directly in visual latent space. It also presents IAD‑LR‑22K, a new instruction dataset with 22,228 image‑question pairs and detailed annotations. Experiments demonstrate that Anomaly‑LR outperforms comparable‑scale methods on multiple IAD benchmarks without needing external references or tools.
By Jaron Yeh, Yen-Wei Chang, Jiang Liu, Shao-Yuan Lo
arXiv:2606. 01992v1 Announce Type: cross Abstract: Industrial anomaly detection has historically been a unimodal task.
By Stefano Samele, Eugenio Lomurno, Teodora Jovanovic, Sanjay Shivakumar Manohar, Alberto Crivellaro, Matteo Matteucci
arXiv:2608.29783v1 Announce Type: new
Abstract: Industrial anomaly detection is a critical component of modern manufacturing. Most traditional unsupervised methods rely on modelling normal feature di...
By Weifei Chen, Honghao Zhang, Zhiyuan You, Xinyi Le
arXiv:2608. 11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.
By Shibo Gao, Peipei Yang, Xu-Yao Zhang, Linlin Huang
Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) can improve recognition accuracy by comparing query images with references at inference time, but these benefits rely on additional retrieval and processing.
Probe‑VAD introduces an ordinal binary‑probing framework that leverages frozen vision‑language models for training‑free video anomaly detection. By querying ten ordered severity thresholds and extracting YES/NO continuation likelihoods, it builds a cumulative severity profile that is converted into a continuous anomaly score with isotonic projection for ordinal consistency. Experiments on public benchmarks show that this simple interface yields superior performance at low computational cost, avoiding the limitations of caption‑based compression or restricted numerical scoring.
By Jiawei Gu, Qilin Zhao, Tengkuo Guo, Zhiming Zhong, Shuangqing Zhang, Fan Lyu, Fang Zhao, Guo-Sen Xie, Caifeng Shan
arXiv:2609.13228v1 Announce Type: new
Abstract: Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasonin...
By Marko Jojic, Zhaonan Li, Ben Zhou
arXiv:2608.23723v1 Announce Type: new
Abstract: Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual f...
By Wenyang Liu, Tianyi Liu, Dongshuo Zhang, Kejun Wu, Adams Wai-Kin Kong
The paper introduces TED (Text-Axis Evidence Decomposition), a post‑hoc scoring method that improves anomaly localization in CLIP‑based detectors without altering the backbone or prompts. TED evaluates whether ambiguous responses are better supported by defect patches or normal patches, thereby distinguishing true defects from visually complex normal regions. Experiments show that TED significantly enhances pixel‑level localization across frozen VLM backbones and adapted hosts, especially under hard‑false‑positive competition.
By JinYoung Kim, Geonho Kim, GiJeong Park, Geonu Lee, YoungJoon Yoo
The paper introduces DPA, a diffusion-based framework that decouples product-agnostic anomaly representations to enable zero-shot anomaly generation. By reusing real anomalies from existing source products and filtering them for plausibility, DPA learns product-irrelevant anomaly embeddings that can be transferred across products. An adaptive mask-guided pipeline and a training-free labeling module further refine the realism and localization of generated anomalies, leading to improved performance on MVTec-AD, VisA, and a new anomaly-transfer benchmark.
By Hang Yao, Yansheng Fu, Ming Liu, Zifei Yan, Yanli Ji, Hongzhi Zhang, Wangmeng Zuo
arXiv:2610.01754v1 Announce Type: cross
Abstract: Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and c...
By Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais
The paper introduces a new task called Quality Anomaly Perception for UGC Image Enhancement (UEAP) and presents the first benchmark dataset, UEAP-4k, featuring fine‑grained annotations of anomaly categories, locations, and severity levels in real‑world user‑generated content. It proposes the Difference‑Fusion Anomaly Perception Method (DFAP‑UGC), which fuses explicit differences between enhanced images and their references using dense spatial querying, regional verification, and quality‑aware ranking to robustly identify localized anomalies. A Locality‑Aware Dynamic Task Prioritization (LADTP) training strategy is also introduced to enable efficient end‑to‑end learning without multi‑stage overhead, and experiments demonstrate that DFAP‑UGC outperforms adapted classical baselines.
By Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi, Tingting Jiang