arXiv:2608.21244v2 Announce Type: replace
Abstract: Anomaly detection aims to identify observations that deviate from normal patterns. Recent work uses pretrained vision-language models (VLMs) for tr...
By Inpyo Song, Jangwon Lee
The paper proposes a lightweight federated multiple‑instance learning (MIL) framework that trains only a compact MIL scorer across distributed clients while using a frozen vision‑language model (VLM) to verify high‑scoring video segments post‑hoc. Two VLM feedback interfaces are explored: a parsed text‑generation interface and a logit‑based interface that derives a continuous anomaly score from next‑token Yes/No probabilities. Experiments on UCF‑Crime with InternVL3.5‑2B and Qwen3‑VL‑2B‑Instruct show that the logit interface consistently improves frame‑level AUC and AP over the MIL baseline without requiring temporal post‑processing, whereas the text‑generation interface is more sensitive to prompts, parsers, and model choice.
By S\'ebastien Thuau, Amira Gran, Siba Haidar, Rachid Chelouah
The paper investigates how vision‑language models (VLMs) can be used for training‑free video anomaly detection (VAD) by answering questions about video segments. It shows that the way a VLM’s output is converted into an anomaly score—either by taking only the most likely answer (generated readout) or by using the full probability distribution (probability readout)—has a significant impact on ranking performance. Across four large VLMs, the probability readout consistently outperforms the generated readout, achieving 5–13 point gains in AUROC or AP, because the generated readout compresses the rank order into only a few distinct scores.
By Inpyo Song, Jangwon Lee
arXiv:2608. 11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.
By Shibo Gao, Peipei Yang, Xu-Yao Zhang, Linlin Huang
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data.
The paper introduces a new task called Quality Anomaly Perception for UGC Image Enhancement (UEAP) and presents the first benchmark dataset, UEAP-4k, featuring fine‑grained annotations of anomaly categories, locations, and severity levels in real‑world user‑generated content. It proposes the Difference‑Fusion Anomaly Perception Method (DFAP‑UGC), which fuses explicit differences between enhanced images and their references using dense spatial querying, regional verification, and quality‑aware ranking to robustly identify localized anomalies. A Locality‑Aware Dynamic Task Prioritization (LADTP) training strategy is also introduced to enable efficient end‑to‑end learning without multi‑stage overhead, and experiments demonstrate that DFAP‑UGC outperforms adapted classical baselines.
By Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi, Tingting Jiang