arXiv Computer Vision
Aug 24

A VLM Answer Is Not an Anomaly Score: Rank Compression in Training-Free Video Anomaly Detection

The paper investigates how vision‑language models (VLMs) can be used for training‑free video anomaly detection (VAD) by answering questions about video segments. It shows that the way a VLM’s output is converted into an anomaly score—either by taking only the most likely answer (generated readout) or by using the full probability distribution (probability readout)—has a significant impact on ranking performance. Across four large VLMs, the probability readout consistently outperforms the generated readout, achieving 5–13 point gains in AUROC or AP, because the generated readout compresses the rank order into only a few distinct scores.

By Inpyo Song, Jangwon Lee
arXiv AI
4d ago

Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection

The paper proposes a lightweight federated multiple‑instance learning (MIL) framework that trains only a compact MIL scorer across distributed clients while using a frozen vision‑language model (VLM) to verify high‑scoring video segments post‑hoc. Two VLM feedback interfaces are explored: a parsed text‑generation interface and a logit‑based interface that derives a continuous anomaly score from next‑token Yes/No probabilities. Experiments on UCF‑Crime with InternVL3.5‑2B and Qwen3‑VL‑2B‑Instruct show that the logit interface consistently improves frame‑level AUC and AP over the MIL baseline without requiring temporal post‑processing, whereas the text‑generation interface is more sensitive to prompts, parsers, and model choice.

By S\'ebastien Thuau, Amira Gran, Siba Haidar, Rachid Chelouah