arXiv AI

MemoVAD: Resource-Efficient Video Anomaly Detection via Dynamic Semantic Memory in Edge Computing Scenarios

arXiv:2606. 07669v1 Announce Type: cross Abstract: Deploying Video Anomaly Detection (VAD) in real-world surveillance faces a fundamental tension between the demand for high-level semantics to ensure effectiveness and the limited computational resources of edge devices.

arXiv AI
6d ago

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Cog-VADU is a training‑free framework that transforms video anomaly detection into a sequential cognitive reasoning task. It uses Chain‑of‑Anomaly Detection Thought Prompting (CoADTP) to create a recurrent reasoning chain across video segments, preserving temporal memory and distinguishing complex anomalies from high‑motion normal activities. A cross‑modal re‑ranking stage aligns textual rationales with visual embeddings to enforce semantic consistency and temporal coherence, yielding competitive zero‑shot performance on multiple VAD benchmarks.

By Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais
arXiv AI
Sep 7

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

The paper introduces an adaptive temporal modeling framework for weakly supervised video anomaly detection that addresses the limitations of rigid Multiple Instance Learning approaches. It presents a Temporal Refinement Module using dynamic positional encoding and a learnable class token to capture long‑range dependencies, and an Event Segmentation Module that identifies event boundaries via temporal discontinuity analysis to produce discriminative event‑level representations. An adaptive similarity‑based fusion strategy replaces fixed top‑k heuristics, dynamically integrating snippet‑level and event‑level anomaly scores into video‑level predictions, and the method outperforms state‑of‑the‑art baselines on two benchmarks.

By Changyi Li, Yu Xiao
arXiv Computer Vision
Sep 30

PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation

PARSEE-VAD is a training‑free online video anomaly detection framework that separates semantic evidence acquisition from score‑state evolution. It uses Proposition‑Aware Reasoning to extract structured propositional evidence from the current causal window and selectively activates more specific queries, while Streaming Evidence Escalation maps this evidence into a compact score‑domain event state and propagates only the bounded state to maintain temporal continuity. Experiments on four benchmarks show strong performance with reduced specialist computation and sparse score‑state propagation, supporting a current‑first principle for streaming multimodal inference.

By Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao
arXiv AI
Jul 13

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

arXiv:2607. 09114v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos.

By Peipei Zhu, Yueqing Niu, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li
arXiv AI
Sep 28

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang
arXiv Computer Vision
Aug 31

Training-Free Temporal Abstraction for General Video Understanding

The paper introduces STITCH, a training‑free method that partitions videos into semantically meaningful temporal chunks using a frozen video‑text backbone. By detecting changes in the embedding sequence of short video windows, STITCH produces reusable temporal abstractions that can be applied to multiple tasks such as event boundary detection, language‑based moment retrieval, and frame selection for vision‑language models. Experiments show that STITCH performs competitively with specialized methods while requiring no task‑specific training, especially when processing is limited to a few frames or tokens.

By Etienne Casanova, Sevan Brodjian, Pietro Perona