Hugging Face Trending Papers

VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-based VAD, which focuses on motion dynamics rather than raw video data.

arXiv AI
Aug 6

VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

arXiv:2608. 05069v1 Announce Type: cross Abstract: Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance.

By Narges Rashvand, Ghazal Alinezhad Noghre, Shanle Yao, Gabriel Maldonado, Hamed Tabkhi
arXiv AI
Jul 13

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

arXiv:2607. 09114v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos.

By Peipei Zhu, Yueqing Niu, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li
arXiv Computer Vision
Sep 23

Latent Dataset Distillation for Human Motion Prediction

The paper introduces a latent dataset distillation framework for human motion prediction, addressing the limitations of traditional gradient matching by incorporating a learned motion prior. Motions are compressed using a residual‑quantized variational autoencoder, and distillation updates only a latent bank while keeping the decoder frozen, ensuring synthetic motions remain plausible. Experiments on Human3.6M, CMU, and 3DPW datasets demonstrate that this method outperforms direct gradient matching in most settings and yields more realistic synthetic motions.

By Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama
arXiv Computer Vision
Sep 16

Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection

Probe‑VAD introduces an ordinal binary‑probing framework that leverages frozen vision‑language models for training‑free video anomaly detection. By querying ten ordered severity thresholds and extracting YES/NO continuation likelihoods, it builds a cumulative severity profile that is converted into a continuous anomaly score with isotonic projection for ordinal consistency. Experiments on public benchmarks show that this simple interface yields superior performance at low computational cost, avoiding the limitations of caption‑based compression or restricted numerical scoring.

By Jiawei Gu, Qilin Zhao, Tengkuo Guo, Zhiming Zhong, Shuangqing Zhang, Fan Lyu, Fang Zhao, Guo-Sen Xie, Caifeng Shan
arXiv AI
Sep 7

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

The paper introduces an adaptive temporal modeling framework for weakly supervised video anomaly detection that addresses the limitations of rigid Multiple Instance Learning approaches. It presents a Temporal Refinement Module using dynamic positional encoding and a learnable class token to capture long‑range dependencies, and an Event Segmentation Module that identifies event boundaries via temporal discontinuity analysis to produce discriminative event‑level representations. An adaptive similarity‑based fusion strategy replaces fixed top‑k heuristics, dynamically integrating snippet‑level and event‑level anomaly scores into video‑level predictions, and the method outperforms state‑of‑the‑art baselines on two benchmarks.

By Changyi Li, Yu Xiao