arXiv:2608.11985v3 Announce Type: replace
Abstract: Frame-level area under the ROC curve (AUC) is the dominant evaluation metric for weakly supervised video anomaly detection (WSVAD). Its standard fo...
By Sara Abdulaziz, Egor Bondarev
The paper investigates how weakly supervised video anomaly detectors, trained with only video‑level labels, are evaluated using frame‑level metrics such as Micro‑AUROC and AP. It shows that these metrics largely measure a detector’s ability to separate different videos rather than correctly ordering anomalous moments within a single video, a phenomenon termed temporal dilution. Experiments demonstrate that a detector can achieve high pooled scores even when it assigns the same score to every frame in a video, indicating that current evaluation practices may overstate temporal localization performance.
By Inpyo Song, Jangwon Lee
arXiv:2608. 19987v1 Announce Type: new Abstract: Skeleton-based Video Anomaly Detection (VAD) offers a robust, privacy-preserving solution for identifying abnormal behaviors.
By Jakub Micorek, Mateusz Kozi\'nski, Horst Possegger
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
arXiv:2607. 03558v1 Announce Type: cross Abstract: Continuous video anomaly detection is dominated by reactive Multiple Instance Learning (MIL) that collapses spatiotemporal features into scalar scores.
By Abu Anas Ibn Samad
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
arXiv:2609.09736v1 Announce Type: new
Abstract: Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a partic...
By Shiwen Zhao, Qi Zhang, Sezer Karaoglu, Theo Gevers, Martin R. Oswald
arXiv:2605.31192v2 Announce Type: replace
Abstract: Generalizable deepfake detection requires complementary forensic and semantic visual evidence. Specialist encoders capture subtle manipulation trac...
By Benedikt Hopf, Zongwei Wu, Radu Timofte
The paper introduces RIFT, a forensic framework for detecting AI-generated videos by exploiting a cross‑scale coupling mismatch between macro‑level temporal dynamics and micro‑level pixel residuals. RIFT comprises a macro stream that models expected temporal evolution, a micro stream that probes residual patterns, and a coupling divergence module that quantifies their conditional dependency. Experiments on VidProM and GenVidBench show near‑perfect F1‑scores and robust performance across different encoders.
By Siyu Li, Jin Yang, Weiheng Liang
arXiv:2605.21957v2 Announce Type: replace
Abstract: Video anomaly detection is critical for public safety and security, yet remains highly challenging despite extensive research due to large variatio...
By Inpyo Song, Jangwon Lee
arXiv:2607. 15400v1 Announce Type: cross Abstract: Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain.
By Tasmiah Haque, Jacob Kosinski, Sumit Mohan, Srinjoy Das, Mohammad Abdullah Al-Mamun
CoRE is a weakly supervised framework that learns fine-grained temporal and entity support for perceived risk in driving videos using only coarse video-level judgments. It first trains a video-level predictor, freezes it, and then uses structured interventions over candidate temporal regions or entity tracks to generate graded prediction-effect targets. These targets train a student model that can predict temporal and entity support directly from the original video, enabling fine-grained evidence localization without requiring detailed annotations.
By Kaiser Hamid, Can Cui, Nade Liang