arXiv AI

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

arXiv Computer Vision
Sep 23

Test-time Reinforcement Learning for Anomalous Video Understanding

The paper introduces a test‑time reinforcement learning framework for anomalous video understanding, addressing challenges such as unreliable pseudo‑labels, inadequate reward design, and collapsed group‑relative advantages. It proposes dual‑query consistency filtering, an entropy‑aware consensus reward, and a virtual negative anchor mechanism to improve sample reliability, reward quality, and policy‑gradient signals. Experiments on VAU‑Bench demonstrate significant performance gains, especially on the ECVA subset where accuracy rises from 75.81% to 90.00%.

By Huining Li, Yuxiang Duan, Jiyang Tan, Qian Li, MingCai Chen, Jian Zhang, Xingdong Sheng, Yuntao Du
arXiv AI
Sep 7

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

The paper introduces an adaptive temporal modeling framework for weakly supervised video anomaly detection that addresses the limitations of rigid Multiple Instance Learning approaches. It presents a Temporal Refinement Module using dynamic positional encoding and a learnable class token to capture long‑range dependencies, and an Event Segmentation Module that identifies event boundaries via temporal discontinuity analysis to produce discriminative event‑level representations. An adaptive similarity‑based fusion strategy replaces fixed top‑k heuristics, dynamically integrating snippet‑level and event‑level anomaly scores into video‑level predictions, and the method outperforms state‑of‑the‑art baselines on two benchmarks.

By Changyi Li, Yu Xiao
arXiv Computer Vision
Sep 23

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

arXiv:2505.01583v2 Announce Type: replace Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...

By Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou, Vivian Wang, Huayu Wang, Hsiang-Wei Huang, Wenhao Chai, Hou-I Liu, Kuang-Ming Chen, Cheng-Yen Yang, Yi-Ling Chen, Vibhav Vineet, Qin Cai, Jenq-Neng Hwang
arXiv AI
Aug 25

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

The paper introduces TimeCatch, a benchmark that evaluates temporal consistency in vision‑language models (VLMs) by treating temporal grounding as an anomaly detection problem. Temporal anomalies are created by swapping consecutive frames, while frame‑level anomalies involve replacing a frame with Gaussian noise. Across synthetic and real‑world datasets, VLMs reliably detect and localize frame‑level anomalies but perform near chance on temporal anomaly detection, whereas humans excel at both tasks.

By Marek Hradil, Danae S\'anchez Villegas
arXiv Computer Vision
1d ago

PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation

PARSEE-VAD is a training‑free online video anomaly detection framework that separates semantic evidence acquisition from score‑state evolution. It uses Proposition‑Aware Reasoning to extract structured propositional evidence from the current causal window and selectively activates more specific queries, while Streaming Evidence Escalation maps this evidence into a compact score‑domain event state and propagates only the bounded state to maintain temporal continuity. Experiments on four benchmarks show strong performance with reduced specialist computation and sparse score‑state propagation, supporting a current‑first principle for streaming multimodal inference.

By Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao
arXiv AI
Jun 12

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding

arXiv:2506. 01274v2 Announce Type: replace-cross Abstract: Recent progress in Large Multi-modal Models (LMMs) has enabled effective vision-language reasoning, yet the ability to video understanding remains constrained by suboptimal frame selection strategies, albeit with the rapid development of video-specialized LMMs.

By Hosu Lee, Junho Kim, Hyunjun Kim, Yong Man Ro
arXiv Computer Vision
Aug 21

Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding

arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.

By Yanxiang Huang, Guohua Gao, Zhaoyang Wei