arXiv Computer Vision By Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao

PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation

Read the original on arXiv Computer Vision →

PARSEE-VAD is a training‑free online video anomaly detection framework that separates semantic evidence acquisition from score‑state evolution. It uses Proposition‑Aware Reasoning to extract structured propositional evidence from the current causal window and selectively activates more specific queries, while Streaming Evidence Escalation maps this evidence into a compact score‑domain event state and propagates only the bounded state to maintain temporal continuity. Experiments on four benchmarks show strong performance with reduced specialist computation and sparse score‑state propagation, supporting a current‑first principle for streaming multimodal inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 20

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

Event-Causal RAG (EC‑RAG) is a lightweight retrieval‑augmented framework designed for reasoning over ultra‑long and streaming videos. It segments video streams into semantically complete events using a dual visual‑audio sentinel mechanism, representing each event as a State‑Event‑State (SES) structure that captures pre‑event, event, and post‑event states. During question answering, bidirectional graph retrieval accesses relevant predecessor and successor events from a dual vector‑graph memory, and answers are generated using both this structured memory and the corresponding video evidence. The authors also introduce ECV‑1H, an hour‑scale long‑video QA benchmark with over 150 hours of untrimmed video and 1,251 human‑annotated QA pairs, where EC‑RAG achieves significant accuracy gains across multiple video foundation models while maintaining efficient streaming memory usage on a single RTX 5090 GPU.

By Peizheng Yan, Yu Zhao, Liang Xie, Juntong Qi, Mingming Wang, Erwei Yin
arXiv AI
3d ago

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang
arXiv AI
1d ago

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.

By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
arXiv Computation and Language
6d ago

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND is a new benchmark that tests multimodal large language models’ ability to track objects, their states, and relationships over time in videos. It evaluates intermediate reasoning by breaking queries into sub‑questions and uses Faithful Accuracy to ensure all parts of an answer are correct. The authors also propose an object‑centric framework that builds structured trajectories and shows reduced hallucinations and better temporal consistency compared to existing models.

By Thong Nguyen, Tri Cao, Khoi Le, Cong-Duy Nguyen, Quynh Vo, See-Kiong Ng, Bryan Hooi Kuen-Yew