arXiv:2608. 13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users.
By Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
arXiv:2608. 11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.
By Shibo Gao, Peipei Yang, Xu-Yao Zhang, Linlin Huang
TAU-Agent is a retrieval‑augmented framework designed for traffic anomaly understanding in transportation videos. It uses a central retrieval agent to coordinate a Video Captioning Tool and an Open‑Vocabulary Tracking Tool, gathering captions, temporal intervals, and object trajectories relevant to a query. The collected evidence, along with sampled frames and the query, is fed into a fine‑tuned vision‑language model that reasons and generates an answer. TAU-Agent was evaluated on the AI City Challenge 2026 benchmarks, achieving notable scores across multiple tracks and ranking second, twelfth, and fifth respectively.
By Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agen...
arXiv:2610.01754v1 Announce Type: cross
Abstract: Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and c...
By Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais
AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.
By Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee
arXiv:2603. 10652v3 Announce Type: replace-cross Abstract: In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion.
By Yangfan He, Changgyu Boo, Jaehong Yoon
arXiv:2608.28675v1 Announce Type: cross
Abstract: Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. I...
By Mingwen Zhang, Jisheng Dang, Minqiang Yang, Bimei Wang, Bin Hu, Tat-Seng Chua
arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.
By Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali
The paper introduces a test‑time reinforcement learning framework for anomalous video understanding, addressing challenges such as unreliable pseudo‑labels, inadequate reward design, and collapsed group‑relative advantages. It proposes dual‑query consistency filtering, an entropy‑aware consensus reward, and a virtual negative anchor mechanism to improve sample reliability, reward quality, and policy‑gradient signals. Experiments on VAU‑Bench demonstrate significant performance gains, especially on the ECVA subset where accuracy rises from 75.81% to 90.00%.
By Huining Li, Yuxiang Duan, Jiyang Tan, Qian Li, MingCai Chen, Jian Zhang, Xingdong Sheng, Yuntao Du
arXiv:2605. 21917v2 Announce Type: replace-cross Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support.
By Han Zhang, Wanting Jiang, Tomasz Kornuta, Tian Zheng, Vidya Murali
The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.
By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang