arXiv:2608. 13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users.
By Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
arXiv:2608. 11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.
By Shibo Gao, Peipei Yang, Xu-Yao Zhang, Linlin Huang
TAU-Agent is a retrieval‑augmented framework designed for traffic anomaly understanding in transportation videos. It uses a central retrieval agent to coordinate a Video Captioning Tool and an Open‑Vocabulary Tracking Tool, gathering captions, temporal intervals, and object trajectories relevant to a query. The collected evidence, along with sampled frames and the query, is fed into a fine‑tuned vision‑language model that reasons and generates an answer. TAU-Agent was evaluated on the AI City Challenge 2026 benchmarks, achieving notable scores across multiple tracks and ranking second, twelfth, and fifth respectively.
By Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agen...
arXiv:2610.01754v1 Announce Type: cross
Abstract: Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and c...
By Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais
AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.
By Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee