arXiv Computer Vision

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

The paper introduces TAR (Traffic Anomaly Reasoning) and its evaluation suite TAR-Bench, designed to train and assess video‑language models on tasks beyond simple anomaly detection. TAR offers 44,040 chain‑of‑thought annotations covering 10 tasks across 3,670 CCTV videos, while TAR‑Bench supplies 960 human‑curated test annotations from 80 clips of 17 YouTube videos. Experiments show that strong question‑answering performance does not guarantee temporal or scene reasoning, and multi‑task fine‑tuning on TAR consistently improves model scores, with a 10‑task model outperforming its zero‑shot baseline by 21.4 points. The datasets serve as the official training and evaluation data for AI City Challenge 2026 Track 3.

arXiv AI
Aug 14

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

arXiv:2608. 13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users.

By Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
arXiv Computer Vision
Aug 27

TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

TAU-Agent is a retrieval‑augmented framework designed for traffic anomaly understanding in transportation videos. It uses a central retrieval agent to coordinate a Video Captioning Tool and an Open‑Vocabulary Tracking Tool, gathering captions, temporal intervals, and object trajectories relevant to a query. The collected evidence, along with sampled frames and the query, is fed into a fine‑tuned vision‑language model that reasons and generates an answer. TAU-Agent was evaluated on the AI City Challenge 2026 benchmarks, achieving notable scores across multiple tracks and ranking second, twelfth, and fifth respectively.

By Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
arXiv AI
Sep 21

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.

By Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee
arXiv Machine Learning
Jul 7

Incentivizing Vision Language Models to Search for Long Video Question Answering

arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.

By Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali
arXiv Computer Vision
Sep 23

Test-time Reinforcement Learning for Anomalous Video Understanding

The paper introduces a test‑time reinforcement learning framework for anomalous video understanding, addressing challenges such as unreliable pseudo‑labels, inadequate reward design, and collapsed group‑relative advantages. It proposes dual‑query consistency filtering, an entropy‑aware consensus reward, and a virtual negative anchor mechanism to improve sample reliability, reward quality, and policy‑gradient signals. Experiments on VAU‑Bench demonstrate significant performance gains, especially on the ECVA subset where accuracy rises from 75.81% to 90.00%.

By Huining Li, Yuxiang Duan, Jiyang Tan, Qian Li, MingCai Chen, Jian Zhang, Xingdong Sheng, Yuntao Du
arXiv AI
Jul 10

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

arXiv:2605. 21917v2 Announce Type: replace-cross Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support.

By Han Zhang, Wanting Jiang, Tomasz Kornuta, Tian Zheng, Vidya Murali
arXiv AI
6d ago

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang