arXiv Computer Vision
Aug 27

TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

TAU-Agent is a retrieval‑augmented framework designed for traffic anomaly understanding in transportation videos. It uses a central retrieval agent to coordinate a Video Captioning Tool and an Open‑Vocabulary Tracking Tool, gathering captions, temporal intervals, and object trajectories relevant to a query. The collected evidence, along with sampled frames and the query, is fed into a fine‑tuned vision‑language model that reasons and generates an answer. TAU-Agent was evaluated on the AI City Challenge 2026 benchmarks, achieving notable scores across multiple tracks and ranking second, twelfth, and fifth respectively.

By Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
arXiv AI
Aug 14

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

arXiv:2608. 13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users.

By Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
arXiv Computer Vision
Sep 18

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

The paper introduces TAR (Traffic Anomaly Reasoning) and its evaluation suite TAR-Bench, designed to train and assess video‑language models on tasks beyond simple anomaly detection. TAR offers 44,040 chain‑of‑thought annotations covering 10 tasks across 3,670 CCTV videos, while TAR‑Bench supplies 960 human‑curated test annotations from 80 clips of 17 YouTube videos. Experiments show that strong question‑answering performance does not guarantee temporal or scene reasoning, and multi‑task fine‑tuning on TAR consistently improves model scores, with a 10‑task model outperforming its zero‑shot baseline by 21.4 points. The datasets serve as the official training and evaluation data for AI City Challenge 2026 Track 3.

By Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Zheng Tang, Varun Praveen, Vidya N. Murali, David C. Anastasiu, Tomasz Kornuta
arXiv AI
6d ago

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang
arXiv AI
Aug 14

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

arXiv:2608. 12843v1 Announce Type: cross Abstract: Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions.

By Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy Minh Nhat Nguyen