arXiv AI

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

arXiv:2608. 13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users.

arXiv Computer Vision
Sep 18

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

The paper introduces TAR (Traffic Anomaly Reasoning) and its evaluation suite TAR-Bench, designed to train and assess video‑language models on tasks beyond simple anomaly detection. TAR offers 44,040 chain‑of‑thought annotations covering 10 tasks across 3,670 CCTV videos, while TAR‑Bench supplies 960 human‑curated test annotations from 80 clips of 17 YouTube videos. Experiments show that strong question‑answering performance does not guarantee temporal or scene reasoning, and multi‑task fine‑tuning on TAR consistently improves model scores, with a 10‑task model outperforming its zero‑shot baseline by 21.4 points. The datasets serve as the official training and evaluation data for AI City Challenge 2026 Track 3.

By Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Zheng Tang, Varun Praveen, Vidya N. Murali, David C. Anastasiu, Tomasz Kornuta
arXiv Computer Vision
Aug 27

TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

TAU-Agent is a retrieval‑augmented framework designed for traffic anomaly understanding in transportation videos. It uses a central retrieval agent to coordinate a Video Captioning Tool and an Open‑Vocabulary Tracking Tool, gathering captions, temporal intervals, and object trajectories relevant to a query. The collected evidence, along with sampled frames and the query, is fed into a fine‑tuned vision‑language model that reasons and generates an answer. TAU-Agent was evaluated on the AI City Challenge 2026 benchmarks, achieving notable scores across multiple tracks and ranking second, twelfth, and fifth respectively.

By Yuqiang Lin, Yan Shi, Sam Lockyer, Harish Tayyar Madabushi, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
arXiv AI
Aug 19

The 10th AI City Challenge

The 10th AI City Challenge, held alongside ECCV 2026, celebrates a decade of benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 inception focused on vehicle detection, classification, and tracking, the challenge has expanded into a comprehensive benchmark suite covering multi‑camera perception, multimodal reasoning, synthetic‑to‑real learning, generative forecasting, and privacy‑preserving evaluation. The 2026 edition saw 325 registered teams from 26 countries, with six main tracks—spanning multi‑camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text‑based person anomaly search, generative traffic video forecasting, and cross‑city object detection—plus two out‑of‑domain leaderboards for fisheye traffic‑violation understanding and pedestrian situated‑intent VQA.

By Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro, Yuxing Wang, Yizhou Wang, Sameer Satish Pusegaonkar, Anqi Li, Nalin Dadhich, Ridham Kachhadiya, Dhanishtha Patil, Haoquan Liang, Jiajun Li, Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Shuyu Yang, Ashutosh Kumar, Rong Wang, Rafael Martin Nieto, Peter Christiansen, Ahmed Abduljawad, Mohanrasu Shanmugam, Nadeem Shaik, Sujit Biswas, Xunlei Wu, Vidya Murali, Rama Chellappa
arXiv Computer Vision
Aug 21

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

arXiv:2608. 19380v1 Announce Type: new Abstract: While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpret traffic accidents.

By Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich
arXiv AI
Jun 16

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

arXiv:2606. 15749v1 Announce Type: cross Abstract: Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics.

By Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen
arXiv AI
Jun 4

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

arXiv:2512. 05277v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.

By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari
arXiv AI
Jun 2

From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model

arXiv:2512. 05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.

By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari
arXiv AI
6d ago

TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

TrafficImag is the first benchmark designed to evaluate counterfactual roadside traffic video generation, combining a large roadside dataset with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is encoded as an actor-level program specifying target actor, intended behavior, legal route, interaction order, and temporal constraints, allowing a unified evaluation across diverse foundation models. The benchmark assesses four validity dimensions—initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation—and reports that the best models achieve 80.4% macro F1 for reasoning and 55.0% end-to-end success when using a complete condition interface.

By Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
arXiv AI
Sep 21

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.

By Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee
arXiv Computer Vision
Sep 17

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.

By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui