arXiv:2607. 13569v1 Announce Type: cross Abstract: Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture.
By Kaicong Huang, Weiheng Oh, Ruimin Ke
arXiv:2504. 11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual surveys, Bluetooth/WiFi tracking, and Automated Passenger Counters, are often costly, device-dependent, or unable to support individual-level matching.
By Kaicong Huang, Talha Azfar, Jack Reilly, Ruimin Ke
arXiv:2403. 09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation.
By Yiming Ma, Victor Sanchez, Tanaya Guha
arXiv:2608. 14724v1 Announce Type: cross Abstract: The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets.
By Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
By Prakhar Bhardwaj, Simone Weikl, Kilian Mang, Elia Jonas Sandtner
The paper proposes a modular training pipeline for zero‑shot cross‑city object detection that combines a multi‑dataset pre‑training strategy with class‑agnostic objectness distillation and a domain‑resilient augmentation stream featuring a Grayworld transformation. Applied to the RF‑DETR detector, the approach reduces cross‑city distribution gaps while using only 16 GB GPU memory, achieving a 24.29‑point mAP improvement and 1st place on the AI City Challenge Track 6 leaderboard. The authors provide code and data at the referenced GitHub repository.
By Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Hyung-Min Jeon, Chi Dai Tran, Son Hong Phan, Duong Khac Vu, Trinh Le Ba Khanh, Jae Wook Jeon
arXiv:2609.05523v1 Announce Type: cross
Abstract: This paper presents a one-stage learning framework that maps monocular roadside-camera images directly to vehicle states in a ground-fixed coordinate...
By Akos T. Kopeczi-Bocz, Tian Mi, Gabor Orosz, Denes Takacs
The 10th AI City Challenge, held alongside ECCV 2026, celebrates a decade of benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 inception focused on vehicle detection, classification, and tracking, the challenge has expanded into a comprehensive benchmark suite covering multi‑camera perception, multimodal reasoning, synthetic‑to‑real learning, generative forecasting, and privacy‑preserving evaluation. The 2026 edition saw 325 registered teams from 26 countries, with six main tracks—spanning multi‑camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text‑based person anomaly search, generative traffic video forecasting, and cross‑city object detection—plus two out‑of‑domain leaderboards for fisheye traffic‑violation understanding and pedestrian situated‑intent VQA.
By Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro, Yuxing Wang, Yizhou Wang, Sameer Satish Pusegaonkar, Anqi Li, Nalin Dadhich, Ridham Kachhadiya, Dhanishtha Patil, Haoquan Liang, Jiajun Li, Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Shuyu Yang, Ashutosh Kumar, Rong Wang, Rafael Martin Nieto, Peter Christiansen, Ahmed Abduljawad, Mohanrasu Shanmugam, Nadeem Shaik, Sujit Biswas, Xunlei Wu, Vidya Murali, Rama Chellappa
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
The paper introduces a vision‑language framework that extracts social indicators from street‑level imagery, converting panoramic views into sidewalk‑facing sideviews with timestamps. Using a VLM‑based activity detection system, it codes each pedestrian across ten observable dimensions, producing a Social Dwelling Index (SDI) that captures grouping, dwelling, activity diversity, and accessibility flags. Applied to over 100,000 sideviews in New York City, the study finds that pedestrian volume and SDI are only weakly correlated, indicating that high foot traffic does not necessarily equate to intense social activity.
By Liu Liu, Andres Sevtsuk
arXiv:2512.07776v2 Announce Type: replace
Abstract: Monitoring critically endangered western lowland gorillas is currently hampered by the immense manual effort required to re-identify individuals fr...
By Maximilian Schall, Felix Leonard Kn\"ofel, Noah Elias K\"onig, Jan Jonas Kubeler, Maximilian von Klinski, Joan Wilhelm Linnemann, Xiaoshi Liu, Iven Jelle Schlegelmilch, Ole Woyciniuk, Alexandra Schild, Dante Wasmuht, Magdalena Bermejo Espinet, German Illera Basas, Gerard de Melo
arXiv:2602. 18047v4 Announce Type: replace-cross Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift while complying with data protection rules that prevent sharing raw imagery.
By Rong Fu, Yibo Meng, Jia Yee Tan, Rui Lu, Jiekai Wu, Simon Fong