arXiv AI By Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

Read the original on arXiv AI →

arXiv:2608. 14724v1 Announce Type: cross Abstract: The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

The paper proposes a modular training pipeline for zero‑shot cross‑city object detection that combines a multi‑dataset pre‑training strategy with class‑agnostic objectness distillation and a domain‑resilient augmentation stream featuring a Grayworld transformation. Applied to the RF‑DETR detector, the approach reduces cross‑city distribution gaps while using only 16 GB GPU memory, achieving a 24.29‑point mAP improvement and 1st place on the AI City Challenge Track 6 leaderboard. The authors provide code and data at the referenced GitHub repository.

By Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Hyung-Min Jeon, Chi Dai Tran, Son Hong Phan, Duong Khac Vu, Trinh Le Ba Khanh, Jae Wook Jeon
arXiv Machine Learning
Jul 23

CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras

arXiv:2602. 18047v4 Announce Type: replace-cross Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift while complying with data protection rules that prevent sharing raw imagery.

By Rong Fu, Yibo Meng, Jia Yee Tan, Rui Lu, Jiekai Wu, Simon Fong
arXiv Computer Vision
Sep 22

Closed-Circuit Television Data as an Emergent Data Source for Urban Rail Platform Crowding Estimation

The paper explores the use of Closed‑Circuit Television (CCTV) footage to estimate urban rail platform crowding in real time. It compares three computer‑vision methods—object detection and counting, crowd‑level classification with a Vision Transformer, and semantic segmentation—to extract crowd-related features. A novel convex ridge regression technique is introduced to convert segmentation outputs into passenger counts, and the methods are evaluated on a privacy‑preserving dataset of over 600 hours of Washington Metropolitan Area Transit Authority (WMATA) video, showing that CCTV alone can provide valuable real‑time crowd estimates.

By Riccardo Fiorista, Awad Abdelhalim, Anson F. Stewart, Gabriel L. Pincus, Ian Thistle, Jinhua Zhao