arXiv AI

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

arXiv:2608. 14724v1 Announce Type: cross Abstract: The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets.

arXiv AI
Aug 26

Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

The paper proposes a modular training pipeline for zero‑shot cross‑city object detection that combines a multi‑dataset pre‑training strategy with class‑agnostic objectness distillation and a domain‑resilient augmentation stream featuring a Grayworld transformation. Applied to the RF‑DETR detector, the approach reduces cross‑city distribution gaps while using only 16 GB GPU memory, achieving a 24.29‑point mAP improvement and 1st place on the AI City Challenge Track 6 leaderboard. The authors provide code and data at the referenced GitHub repository.

By Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Hyung-Min Jeon, Chi Dai Tran, Son Hong Phan, Duong Khac Vu, Trinh Le Ba Khanh, Jae Wook Jeon
arXiv Machine Learning
Jul 23

CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras

arXiv:2602. 18047v4 Announce Type: replace-cross Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift while complying with data protection rules that prevent sharing raw imagery.

By Rong Fu, Yibo Meng, Jia Yee Tan, Rui Lu, Jiekai Wu, Simon Fong
arXiv Computer Vision
Sep 22

Closed-Circuit Television Data as an Emergent Data Source for Urban Rail Platform Crowding Estimation

The paper explores the use of Closed‑Circuit Television (CCTV) footage to estimate urban rail platform crowding in real time. It compares three computer‑vision methods—object detection and counting, crowd‑level classification with a Vision Transformer, and semantic segmentation—to extract crowd-related features. A novel convex ridge regression technique is introduced to convert segmentation outputs into passenger counts, and the methods are evaluated on a privacy‑preserving dataset of over 600 hours of Washington Metropolitan Area Transit Authority (WMATA) video, showing that CCTV alone can provide valuable real‑time crowd estimates.

By Riccardo Fiorista, Awad Abdelhalim, Anson F. Stewart, Gabriel L. Pincus, Ian Thistle, Jinhua Zhao
arXiv AI
Aug 28

Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations

The paper argues that evaluating privacy‑enhancing technologies (PETs) solely through image classification is insufficient because classification remains robust to many geometric and local perturbations. It proposes a compute‑aware multi‑task protocol that uses lightweight proxy tasks to assess PETs across various transformations, revealing that PETs with similar classification accuracy can perform very differently on other vision tasks. The study demonstrates the necessity of broader evaluation metrics beyond classification to truly gauge PET effectiveness.

By Leon Ranke, Wolfgang H\"ubner, Ronny Hug, Michael Arens, J\"urgen Beyerer
arXiv Machine Learning
Aug 20

On the Robustness of Vision-Language Models in Zero-shot Privacy Classification

The paper investigates whether large, instruction‑following Vision‑Language Models (VLMs) can reliably perform zero‑shot image privacy classification. It compares three open‑source VLMs to specialized privacy models on two public benchmarks, evaluating accuracy, robustness to image degradations (compression, lighting changes, noise), inference speed, and parameter count. The findings show that while VLMs remain robust to perturbations, they are less accurate and significantly slower than smaller, purpose‑built privacy models, indicating that scaling alone does not guarantee effective privacy classification.

By Alina Elena Baia, Alessio Xompero, Andrea Cavallaro
arXiv Computer Vision
Sep 14

MPT: Missing Prototype Tracking via Barycentric Reconstruction in Vehicular Federated Learning

arXiv:2609.12771v1 Announce Type: cross Abstract: Cross-vehicle federated learning enables vehicles to collaboratively improve perception models while keeping locally collected driving data private....

By Hanju Jang (Yonsei University), Gyeongmin Han (Yonsei University), Sungmin Lee (Yonsei University), Kichang Lee (Yonsei University), Chunghan Lee (Toyota Motor Corporation), JeongGil Ko (Yonsei University)