arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
The paper introduces a vision‑language framework that extracts social indicators from street‑level imagery, converting panoramic views into sidewalk‑facing sideviews with timestamps. Using a VLM‑based activity detection system, it codes each pedestrian across ten observable dimensions, producing a Social Dwelling Index (SDI) that captures grouping, dwelling, activity diversity, and accessibility flags. Applied to over 100,000 sideviews in New York City, the study finds that pedestrian volume and SDI are only weakly correlated, indicating that high foot traffic does not necessarily equate to intense social activity.
By Liu Liu, Andres Sevtsuk
arXiv:2608. 14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults.
By Houhao Liang, Kresimir Friganovic, Joanne Kua, Noor Hafizah Ismail, Su Su, Bryan Yijia Tan, Navrag B. Singh, Panos Mavros
The paper introduces an explainable pedestrian routing method that estimates perceived safety by first generating a natural‑language caption from street‑level images and then deriving risk scores solely from structured features of that caption. Benchmarking nine captioning setups against a CLIP image‑embedding baseline shows comparable performance, and the system was deployed on over 650,000 images across 36 wards in Manchester and Huddersfield. Field validation with 3,669 ratings from 70 participants revealed a modest but statistically significant correlation (r = 0.262) with human judgments, while a stronger supervised benchmark did not translate into better real‑world performance.
By Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi
We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment.
arXiv:2607. 23822v1 Announce Type: new Abstract: Driving style captures stable, driver-specific patterns in how a vehicle is driven.
By Yuhang Wang, Lingyao Li, Hao Zhou
MTF‑Net is a Multi‑Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. It fuses four modalities—bounding‑box dynamics, human pose keypoints, local context, and scene‑level semantics—within a recurrent framework enhanced by gated linear units (GLUs) and an attention‑guided fusion head. Evaluations on the PIE and JAAD benchmarks show that MTF‑Net outperforms recent transformer‑ and graph‑based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD while maintaining real‑time performance.
By Md Mahfuzur Rahman, Pengzhan Zhou, A. F. M. Abdun Noor, Md Imam Ahasan, Md Mustafizur Rahman, Fang Qu
arXiv:2602.10124v2 Announce Type: replace-cross
Abstract: Cycling is reported by an average of 35% of adults at least once per week across 28 countries, and as vulnerable road users directly exposed...
By Haining Ding, Chenxi Wang, Simon Ladouce, Michal Gath-Morad
arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.
By Joan Perez, Giovanni Fusco
The 10th AI City Challenge, held alongside ECCV 2026, celebrates a decade of benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 inception focused on vehicle detection, classification, and tracking, the challenge has expanded into a comprehensive benchmark suite covering multi‑camera perception, multimodal reasoning, synthetic‑to‑real learning, generative forecasting, and privacy‑preserving evaluation. The 2026 edition saw 325 registered teams from 26 countries, with six main tracks—spanning multi‑camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text‑based person anomaly search, generative traffic video forecasting, and cross‑city object detection—plus two out‑of‑domain leaderboards for fisheye traffic‑violation understanding and pedestrian situated‑intent VQA.
By Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro, Yuxing Wang, Yizhou Wang, Sameer Satish Pusegaonkar, Anqi Li, Nalin Dadhich, Ridham Kachhadiya, Dhanishtha Patil, Haoquan Liang, Jiajun Li, Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Shuyu Yang, Ashutosh Kumar, Rong Wang, Rafael Martin Nieto, Peter Christiansen, Ahmed Abduljawad, Mohanrasu Shanmugam, Nadeem Shaik, Sujit Biswas, Xunlei Wu, Vidya Murali, Rama Chellappa
The study evaluates whether vision‑language models (VLMs) can reliably assess sidewalk accessibility attributes—effective width, longitudinal slope, cross slope, and pavement condition—from pedestrian‑level images. Using sampling‑based conformal prediction on 514 images from Seoul, the authors find that calibrated models achieve nominal 90% coverage, but only effective width yields informative estimates; other attributes remain too uncertain for compliance assessment. The work also demonstrates that raw sampling dispersion is not a trustworthy uncertainty measure without calibration and releases annotated images with ground‑truth measurements.
By Seung Jae Lieu, Diego Morra, Chiara Cadoni, Wonseop Song, Martina Mazzarello, Carlo Ratti