arXiv Machine Learning

Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning

arXiv:2608. 06934v1 Announce Type: new Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences.

arXiv Machine Learning
Sep 22

MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments

arXiv:2505.12254v3 Announce Type: replace-cross Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...

By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv Computer Vision
Sep 25

A Vision-Language Framework for Measuring Social Life on Sidewalks

The paper introduces a vision‑language framework that extracts social indicators from street‑level imagery, converting panoramic views into sidewalk‑facing sideviews with timestamps. Using a VLM‑based activity detection system, it codes each pedestrian across ten observable dimensions, producing a Social Dwelling Index (SDI) that captures grouping, dwelling, activity diversity, and accessibility flags. Applied to over 100,000 sideviews in New York City, the study finds that pedestrian volume and SDI are only weakly correlated, indicating that high foot traffic does not necessarily equate to intense social activity.

By Liu Liu, Andres Sevtsuk
arXiv Machine Learning
Aug 18

In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults

arXiv:2608. 14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults.

By Houhao Liang, Kresimir Friganovic, Joanne Kua, Noor Hafizah Ismail, Su Su, Bryan Yijia Tan, Navrag B. Singh, Panos Mavros
arXiv AI
3d ago

Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing

The paper introduces an explainable pedestrian routing method that estimates perceived safety by first generating a natural‑language caption from street‑level images and then deriving risk scores solely from structured features of that caption. Benchmarking nine captioning setups against a CLIP image‑embedding baseline shows comparable performance, and the system was deployed on over 650,000 images across 36 wards in Manchester and Huddersfield. Field validation with 3,669 ratings from 70 participants revealed a modest but statistically significant correlation (r = 0.262) with human judgments, while a stronger supervised benchmark did not translate into better real‑world performance.

By Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi
Hugging Face Trending Papers
Jul 23

Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment.

arXiv Computer Vision
Sep 18

MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

MTF‑Net is a Multi‑Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. It fuses four modalities—bounding‑box dynamics, human pose keypoints, local context, and scene‑level semantics—within a recurrent framework enhanced by gated linear units (GLUs) and an attention‑guided fusion head. Evaluations on the PIE and JAAD benchmarks show that MTF‑Net outperforms recent transformer‑ and graph‑based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD while maintaining real‑time performance.

By Md Mahfuzur Rahman, Pengzhan Zhou, A. F. M. Abdun Noor, Md Imam Ahasan, Md Mustafizur Rahman, Fang Qu
arXiv Machine Learning
Aug 21

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.

By Joan Perez, Giovanni Fusco
arXiv AI
Aug 19

The 10th AI City Challenge

The 10th AI City Challenge, held alongside ECCV 2026, celebrates a decade of benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 inception focused on vehicle detection, classification, and tracking, the challenge has expanded into a comprehensive benchmark suite covering multi‑camera perception, multimodal reasoning, synthetic‑to‑real learning, generative forecasting, and privacy‑preserving evaluation. The 2026 edition saw 325 registered teams from 26 countries, with six main tracks—spanning multi‑camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text‑based person anomaly search, generative traffic video forecasting, and cross‑city object detection—plus two out‑of‑domain leaderboards for fisheye traffic‑violation understanding and pedestrian situated‑intent VQA.

By Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro, Yuxing Wang, Yizhou Wang, Sameer Satish Pusegaonkar, Anqi Li, Nalin Dadhich, Ridham Kachhadiya, Dhanishtha Patil, Haoquan Liang, Jiajun Li, Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Shuyu Yang, Ashutosh Kumar, Rong Wang, Rafael Martin Nieto, Peter Christiansen, Ahmed Abduljawad, Mohanrasu Shanmugam, Nadeem Shaik, Sujit Biswas, Xunlei Wu, Vidya Murali, Rama Chellappa
arXiv Machine Learning
Sep 17

Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?

The study evaluates whether vision‑language models (VLMs) can reliably assess sidewalk accessibility attributes—effective width, longitudinal slope, cross slope, and pavement condition—from pedestrian‑level images. Using sampling‑based conformal prediction on 514 images from Seoul, the authors find that calibrated models achieve nominal 90% coverage, but only effective width yields informative estimates; other attributes remain too uncertain for compliance assessment. The work also demonstrates that raw sampling dispersion is not a trustworthy uncertainty measure without calibration and releases annotated images with ground‑truth measurements.

By Seung Jae Lieu, Diego Morra, Chiara Cadoni, Wonseop Song, Martina Mazzarello, Carlo Ratti