arXiv:2510. 03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastructure-based measures are often insufficient in dynamic urban environments.
By Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Muhammad Monjurul Karim, Kehua Chen, Chenxi Liu, Mehrdad Nasri, Yinhai Wang
arXiv:2603. 22531v2 Announce Type: replace Abstract: Sidewalk width is an important indicator of pedestrian accessibility, comfort, and network quality, yet large-scale width data remain scarce in most cities.
By Kaizhen Tan, Fan Zhang
The study evaluates whether vision‑language models (VLMs) can reliably assess sidewalk accessibility attributes—effective width, longitudinal slope, cross slope, and pavement condition—from pedestrian‑level images. Using sampling‑based conformal prediction on 514 images from Seoul, the authors find that calibrated models achieve nominal 90% coverage, but only effective width yields informative estimates; other attributes remain too uncertain for compliance assessment. The work also demonstrates that raw sampling dispersion is not a trustworthy uncertainty measure without calibration and releases annotated images with ground‑truth measurements.
By Seung Jae Lieu, Diego Morra, Chiara Cadoni, Wonseop Song, Martina Mazzarello, Carlo Ratti
iOSPointMapper is a mobile app that performs real‑time, privacy‑conscious sidewalk mapping using on‑device semantic segmentation, LiDAR depth estimation, and fused GPS/IMU data on recent iPhones and iPads. It detects and localizes sidewalk‑relevant features such as traffic signs, traffic lights, and poles, and includes a user‑guided annotation interface for validating outputs before submission. The anonymized data is transmitted to the Transportation Data Exchange Initiative (TDEI), where it integrates with broader multimodal transportation datasets, and evaluations show the app’s potential for enhanced pedestrian mapping.
By Himanshu Naidu, Yuxiang Zhang, Sachin Mehta, Anat Caspi
The paper introduces an explainable pedestrian routing method that estimates perceived safety by first generating a natural‑language caption from street‑level images and then deriving risk scores solely from structured features of that caption. Benchmarking nine captioning setups against a CLIP image‑embedding baseline shows comparable performance, and the system was deployed on over 650,000 images across 36 wards in Manchester and Huddersfield. Field validation with 3,669 ratings from 70 participants revealed a modest but statistically significant correlation (r = 0.262) with human judgments, while a stronger supervised benchmark did not translate into better real‑world performance.
By Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi
The paper introduces Smol‑VL‑BLV, a compact vision‑language model designed for blind and low‑vision users. It employs a 500M decoder transformer with teacher‑student distillation and Group Relative Policy Optimization to add spatial detail, directional cues, and hazard detection to post‑training. After a lightweight finetuning step, the model achieves significant gains on spatial, social, OCR, and VQA benchmarks while remaining under 450 MB and running entirely offline on a mid‑range Android phone.
By Rishabh Choudhary, Shreyansh Raj, Umesh Goyal, Shubh Kashyap, Shrestha Kumar, Sushovan Jena, Komal Kumar, Hisham Cholakkal, Aditya Nigam