Multimodal Large Language Models Predict Urban Safety Perception but Encode Non-Neutral Demographic Priors
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Large language models (LLMs) are increasingly used to guide urban safety decisions, but this study shows that their judgments are more influenced by neighborhood names than by geographic coordinates. Across seven instruct‑tuned models tested on 186 neighborhoods in Los Angeles and Chicago, name‑based ratings varied significantly and correlated with the proportion of locally dominant marginalized groups, while coordinate‑only ratings remained largely flat. The research finds that removing neighborhood names reduces both bias and accuracy, highlighting the complex role of demographic stereotypes and crime signals in LLM safety assessments.
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
arXiv:2606. 00871v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing, mapping, and public consultation.
The paper investigates how large language and visual‑language models used in autonomous vehicles inherit human driver biases when deciding whether to yield to pedestrians. It introduces two new bias‑testing methods—All Else Being Equal and Self‑Consistency tests—to evaluate these models. Results reveal that the models’ yielding decisions are influenced by pedestrian attributes such as gender, ethnicity, religion, disability, age, skin tone, and socio‑economic status, with varying patterns across models.
arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.
The paper introduces MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile, to evaluate personalized safety in vision‑language models (VLMs). Eight leading VLMs were tested and found to almost always respond directly (86‑99%) without seeking missing context, scoring no higher than 2.6/5 on personalized safety. The authors identify a phenomenon called visual dominance, where visual information enters text representations early and suppresses textual risk signals, and propose PRISM, a lightweight input monitor that predicts when a query should be deferred, achieving 0.978 AUC and outperforming all tested models on the safety‑utility Pareto frontier.