arXiv AI

The Geography of Algorithmic Judgment: LLM Intermediaries, Place Identity, and Racial Steering in Housing Search

arXiv:2606. 06694v1 Announce Type: cross Abstract: Large language models (LLMs) are rapidly assuming an intermediary role in housing search through the integration of listing platforms within conversational interfaces, mediating access to information, search, and recommendations within urban settings.

arXiv AI
Aug 28

Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments

Large language models (LLMs) are increasingly used to guide urban safety decisions, but this study shows that their judgments are more influenced by neighborhood names than by geographic coordinates. Across seven instruct‑tuned models tested on 186 neighborhoods in Los Angeles and Chicago, name‑based ratings varied significantly and correlated with the proportion of locally dominant marginalized groups, while coordinate‑only ratings remained largely flat. The research finds that removing neighborhood names reduces both bias and accuracy, highlighting the complex role of demographic stereotypes and crime signals in LLM safety assessments.

By Huy Nguyen, Yue Lin
arXiv Machine Learning
2d ago

Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning

The paper introduces a network‑based spatial context retrieval pipeline that uses pedestrian street networks and open data (OpenStreetMap, GHS‑POP) to generate compact spatial briefs for open‑weight large language models. It then builds a faithfulness benchmark that labels each model claim by its source—whether grounded in the brief or drawn from training knowledge—and tests models against planted false premises across multiple cities and model configurations. The study finds that model family and generation influence resistance to false premises more than model size, revealing dimensions of spatial reasoning not captured by traditional correctness metrics.

By Joan Perez
arXiv Machine Learning
Sep 2

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

The study investigates whether large language models (LLMs) can predict neighborhood-level human mobility without training data. Using anonymized Cuebiq data across four U.S. metropolitan areas, the authors compare zero‑shot LLM predictions to supervised baselines for various mobility outcomes and assess structural alignment with empirical trends. Results show supervised models outperform LLMs (average accuracy 0.580 vs. 0.435), with LLMs relying on coarse, stable priors that may exhibit biased treatment of protected-group predictors.

By Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez
arXiv AI
Aug 26

PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

PlaceSeek is a human‑centered geospatial retrieval framework that maps natural‑language queries to street‑view images by decomposing queries into functional and affective sub‑intents. It uses a Semantic Grounding Module to verify that candidate images contain the physical evidence needed for the intended activity, and an Affective Alignment Module to re‑rank these candidates based on human urban perception judgments. Evaluated on 31,956 Milan street‑view locations, PlaceSeek achieves high precision and ranking metrics, outperforming several vision‑language baselines and demonstrating the importance of both physical grounding and affective alignment for complex urban spatial queries.

By Ziqi Cui, Shangyu Lou
arXiv Machine Learning
Jul 21

STRATA: A Name-and-Geography Race Inference Model for Fair Lending and Housing Equity Applications

arXiv:2504. 21259v2 Announce Type: replace-cross Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records.

By S. Chalavadi, A. Pastor, T. Leitch
arXiv AI
Sep 16

Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI

The paper reviews ethical and privacy risks of large language model (LLM)–enabled geospatial artificial intelligence (GeoAI), identifying eight recurring issues such as data provenance, spatial privacy, algorithmic bias, and technical risks. It evaluates current responses, noting many remain largely unaddressed or conceptual, and proposes a governance‑aware architecture with enforceable controls illustrated by a flood‑response routing scenario. The authors call for empirical validation, spatially specific interpretability tools, and workforce training to address these emerging risks.

By Maya Subramanian, Devika Jain
arXiv AI
Sep 17

From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale

The paper investigates covert dialect bias in large language models (LLMs) by analyzing how internal probability distributions associate different English varieties—Standard American English, African American Vernacular English, Nigerian Standard English, and Nigerian Pidgin—with housing-related adjectives. Using 260 meaning‑matched sentence quadruples and log‑probability scoring across ten open‑weight LLMs, the study finds that AAVE and NP are consistently linked to more negative adjectives than SAE, with NP experiencing the greatest penalty. The bias varies by context and stereotype cluster, and Nigerian Standard English shows a context‑dependent shift, being favored in formal tenant screening but penalized in more socially proximate scenarios.

By Chowdhury Mohammad Abdullah, Rita Orji
arXiv AI
Aug 25

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

GeoRisk-RAG is a hierarchy‑aware framework that improves the reliability of Retrieval‑Augmented Generation (RAG) by incorporating geographic validity. It estimates geographic applicability using a Directed Acyclic Graph (DAG)‑based distance during context retrieval, enabling selective answering. Experiments on a wildfire‑related QA dataset show that GeoRisk‑RAG reduces false confidence rates from ~0.090 to 0.009 and aligns better with human preferences.

By Meenu Ravi, Shailik Sarkar, Lulwah AlKulaib, Yordanos Tessema, Chang-Tien Lu