arXiv Computer Vision

Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications

The paper introduces GBA-GCs, a large-scale multimodal benchmark for identifying gated and open residential compounds in China’s Greater Bay Area, comprising 37,444 compounds with satellite imagery, metadata, and verified labels. It presents MCGC, a vision-centric multimodal framework that fuses imagery, text, and structured data to accurately classify gated communities, outperforming existing baselines. Using the model, the authors map gated communities across the metropolitan area and uncover equity-related patterns such as clustered gated zones, privatized green space, and diminished pedestrian connectivity.

Hugging Face Trending Papers
Jul 22

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment.

arXiv Machine Learning
Sep 22

MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments

arXiv:2505.12254v3 Announce Type: replace-cross Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...

By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv Computer Vision
Sep 18

Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning

The paper introduces an adaptive region‑dividing strategy that projects a 3D point cloud onto a bird’s‑eye‑view plane to detect building regions, then back‑projects bounding boxes to create structure‑aligned training blocks for unified scene‑level evaluation. It also proposes a fine‑grained classification model using a point transformer classifier and a spatially‑supervised contrastive loss to improve inter‑class discriminability, addressing class imbalance with a weighted cross‑entropy. Experiments on UrbanBIS and STPLS3D datasets show the method outperforms state‑of‑the‑art approaches in both building instance segmentation and fine‑grained classification.

By Weiyuan Zhang, Qi Zhang, Hui Huang
arXiv Machine Learning
4d ago

DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology

DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.

By Joshua Dimasaka, Christian Gei{\ss}, Emily So
arXiv AI
Sep 16

Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM

The paper presents a multimodal machine learning framework that classifies Emirati residential architectural styles by combining visual features from images and textual descriptions using OpenAI's CLIP model. The unified 512‑dimensional embeddings are reduced with UMAP, clustered with K‑Means, and then used to train an SVM classifier, achieving a 98% accuracy across eight style clusters. This approach outperforms previous studies and demonstrates the value of integrating visual and textual data for cultural heritage analysis.

By Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir
arXiv AI
Aug 26

PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

PlaceSeek is a human‑centered geospatial retrieval framework that maps natural‑language queries to street‑view images by decomposing queries into functional and affective sub‑intents. It uses a Semantic Grounding Module to verify that candidate images contain the physical evidence needed for the intended activity, and an Affective Alignment Module to re‑rank these candidates based on human urban perception judgments. Evaluated on 31,956 Milan street‑view locations, PlaceSeek achieves high precision and ranking metrics, outperforming several vision‑language baselines and demonstrating the importance of both physical grounding and affective alignment for complex urban spatial queries.

By Ziqi Cui, Shangyu Lou
arXiv Machine Learning
Sep 1

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

BEACON is a tri‑modal contrastive learning framework that enriches AlphaEarth embeddings by aligning physical representations from Earth‑observation imagery with semantic POI text and human behavioral POI visitation data, while keeping the deployed model image‑only. In a Houston case study, BEACON outperformed six baselines on nine downstream tasks, achieving up to 43% higher R² for obesity prevalence, 34% for poor mental health, and 22% for median household income under a linear probe.

By Hao Tian, Heng Cai, Yifan Yang
arXiv Computer Vision
1d ago

Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing

The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.

By Kaizhen Tan