Hugging Face Trending Papers

Remote-Sensing City Layout Extraction with MLLM

Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector boundaries. Such outputs locate classes and support image-plane scoring, yet they do not by themselves constitute an executable layout that retains object identities, typed relations, topology, and regeneration rules.

arXiv Computer Vision
Sep 1

SVI2LoD3: Agent-Driven Reconstruction of LoD3 Facade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models

arXiv:2608.29992v1 Announce Type: new Abstract: This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of facade openings in 3D city models, producing directly usable Ci...

By Elmehdi Kanna, Lukas Arzoumanidis, Huynh Duc An Son Nguyen, Youness Dehbi
arXiv Machine Learning
Aug 21

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.

By Joan Perez, Giovanni Fusco
arXiv Computer Vision
2d ago

Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing

The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.

By Kaizhen Tan
arXiv AI
2d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv AI
Sep 10

Knowledge-Guided Vision-Language Inference for Image-Based Urban Flood Depth Estimation

The paper introduces FloodVision, a knowledge-guided vision‑language framework that estimates urban flood depth from a single RGB image. It combines a general‑purpose vision‑language model with FloodKG, a domain knowledge base that encodes canonical object dimensions and component landmarks to promote component‑level reasoning. On 654 crowdsourced New York flood images, FloodVision cuts mean absolute error from 15.62 cm to 8.75 cm and median error from 14.35 cm to 7.75 cm, outperforming the VLM‑only baseline in 69.3 % of cases.

By Zhangding Liu, Neda Mohammadi, John E. Taylor