arXiv Machine Learning By Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu

Observing Health Outcomes Using Remote Sensing Imagery and Geo-Context Guided Visual Transformer

Read the original on arXiv Machine Learning →

The paper introduces the Geo-Context Guided Visual Transformer, a model that augments remote sensing image analysis with geospatial embeddings and an asymmetric attention module. By converting heterogeneous geospatial variables into patch-aligned representations and assigning geospatial roles to attention heads, the approach improves disease prevalence prediction over existing vision-language and graph-based baselines. Ablation and visualization studies demonstrate its effectiveness and interpretability for health-related remote sensing tasks, especially when comprehensive geospatial data are scarce.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
1d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
Hugging Face Trending Papers
Aug 11

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.