arXiv AI By Jiahui Cui, Yan Zhao, Kan Wei, Enze Zhu, Peirong Zhang, Lei Wang, Yiru Wang

GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views

Read the original on arXiv AI →

arXiv:2608. 09270v1 Announce Type: cross Abstract: Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

GrabVG is a visual grounding framework for UAV imagery that tackles the challenges of small, densely packed, and visually similar objects. It splits the task into preattentive hypothesis search and graph‑attentive feature binding, using distillation‑guided proposals and a sparse graph to capture intra‑ and inter‑instance relationships. Experiments on AerialVG and AerialSense show that GrabVG achieves higher accuracy and speed, outperforming baselines by significant margins.

By Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu, Jiaxu Tian, Yuheng Li, Guangqian Guo, Shan Gao
arXiv Computer Vision
Aug 28

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

UniGeo is a multimodal large language model designed for text-guided drone geo‑localization, enabling the identification of target regions in large image galleries from natural‑language descriptions. It integrates geo‑semantic understanding, cross‑view semantic generation, and candidate‑level verification within a shared vision‑language framework, establishing stable correspondences among local scene elements, spatial relations, and language. A multi‑stage training strategy progressively refines geo‑semantic learning, cross‑view mapping, and fine‑grained verification, yielding significant performance gains on GeoText‑1652, with R@10 and mAP improvements of 13.59 and 2.83 percentage points respectively.

By Jiahao Wen, Hang Yu, Zhedong Zheng
Hugging Face Trending Papers
Aug 19

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

GrabVG is a visual grounding framework for UAV imagery that tackles the challenges of small, densely packed, and visually similar objects by separating the task into preattentive hypothesis search and graph-attentive feature binding. It first generates a compact set of reliable object hypotheses using distillation-guided proposal induction and text-aware filtering, then constructs a sparse graph where language-guided visual cues and inter-instance topological relationships are jointly bound and propagated via graph attention. Experiments on AerialVG and AerialSense demonstrate that GrabVG achieves a strong accuracy–speed trade‑off, reaching 67.31% and 80.34% Acc@0.5 and outperforming baselines by 10.55 and 8.76 percentage points.