arXiv AI

X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

arXiv:2608. 16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images.

Hugging Face Trending Papers
Jul 22

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.

arXiv Computer Vision
Sep 3

A Top-Down Framework for Metric-Scale Athlete Localization from Single Broadcast Frames

The paper introduces a top‑down framework for accurately locating athletes in metric world coordinates using a single calibrated broadcast frame. It presents three main contributions: a Boundary‑Aware Adaptive Tiling method that expands tile boundaries to avoid splitting athletes across tiles, a specialized two‑keypoint estimator based on RTMPose‑X for pelvis and ground projection points, and a deterministic lift of 2D projections into 3D world coordinates via camera‑calibrated ray casting. The approach achieves a LocSim score of 97.44 and an mAP of 0.9128, surpassing the baseline by over 21 % on a public test set.

By Thanh-Khoi Nguyen, Hoang-Phuc Nguyen, Linh-Huynh, Minh-Triet Tran
arXiv Computer Vision
Sep 3

GeoStore: Finding Small Storefronts in Large Scenes -- A Fine-Grained POI Localization Benchmark with Global-to-Local Asymmetric Matching

GeoStore is a new benchmark for fine‑grained point‑of‑interest (POI) localization that matches close‑up storefront photos against large geo‑tagged street‑view images, a task distinct from traditional visual place recognition. The paper shows that global‑descriptor methods designed for symmetric matching perform poorly on this asymmetric problem, and introduces GLAM, a Global‑to‑Local Asymmetric Matching approach that combines a global retrieval anchor with a lightweight local re‑ranking using pooled region tokens. GLAM achieves higher Recall@1/5/10 and mAP than strong baselines while using far fewer re‑ranking features and significantly lower per‑pair matching cost.

By Lu Han, Xiting Sun, Hao Wang, Zhiqiang Cao, Ruihuan Du, Ziquan Zeng, Chunlong Lv
arXiv Computer Vision
Sep 21

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.

By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
arXiv Computer Vision
Aug 31

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

The paper introduces Parallel Tube Decoding (PTD), a generative approach for spatio‑temporal video grounding that splits the task into a temporal block and simultaneous time‑conditioned spatial blocks, eliminating token‑level and trajectory‑level dependencies. PTD uses Decoupled Block Attention to allow parallel spatial generation while maintaining shared video‑query context, and incorporates localization‑aware policy optimization for temporal boundaries and spatial geometry. Experiments on VidSTG show PTD cuts tube completion latency by 79× and boosts spatial decoding throughput by 92× compared to autoregressive decoding, while improving grounding accuracy and performing well on related tasks such as temporal grounding, VideoQA, and referring video object tracking.

By Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
arXiv Computer Vision
Sep 21

Multi-viewpoint Geo-localization with Event Cameras

The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.

By Adam D. Hines, Michael Milford, Tobias Fischer