arXiv Computer Vision By Thanh-Khoi Nguyen, Hoang-Phuc Nguyen, Linh-Huynh, Minh-Triet Tran

A Top-Down Framework for Metric-Scale Athlete Localization from Single Broadcast Frames

Read the original on arXiv Computer Vision →

The paper introduces a top‑down framework for accurately locating athletes in metric world coordinates using a single calibrated broadcast frame. It presents three main contributions: a Boundary‑Aware Adaptive Tiling method that expands tile boundaries to avoid splitting athletes across tiles, a specialized two‑keypoint estimator based on RTMPose‑X for pelvis and ground projection points, and a deterministic lift of 2D projections into 3D world coordinates via camera‑calibrated ray casting. The approach achieves a LocSim score of 97.44 and an mAP of 0.9128, surpassing the baseline by over 21 % on a public test set.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 7

ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization

ARC‑Loc introduces a new cross‑view localization method that bypasses heavy Bird’s‑Eye‑View transformations and external depth models. By converting ground keypoints into azimuthal rays on a satellite map and exploiting their convergence at the user’s location, the approach uses a minimal Azimuthal Ray Convergence solver and an ARC loss to directly match ground and satellite images. Experiments on VIGOR and KITTI show that ARC‑Loc achieves competitive accuracy while offering faster, memory‑efficient inference and easy integration with existing frameworks.

By Hyeongsik Kim, Mincheol Kim, Heejoon Moon, Je Hyeong Hong
arXiv Computer Vision
Sep 21

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.

By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang