Hugging Face Trending Papers

Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and uses context-aware sampling to build compact, scene‑specific training sets for ship detection. By enriching frames with weather and solar metadata and applying diversity sampling, the method reduces 40,718 candidate frames to just 220 for annotation, achieving a 99.5% reduction. Fine‑tuned YOLOv26‑m on this curated subset attains high detection performance (AP50 ≈ 94.8% and AP50‑95 ≈ 75.1%).

arXiv AI
Sep 25

Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and then selects a context‑aware, diverse subset for ship detection training. By combining SuperPoint‑LightGlue localization, weather and solar‑state metadata, and diversity sampling, the method reduces 40,718 candidate frames to just 220 images for annotation. Fine‑tuned YOLOv26 on this compact set achieves high detection performance (AP50 ≈ 94.8%) while cutting annotation effort by 99.5%.

By Ignat Romanov, Andreas Hadjipieris, Neofytos Dimitriou
Hugging Face Trending Papers
Jul 22

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.

arXiv AI
4d ago

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.

By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv Computer Vision
Aug 27

OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization

OpenCVL is a large, open dataset for fine-grained cross-view localization, comprising 617,388 ground‑aerial image pairs from 41 European cities. It blends high‑end sensor data with diverse in‑the‑wild images and includes a curation framework to correct pose annotations, enabling reliable evaluation. The dataset also offers cross‑area and snowy test sets to probe generalization, and experiments show that adding noisy in‑the‑wild data improves model performance on clean tests.

By Zimin Xia, Mubariz Zaffar, Junsheng Fu, Alexandre Alahi, Julian F. P. Kooij
arXiv Computer Vision
Sep 17

PDA++: Field-Aligned Planning and Scene-Adaptive Insertion in Remote Sensing

PDA++ is a unified, environment‑aware object insertion framework for remote sensing imagery that improves few‑shot and long‑tail recognition. It operates in three stages: Planning, which selects scene‑compatible poses using an affordance field; Decoupling, which conditions the background on pose to preserve object identity while adapting to the scene; and Assimilation, which aligns multi‑scale texture distributions via optimal transport to enhance local coherence. The method achieves a whole‑image FID of 6.28 and boosts average few‑shot recognition mAP50 by 17.69 points on optical data, while also improving ship detection on SAR imagery and maintaining performance under cross‑dataset transfer and amorphous‑target insertion.

By Xianchi Dong, Yingyan Hou, Chao Ren, Wanxuan Lu, Zihan Wei, Hongfeng Yu, Yixiao Wang, Chubo Deng, Xian Sun