arXiv:2601. 12507v2 Announce Type: replace-cross Abstract: Low-resolution remote sensing small object detection is limited by both missing visual details and the ambiguity of how details serve detection.
By Ruo Qi, Linhui Dai, Yusong Qin, Chaolei Yang, Yanshan Li
MINER is a training‑free inference framework that enhances frozen dual‑encoder models for text‑to‑image retrieval when queries refer to small, visually subordinate objects in cluttered scenes. It augments the global image embedding with a bank of region‑level embeddings and applies hubness‑correcting similarity rescoring, thereby recovering visual evidence that global pooling underweights. The authors introduce ROCS, a benchmark derived from Flickr30K and MS COCO, and demonstrate that MINER improves retrieval performance across CLIP, SigLIP, and SigLIP 2 backbones on both ROCS and standard splits, attributing gains mainly to broader spatial coverage rather than precise crop placement.
By Abdulmalik Alquwayfili, Faisal AlMeshal, Jumanah Almajnouni, Huda Abdulhadi Alamri, Muhammad Kamran J Khan
arXiv:2607.20116v2 Announce Type: replace
Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, diff...
By Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.
arXiv:2511. 12810v2 Announce Type: replace-cross Abstract: Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blend seamlessly into their environments due to high similarity in color, texture, and size.
By Leena Alghamdi, Muhammad Usman, Hafeez Anwar, Abdul Bais, Saeed Anwar
Semantic segmentation of remote sensing imagery requires models that capture both global context and local detail under tight computational budgets. Prior work typically optimizes for one of these axes: attention for global context, convolution for local detail, or compactness for efficiency.