arXiv Computer Vision By Yupeng Zhang, Fangzhuo Gao, Juntao Cheng, Ziyi Zhao, Liang Wan, Ruize Han

LiG-DETR: Local-in-Global Reassembly in Latent Space for Aerial Object Detection

Read the original on arXiv Computer Vision →

LiG-DETR introduces a Global-Local Reassembly framework for aerial object detection that captures high‑fidelity local features before compression and integrates them into a unified end‑to‑end DETR decoder. The method uses a shared encoder to extract both global and locally magnified features, reassembles the local features according to their spatial positions, and employs Context‑Preserved Selective Reassembly and Density‑Aware Adaptive Query Allocation to reduce redundant computation. Experiments demonstrate significant improvements on small and medium objects while maintaining strong performance on large objects, with favorable accuracy–efficiency trade‑offs and better cross‑domain generalization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 24

MINER: Multi-crop INference-time Enhancement for Rare-Object Retrieval with Frozen Dual Encoders

MINER is a training‑free inference framework that enhances frozen dual‑encoder models for text‑to‑image retrieval when queries refer to small, visually subordinate objects in cluttered scenes. It augments the global image embedding with a bank of region‑level embeddings and applies hubness‑correcting similarity rescoring, thereby recovering visual evidence that global pooling underweights. The authors introduce ROCS, a benchmark derived from Flickr30K and MS COCO, and demonstrate that MINER improves retrieval performance across CLIP, SigLIP, and SigLIP 2 backbones on both ROCS and standard splits, attributing gains mainly to broader spatial coverage rather than precise crop placement.

By Abdulmalik Alquwayfili, Faisal AlMeshal, Jumanah Almajnouni, Huda Abdulhadi Alamri, Muhammad Kamran J Khan
Hugging Face Trending Papers
Jul 22

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.