The paper introduces TFA, a training‑free aggregation technique that calibrates frozen visual foundation models for visual place recognition. TFA uses cross‑codebook agreement, retrieval coverage, and spectral statistics to adjust residual assignment, spectral shaping, and global‑feature fusion without requiring place labels or task‑specific weights. Experiments with a DINOv2‑B backbone show significant Recall@1 gains over existing training‑free methods across multiple benchmarks, demonstrating that reliability‑guided aggregation can unlock additional retrieval performance from frozen representations.
By Xin Li, Zhimin Mao, Shang Wang, Siyuan Duan, Geng Zhang
arXiv:2607.20116v2 Announce Type: replace
Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, diff...
By Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation.
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature...
The paper investigates how the choice of retrieval encoder affects cache‑based test‑time adaptation for CLIP. By keeping the memory fixed and varying the retrieval space across sixteen encoders, the authors show that retrieval space can dramatically alter performance, with gains ranging from +0.44 to +19.7 points on ImageNet‑A. They introduce MARC, a training‑free system that pairs frozen CLIP with DINOv2‑B for retrieval, achieving superior out‑of‑distribution accuracy and efficiency compared to prior methods.
By Mahir Shahriar Tamim, Md. Samiul Alim, Azmine Toushik Wasi, Shahriyar Zaman Ridoy, Meharun Nesa, Mohammad Abu Yousuf, Alex Lamb, Mohammad Ali Moni
The paper addresses lifelong aerial autonomy by treating visual place recognition (VPR) as a mission‑based domain‑incremental learning problem. It introduces a heterogeneous memory framework that first trains on a static satellite exemplar memory and then uses a bounded replay buffer to retain selected airborne observations across missions. The proposed DBS‑Hybrid replay strategy, which blends prototype‑based diversity trimming with representative‑first feature‑space coverage, outperforms baseline methods in accuracy, generalization, and knowledge retention across multiple UAV missions.
By Xingyu Shao, Zhiqiang Yan, Liangzheng Sun, Mengfan He, Chao Chen, Jinhui Zhang, Chunyu Li, Ziyang Meng
arXiv:2607. 09825v1 Announce Type: cross Abstract: Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid.
By Yi Li (TU Darmstadt), Alexandre Chapin (LIRIS), Liming Chen (LIRIS), Jan Peters (TU Darmstadt), Alap Kshirsagar (IIT Delhi, ADU)
arXiv:2607.22068v2 Announce Type: replace-cross
Abstract: Multi-branch architectures and CNN-Transformer fusion are widely believed to improve vehicle re-identification (Re-ID) by combining complemen...
By Yu Wang, Hongyu Yang
GeoStore is a new benchmark for fine‑grained point‑of‑interest (POI) localization that matches close‑up storefront photos against large geo‑tagged street‑view images, a task distinct from traditional visual place recognition. The paper shows that global‑descriptor methods designed for symmetric matching perform poorly on this asymmetric problem, and introduces GLAM, a Global‑to‑Local Asymmetric Matching approach that combines a global retrieval anchor with a lightweight local re‑ranking using pooled region tokens. GLAM achieves higher Recall@1/5/10 and mAP than strong baselines while using far fewer re‑ranking features and significantly lower per‑pair matching cost.
By Lu Han, Xiting Sun, Hao Wang, Zhiqiang Cao, Ruihuan Du, Ziquan Zeng, Chunlong Lv
arXiv:2607. 22068v1 Announce Type: cross Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combining complementary representations.
By Yu Wang, Hongyu Yang
SCOUT is a frozen‑encoder approach for sim‑to‑real text‑based person retrieval that predicts cross‑modal embeddings instead of fine‑tuning cross‑encoders. It uses a trainable predictor to map patch tokens from a frozen video encoder (V‑JEPA) into the embedding space of a frozen text encoder (EmbeddingGemma), guided by a bidirectional InfoNCE objective. The method achieves state‑of‑the‑art results on the AI City Challenge 2026 Track 4, with a full retrieve‑fuse‑rerank pipeline reaching 84.25 mAP@10 and a single frozen model alone scoring 60.63, while training costs are modest (≈95 GPU‑hours).
By Abdarahmane Traor\'e, Andy Couturier, \'Eric Hervet
MINER is a training‑free inference framework that enhances frozen dual‑encoder models for text‑to‑image retrieval when queries refer to small, visually subordinate objects in cluttered scenes. It augments the global image embedding with a bank of region‑level embeddings and applies hubness‑correcting similarity rescoring, thereby recovering visual evidence that global pooling underweights. The authors introduce ROCS, a benchmark derived from Flickr30K and MS COCO, and demonstrate that MINER improves retrieval performance across CLIP, SigLIP, and SigLIP 2 backbones on both ROCS and standard splits, attributing gains mainly to broader spatial coverage rather than precise crop placement.
By Abdulmalik Alquwayfili, Faisal AlMeshal, Jumanah Almajnouni, Huda Abdulhadi Alamri, Muhammad Kamran J Khan