arXiv:2607.02486v2 Announce Type: replace
Abstract: Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet it...
By Yejun Zhang, Xinjue Wang, Zihan Wang, Esa Rahtu, Juho Kannala
EviViT is a lightweight attachment for pretrained vision transformers that learns where to focus detail in high‑resolution images. It uses human visual‑search traces to supervise a question‑conditioned evidence density, guiding regional re‑reading and efficient visual token allocation. The method connects regional features to the global scene via a sparse, coordinate‑aware bridge, improving fine‑grained accuracy across nine host models while using fewer tokens than global‑only processing.
By Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv:2608.28216v1 Announce Type: new
Abstract: Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is ab...
By Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Md. Sadman Haque, Mohd Ariful Haque
arXiv:2607.20116v2 Announce Type: replace
Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, diff...
By Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang
arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.
By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja