arXiv:2608. 09519v1 Announce Type: cross Abstract: We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware.
By Lazar {\DJ}okovi\'c, Aimee Lin
arXiv:2609.37013v1 Announce Type: cross
Abstract: Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire rele...
By Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise
arXiv:2606.29167v2 Announce Type: replace
Abstract: Dense correspondence on in-the-wild 3D scans must handle severe non-isometric deformation, partial observations, topology artifacts, irregular disc...
By Qilong Liu, Qinfeng Xiao, Chenyuan Yi, Yongsheng Lin, Liying Zhang, Kit-lun Yick
Classical image correspondence is solved at the level of sparse keypoints or dense pixels, but the systems that consume these matches - object-level mapping, topological navigation, scene-graph maintenance - reason about whole objects. Recent work narrows this gap by matchng directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are obtained by pooling features from large 3D foundation models over the masks.
arXiv:2608.22300v1 Announce Type: new
Abstract: Co-registration underlies nearly every multi-temporal and multi-sensor use of optical satellite imagery, and operational products still carry documente...
By Shoukun Sun, Zhe Wang, Sanaz Salati, Jiyin Zhang, Hui Wang, Xiaogang Ma
The paper introduces TFA, a training‑free aggregation technique that calibrates frozen visual foundation models for visual place recognition. TFA uses cross‑codebook agreement, retrieval coverage, and spectral statistics to adjust residual assignment, spectral shaping, and global‑feature fusion without requiring place labels or task‑specific weights. Experiments with a DINOv2‑B backbone show significant Recall@1 gains over existing training‑free methods across multiple benchmarks, demonstrating that reliability‑guided aggregation can unlock additional retrieval performance from frozen representations.
By Xin Li, Zhimin Mao, Shang Wang, Siyuan Duan, Geng Zhang
arXiv:2606.03406v2 Announce Type: replace
Abstract: Reliable correspondence estimation supports image processing and 3D vision tasks, including Structure from Motion, visual localization, and image r...
By Xu Pan, Zhen Pang, Qiyuan Ma, Wei Ji, Shuhan Shen, Xianwei Zheng
arXiv:2608. 16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images.
By Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
The paper introduces the first systematic robustness benchmark for local invisible image watermarking, evaluating five methods across 55 image transformations that include signal distortions, coordinate misalignments, indirect local edits, and direct watermark edits. Results show that all methods are vulnerable to some transformation, with MaskWM achieving the best payload recovery and localization but at the cost of image quality. The study highlights that robustness varies strongly with transformation type, especially noting that geometric misalignment and generative local edits can completely disrupt payload recovery.
By Kai Yao, Bence Szil\'agyi, Sebesty\'en Kamp, M\'at\'e Po\'or, M\'at\'e Szilveszter, Matyas K. Zsoldos, Marc Juarez
arXiv:2506. 12697v3 Announce Type: replace-cross Abstract: Small-object detection in Unmanned Aerial Vehicle (UAV) imagery requires preserving weak local evidence while using broader context to separate tiny foreground targets from cluttered backgrounds.
By Yuxiang Wang, Xuecheng Bai, Chuanzhi Xu, Ying Zhou, Weidong Cai
arXiv:2608.29733v1 Announce Type: new
Abstract: Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure-from-motion (SfM): visually similar but physically disti...
By Gonglin Chen, Ben Southall, Hanyuan Xiao, Wenbin Teng, Haolin Xiong, Tianwen Fu, Junyi Ouyang, Kshitij Singh Minhas, Supun Samarasekera, Rakesh Kumar, Yajie Zhao
Frozen visual foundation models provide transferable features for visual place recognition, but fixed aggregation can suppress useful distinctions in new environments. We introduce TFA, a reliability-...