arXiv AI By Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya, David Hall, Kaushik Roy, Peyman Moghadam

Geometry-Conditioned Visual Place Recognition in Natural Environments

Read the original on arXiv AI →

The paper introduces Depth‑Aware Distillation (DAD), a method that conditions a pretrained Vision Foundation Model’s token representations on geometry inferred by a Geometric Foundation Model, without using a depth sensor. DAD projects image‑aligned depth into the VFM token space and selectively modulates visual representations through channel‑wise geometric conditioning. On the WildCross benchmark, DAD raises average inter‑sequence Recall@1 from 61.41% to 66.37% and Recall@5 from 65.86% to 72.49%, especially improving performance under reverse traversal and long‑term appearance variation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 7

AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition

AdaptVPR introduces a route-aware generative augmentation framework that creates hard positive examples for Visual Place Recognition (VPR) training. It uses a vision‑language model to assess scene editability, a rule‑based scheduler to select generation routes, and a VPR‑oriented verification scheme to ensure geometric consistency and appearance diversity. The resulting AdaptCities dataset contains 160K verified synthetic hard positives, leading to consistent performance gains across VPR baselines, including up to 9.2% improvement in R@1 under challenging domain shifts.

By Shunpeng Chen, Jingyi Zhang, Changwei Wang, Shengpeng Xu, Yukun Song, Xingtian Pei, Jinzhou Lin, Li Guo, Shibiao Xu
arXiv Machine Learning
Aug 14

SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models

arXiv:2601. 11729v2 Announce Type: replace-cross Abstract: Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems.

By Turhan Can Kargin, Wojciech Jasi\'nski, Adam Pardyl, Bartosz Zieli\'nski, Marcin Przewi\k{e}\'zlikowski
arXiv Computer Vision
Sep 23

Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

The paper introduces TFA, a training‑free aggregation technique that calibrates frozen visual foundation models for visual place recognition. TFA uses cross‑codebook agreement, retrieval coverage, and spectral statistics to adjust residual assignment, spectral shaping, and global‑feature fusion without requiring place labels or task‑specific weights. Experiments with a DINOv2‑B backbone show significant Recall@1 gains over existing training‑free methods across multiple benchmarks, demonstrating that reliability‑guided aggregation can unlock additional retrieval performance from frozen representations.

By Xin Li, Zhimin Mao, Shang Wang, Siyuan Duan, Geng Zhang