Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs).
arXiv:2606.03406v2 Announce Type: replace
Abstract: Reliable correspondence estimation supports image processing and 3D vision tasks, including Structure from Motion, visual localization, and image r...
By Xu Pan, Zhen Pang, Qiyuan Ma, Wei Ji, Shuhan Shen, Xianwei Zheng
SOCO is a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. It also includes keypoint language descriptions, enabling evaluation of large vision‑language models and their fine‑grained part‑level understanding. Experiments show that vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories, LVLMs excel at text‑prompted part localization but lag in visual‑reference matching, and correspondence performance predicts dense downstream tasks more strongly than ImageNet classification.
By Olaf D\"unkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann, Christian Theobalt, Adam Kortylewski
Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene...
arXiv:2609.39748v1 Announce Type: new
Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
By Yongjian Zhang, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, Yulan Guo
arXiv:2512.15708v2 Announce Type: replace
Abstract: Foundation models are vital tools in various Computer Vision applications. They take as input a single RGB image and output a deep feature represen...
By Leo Segre, Or Hirschorn, Shai Avidan
arXiv:2608.23850v1 Announce Type: new
Abstract: Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-v...
By Jeong-gi Kwak, Sho Kagami, Yuki Ono, Kwang Moo Yi
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
By Prakhar Bhardwaj, Simone Weikl, Kilian Mang, Elia Jonas Sandtner
arXiv:2610.01098v1 Announce Type: new
Abstract: Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D recons...
By Hanyuan Xiao, Gonglin Chen, Haolin Xiong, Wenbin Teng, Haiwei Chen, Yajie Zhao
WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.
By Jerrin Bright, John Zelek
The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.
By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
Two-view correspondence learning aims to distinguish true correspondences (inliers) from false ones (outliers) in image pairs by leveraging their underlying differences. Existing methods mainly rely on coordinate-based geometric consistency.