DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests. M3GA-Wild unifies and extends existing forest localisation datasets, providing a...
M3GA-Wild is a new benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests, combining synchronized RGB imagery and LiDAR from ground traversals with high‑resolution aerial imagery and multi‑altitude LiDAR over 370 hectares. The dataset includes accurate geo‑referenced 6‑DoF poses and spans 36 km of forest traversals, enabling systematic evaluation of visual, LiDAR, cross‑modal, and multi‑modal methods. Baseline experiments show LiDAR outperforms vision‑only approaches under severe viewpoint changes, while current multi‑modal fusion offers limited gains due to poor cross‑modal alignment, highlighting challenges in cross‑platform localisation and domain gaps.
By Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani
The paper introduces DPA-I2P, a depth-guided projective alignment method for image-to-point-cloud registration in autonomous driving. It employs Ray-Conditioned Metric Depth Encoding and Projection-Consistent Vision Lifting to align depth and visual cues geometrically, and uses Cross-Modal Query Pruning to enhance matching stability. Experiments on KITTI and nuScenes show significant reductions in rotation and translation errors compared to existing implicit baselines.
By Wenxin Zhang, Hang Li, Zhiwei Xu, Qiankun Dong, Gang Wang, Tao Li
The paper introduces DPA-I2P, a depth‑guided projective alignment method for image‑to‑point‑cloud registration in autonomous driving. It employs Ray‑Conditioned Metric Depth Encoding and Projection‑Consistent Vision Lifting to align depth and visual cues geometrically, and uses Cross‑Modal Query Pruning to filter unreliable matches during refinement. Experiments on KITTI and nuScenes show significant improvements, reducing rotation and translation errors by up to 55.6% compared to existing implicit baselines.
arXiv:2609.21522v1 Announce Type: new
Abstract: Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representation...
By Hang Cheng, Yan Chen, Mingyu Fan, Long Zeng