The paper introduces DPA-I2P, a depth-guided projective alignment method for image-to-point-cloud registration in autonomous driving. It employs Ray-Conditioned Metric Depth Encoding and Projection-Consistent Vision Lifting to align depth and visual cues geometrically, and uses Cross-Modal Query Pruning to enhance matching stability. Experiments on KITTI and nuScenes show significant reductions in rotation and translation errors compared to existing implicit baselines.
By Wenxin Zhang, Hang Li, Zhiwei Xu, Qiankun Dong, Gang Wang, Tao Li
The paper introduces DPA-I2P, a depth‑guided projective alignment method for image‑to‑point‑cloud registration in autonomous driving. It employs Ray‑Conditioned Metric Depth Encoding and Projection‑Consistent Vision Lifting to align depth and visual cues geometrically, and uses Cross‑Modal Query Pruning to filter unreliable matches during refinement. Experiments on KITTI and nuScenes show significant improvements, reducing rotation and translation errors by up to 55.6% compared to existing implicit baselines.
DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.
Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations.
arXiv:2609.22716v1 Announce Type: new
Abstract: Image-to-LiDAR registration estimates the camera pose of an image with respect to a LiDAR point cloud. It has diverse applications in autonomous drivin...
By Zijun Li, Xiaotian Sun, Xuelun Shen, Yao Dai, Sheng Ao, Yangyang Shi, Jakob Engel, Zhipeng Cai, Cheng Wang
arXiv:2512.18991v3 Announce Type: replace-cross
Abstract: Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds...
By Gyeongrok Oh, Youngdong Jang, Jonghyun Choi, Suk-Ju Kang, Guang Lin, Sangpil Kim
SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.
By Yuting Zhao, Ziyi Zheng, Shuxiao Li
arXiv:2608. 19536v1 Announce Type: cross Abstract: Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics.
By Eunsoo Im, Junghun Suh, Gyeonggwan Lee, Seunghwan Hong
The paper introduces TempLoc, a Temporal‑aware Localization framework that improves outdoor LiDAR relocalization by leveraging spatio‑temporal consistency across scans. It first predicts point‑wise global coordinates with uncertainties, then estimates inter‑frame correspondences using an attention‑based Prior Coordinate Generation module, and finally fuses these predictions in an uncertainty‑guided manner to produce a more accurate global 6‑DoF pose. Experiments on the NCLT and Oxford RobotCar datasets show that TempLoc significantly outperforms existing state‑of‑the‑art methods.
By Minghang Zhu, Zhijing Wang, Yuxin Guo, Chen Liu, Yongshu Huang, Wen Li, Sheng Ao, Cheng Wang
arXiv:2609.25375v1 Announce Type: cross
Abstract: Global point-cloud registration remains challenging when limited overlap, repetitive geometry, and sensor noise produce correspondence sets dominated...
By Abolfazl Babanazari, Carson Cramer, Tyler Summers, Carlos Nieto, Kaveh Fathian
Cross-modal place recognition (CMPR) aims to identify the same location across heterogeneous sensing modalities, such as vision and LiDAR. Existing methods commonly bridge the modality gap using complex alignment modules, multi-stage training, or full fine-tuning of pretrained backbones.
arXiv:2607. 26645v1 Announce Type: cross Abstract: Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion.
By Wenzhe He, Meng Wang, JiaWei Qian, Jinfeng Xu, Ying Liu, Ruihui Li