The paper introduces Mask 2D-3D, an Adaptive Dual-Masked Autoencoder Network designed for image-to-point cloud registration. It proposes an Intermodal Dual-MAE Framework (ID-MAE) with a Similarity-based RL Masking Strategy (SRLM) that adaptively masks informative positions using cross-modal similarity and reinforcement learning. Experiments on RGB-D Scenes v2 and 7-Scenes benchmarks demonstrate state-of-the-art performance in this registration task.
By Zhixin Cheng, Jiacheng Deng, Xiaotian Yin, Baoqun Yin, Richang Hong, Tianzhu Zhang
DMM-Align introduces a closed‑loop framework for 2D‑3D registration that jointly refines correspondences, estimates pose, and learns representations using a shared differentiable geometric state. The method employs two diffusion processes: a geometry‑aware diffusion that improves the soft matching matrix for robust correspondence estimation, and a geometry‑conditioned diffusion teacher that feeds pose‑induced supervision back into feature learning. Experiments on 7‑Scenes and RGB‑D Scenes V2 show that DMM‑Align outperforms strong baselines, particularly in low‑overlap and heavily occluded scenarios, demonstrating the value of closed‑loop geometric feedback.
By Chongjian Wang, Junjie Gao
The paper introduces DPA-I2P, a depth-guided projective alignment method for image-to-point-cloud registration in autonomous driving. It employs Ray-Conditioned Metric Depth Encoding and Projection-Consistent Vision Lifting to align depth and visual cues geometrically, and uses Cross-Modal Query Pruning to enhance matching stability. Experiments on KITTI and nuScenes show significant reductions in rotation and translation errors compared to existing implicit baselines.
By Wenxin Zhang, Hang Li, Zhiwei Xu, Qiankun Dong, Gang Wang, Tao Li
The paper introduces DPA-I2P, a depth‑guided projective alignment method for image‑to‑point‑cloud registration in autonomous driving. It employs Ray‑Conditioned Metric Depth Encoding and Projection‑Consistent Vision Lifting to align depth and visual cues geometrically, and uses Cross‑Modal Query Pruning to filter unreliable matches during refinement. Experiments on KITTI and nuScenes show significant improvements, reducing rotation and translation errors by up to 55.6% compared to existing implicit baselines.
GRIP is a pose‑conditioned refinement framework that improves pixel‑to‑point matching for image‑to‑point‑cloud registration. It mitigates the mismatch between grid‑based image descriptors and unordered point cloud descriptors by softly rendering learned 3D point features onto the image grid using Gaussian feature splatting. The resulting rendered point‑derived feature map is fused with image features via a pixel‑aligned transformer, enabling visual semantic and geometric cues to interact in a shared 2D representation, which is then decoded and propagated to finer resolutions for dense correspondence estimation and final pose refinement. Experiments on RGB‑D Scenes V2 and 7 Scenes show state‑of‑the‑art inlier ratios and competitive registration recall, especially under stricter evaluation thresholds.
By Karim Slimani, Catherine Achard, Eric Marchand, Brahim Tamadazte
arXiv:2607. 03612v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success.
By Jianing Deng, Yuanzhe Li, Jialu Wang, Song Wang, Tianlong Chen, Huanrui Yang, Jingtong Hu
Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations.
arXiv:2606. 27818v1 Announce Type: cross Abstract: We present MMD-Reg, a novel correspondence-free approach to point-cloud registration that is differentiable and has linear computational complexity in the number of points.
By Rixon Crane, Fahira Afzal Maken, Nicholas Lawrance, Stanislav Funiak, Kasra Khosoussi, Ming Xu, Russell Tsuchida
arXiv:2608. 19536v1 Announce Type: cross Abstract: Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics.
By Eunsoo Im, Junghun Suh, Gyeonggwan Lee, Seunghwan Hong
arXiv:2512. 16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving.
By Sicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu, Fang Li, Shengyin Jiang, Long Chen, Zhi-Xin Yang, Jiwen Lu
arXiv:2604. 00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks.
By Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee
arXiv:2608.30279v1 Announce Type: new
Abstract: Point cloud video representation learning is crucial for 3D dynamic scene understanding. In this paper, we propose MoSaiC, a novel Motion-Saliency Comp...
By Wei Wang, Yiding Sun, Yuyan Wang, Zhuoyue Zhang, Zhengqiao Li, Dongfu Yin, Chen Li