arXiv Computer Vision

Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement

FreeInpaint is a feed‑forward 3D inpainting framework that reconstructs complete, geometrically consistent scenes directly from unposed multi‑view images with masked regions. It extends a 3D foundation model by propagating masked areas across views, using a Learnable Mask Attention mechanism to maintain reliable cross‑view correspondences and a Support Token Refinement strategy that injects diffusion‑generated auxiliary tokens for high‑fidelity completion. Experiments on diverse datasets show that FreeInpaint delivers superior inpainting quality without requiring pre‑computed camera poses, while maintaining fast inference speed.

arXiv Computer Vision
1d ago

PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors

The paper introduces PointVGGT, a zero‑shot framework for multiview RGB‑D point cloud registration that replaces the traditional pairwise‑then‑global pipeline. It employs a foundation‑then‑refinement paradigm, using visual geometry foundation models to directly recover metrically consistent global poses and then refining them with voxelized spatial hashing and IRLS‑based bundle adjustment. Experiments on indoor, object‑centric, and outdoor datasets demonstrate superior registration accuracy and computational efficiency without any training.

By Haobo Jiang, Liang Yu, Jianmin Zheng
arXiv Computer Vision
Sep 3

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.

By Jung-Hee Kim, Xiaoming Liu