arXiv Computer Vision By Haobo Jiang, Liang Yu, Jianmin Zheng

PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors

Read the original on arXiv Computer Vision →

The paper introduces PointVGGT, a zero‑shot framework for multiview RGB‑D point cloud registration that replaces the traditional pairwise‑then‑global pipeline. It employs a foundation‑then‑refinement paradigm, using visual geometry foundation models to directly recover metrically consistent global poses and then refining them with voxelized spatial hashing and IRLS‑based bundle adjustment. Experiments on indoor, object‑centric, and outdoor datasets demonstrate superior registration accuracy and computational efficiency without any training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv Computer Vision
1d ago

Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement

FreeInpaint is a feed‑forward 3D inpainting framework that reconstructs complete, geometrically consistent scenes directly from unposed multi‑view images with masked regions. It extends a 3D foundation model by propagating masked areas across views, using a Learnable Mask Attention mechanism to maintain reliable cross‑view correspondences and a Support Token Refinement strategy that injects diffusion‑generated auxiliary tokens for high‑fidelity completion. Experiments on diverse datasets show that FreeInpaint delivers superior inpainting quality without requiring pre‑computed camera poses, while maintaining fast inference speed.

By Jingyi Pan, Dan Xu, Qiong Luo
arXiv Computer Vision
5d ago

SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

SymRegFlow is a symmetry‑regularized flow‑matching framework that enables multi‑view‑consistent video generation across continuously varying camera poses without requiring ground‑truth novel‑view RGB supervision. The method geometrically warps source views into noisy anchors and uses masked dual‑anchor supervision combined with cross‑anchor denoising‑output consistency to reduce anchor‑specific errors. Experiments on Cosmos‑Drive‑Dreams and nuScenes show that SymRegFlow achieves superior video quality, achieving the lowest FVD and FVMD scores and improving FID and instance preservation compared to existing baselines.

By Xi Ye, Yuzhu Wang, Xiaoyang Liu, Jiayi Wang, Yangyang Xu, Ruyu Wang, Wenlin Chen, Duo Su, Jun Zhu