arXiv Computer Vision

PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors

The paper introduces PointVGGT, a zero‑shot framework for multiview RGB‑D point cloud registration that replaces the traditional pairwise‑then‑global pipeline. It employs a foundation‑then‑refinement paradigm, using visual geometry foundation models to directly recover metrically consistent global poses and then refining them with voxelized spatial hashing and IRLS‑based bundle adjustment. Experiments on indoor, object‑centric, and outdoor datasets demonstrate superior registration accuracy and computational efficiency without any training.

arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv Computer Vision
1d ago

Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement

FreeInpaint is a feed‑forward 3D inpainting framework that reconstructs complete, geometrically consistent scenes directly from unposed multi‑view images with masked regions. It extends a 3D foundation model by propagating masked areas across views, using a Learnable Mask Attention mechanism to maintain reliable cross‑view correspondences and a Support Token Refinement strategy that injects diffusion‑generated auxiliary tokens for high‑fidelity completion. Experiments on diverse datasets show that FreeInpaint delivers superior inpainting quality without requiring pre‑computed camera poses, while maintaining fast inference speed.

By Jingyi Pan, Dan Xu, Qiong Luo
arXiv Computer Vision
5d ago

SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

SymRegFlow is a symmetry‑regularized flow‑matching framework that enables multi‑view‑consistent video generation across continuously varying camera poses without requiring ground‑truth novel‑view RGB supervision. The method geometrically warps source views into noisy anchors and uses masked dual‑anchor supervision combined with cross‑anchor denoising‑output consistency to reduce anchor‑specific errors. Experiments on Cosmos‑Drive‑Dreams and nuScenes show that SymRegFlow achieves superior video quality, achieving the lowest FVD and FVMD scores and improving FID and instance preservation compared to existing baselines.

By Xi Ye, Yuzhu Wang, Xiaoyang Liu, Jiayi Wang, Yangyang Xu, Ruyu Wang, Wenlin Chen, Duo Su, Jun Zhu
arXiv Computer Vision
3d ago

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

Con-DSO introduces a consistency-aware RGB‑D direct sparse odometry framework that learns pixel‑level photometric and geometric uncertainty from adjacent RGB‑D frame pairs. The network predicts uncertainties that are converted into pairwise quality scores, guiding support‑pixel selection and forming a host‑side quality prior for keyframe tracking. Experiments on five public benchmarks show that this approach reduces absolute trajectory error by over 20% on ICL‑NUIM and by 50–80% on other datasets, improving robustness in challenging environments.

By Haolan Zhang, Thanh Nguyen Canh, Chenghao Li, Ziyan Gao, Xiongwen Jiang, Nak Young Chong
arXiv Computer Vision
Sep 17

Category Level 6D Object Pose Estimation from a Single RGB Image using Diffusion

The paper presents a generative framework that estimates category-level 6D pose and 3D size of objects from a single RGB image, using score-based diffusion models to produce a multi-hypothesis pose distribution. It replaces costly likelihood pruning with a Mean Shift approach to isolate the mode as the final pose estimate, achieving state-of-the-art results on the REAL275 benchmark. The method also decouples detection from pose estimation, enabling robust zero-shot generalisation on the Wild6D dataset and extending naturally to video sequences by propagating the pose distribution over time.

By Adam Bethell, Ravi Garg, Ian Reid
arXiv Computer Vision
Aug 31

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

SUFLECA is a weakly supervised framework that improves zero‑shot CAD‑to‑image alignment by scaling geometry‑grounded feature learning using Normalized Object Coordinates across up to 12 real and synthetic datasets. It introduces a geometrically consistent matching algorithm that reliably establishes CAD‑to‑image correspondences, enabling accurate, sub‑second alignment without iterative pose refinement. On the ScanNet25k benchmark, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero‑shot baseline by 9.7/12.5 percentage points and surpassing existing pose‑supervised methods for the first time.

By Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez
Hugging Face Trending Papers
Jun 29

Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes

Metric feed-forward 3D reconstruction for panoramic data remains under-explored due to the lack of large-scale panoramic RGB-D training data. We present Realsee3D, a hybrid dataset of 10K indoor scenes (1K real, 9K synthetic) with 299K panoramic viewpoints and precise metric annotations, and Argus, a feed-forward network trained on it for metric panoramic 3D reconstruction.