arXiv Computer Vision

Multi-View Structure-from-Motion Enables Oriented Projective Shape Analysis in Three Dimensions

arXiv Computer Vision
Sep 3

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
arXiv Computer Vision
Sep 15

G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity

The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that remains consistent across different camera projections. By parameterizing rotary phases with camera-local ray angles, G-ray achieves projection-invariant positional consistency and can be integrated with existing encodings without extra learned parameters. Experiments on 3D reconstruction and novel-view synthesis benchmarks show that G-ray improves performance, notably reducing mean pointmap relative error by 45.8% over MapAnything.

By Shuo Zhang, Xin Su, Wei Wang, Jun Liu, Xinrui Zeng, Yongsen Chen, Chenjie Wang, Guibo Zhu, Jinqiao Wang, Bin Luo, Liangpei Zhang
Google AI Blog
Mar 18, 2024

MELON: Reconstructing 3D objects from images with unknown poses

Posted by Mark Matthews, Senior Software Engineer, and Dmitry Lagun, Research Scientist, Google Research A person's prior experience and understanding of the world generally enables them to easily infer what an object looks like in whole, even if only looking at a few 2D pictures of it. Yet the capacity for a computer to reconstruct the shape of an object in 3D given only a few images has remained a difficult algorithmic problem for years.

By Google AI
arXiv Machine Learning
Jul 8

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

arXiv:2601. 15275v3 Announce Type: replace-cross Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene.

By Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani
arXiv Computer Vision
Sep 4

Stable and Scalable Bundle Adjustment of Holistic 3D Structures

The paper introduces a unified bundle adjustment framework that jointly optimizes camera parameters, sparse 3D points, and richer geometric features such as lines, coplanarity, and parallelism. It classifies features into scalable ones with direct 2D measurements and higher‑order groups that can be treated as camera‑like entities, allowing group constraints and cross‑feature relations to be expressed via 2D reprojection errors. This approach preserves the sparsity of classical point‑based BA, maintains numerical stability, and achieves runtime comparable to point‑only BA while producing richer 3D structures and improved accuracy.

By Shaohui Liu, R\'emi Pautrat, Daniel Barath, Richard Hartley, Viktor Larsson, Marc Pollefeys
arXiv AI
Sep 7

Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models

The paper introduces a CAD‑free 3D shape prior that enhances object recognition by reconstructing each object with 3D Gaussian Splatting (3DGS) from short RGB‑D scans and fusing the resulting shape prototype with frozen DINOv2 image features. Experiments on T‑LESS and HOPE datasets show that geometry alone can match or exceed CAD‑based recognition, and that the combined approach improves performance, especially on shape‑distinctive or partially occluded objects. The study demonstrates that the benefit comes from the geometric information rather than rendered pixels, and that the prior is complementary to frozen vision features.

By Chenxi Tao, Seung-Kyum Choi
arXiv Computer Vision
Sep 3

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

The paper introduces Self-Geometry, a plug‑and‑play test‑time adaptation framework that enforces explicit multi‑view geometric constraints on Vision Foundation Models (VFMs) using 2D pixel correspondences as pseudo ground truth. It combines Geometric Disentanglement Optimization—mixing Multi‑View and Epipolar Consistency losses with Gradient Disentanglement—to avoid gradient conflicts, a Frame Angular‑Neighbor sampler based on SO(3) geodesic distances to select informative views, and a Lightweight TTA module that adapts VFMs via LoRA. Experiments on six VFMs and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom) show consistent improvements in pose and geometry estimation.

By Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
arXiv Computer Vision
Sep 17

CADSplat: Sparse-View 3D Gaussian Splatting Aided by CAD Models for Robust, Photorealistic Digital-Twin Reconstruction

CADSplat is a framework that reconstructs photorealistic, geometrically accurate digital twins from fewer than 15 wide‑baseline images by regularizing 3D Gaussian Splatting with an explicit CAD shape prior. It matches segmented object silhouettes to a CAD library to retrieve a suitable model and camera poses, then anchors Gaussian primitives to the model’s surface and jointly optimizes splat parameters, registration, and a non‑rigid deformation field. Experiments on two real‑world datasets show CADSplat outperforms baselines, especially in sparse and self‑occluded scenarios, and its gains mainly stem from constraining splats to a surface rather than the CAD shape itself.

By Kristof Overdulve, Lode Jorissen, Nick Michiels
arXiv Computer Vision
Sep 2

Online camera-pose-free stereo endoscopic tissue deformation recovery with tissue-invariant vision-biomechanics consistency

The paper presents a camera‑pose‑free stereo endoscopic method for recovering tissue deformation by modeling geometry as a 3D point‑derivative map and deformation as a 3D displacement‑local deformation map. It optimizes inter‑frame deformation in a camera‑centric setting, eliminating the need for camera pose estimation, and introduces a canonical map for online geometry and deformation optimization. Experiments on in‑vivo and ex‑vivo laparoscopic data show accurate 3D reconstruction (≈0.37–0.39 mm surface distance) even under occlusion, and the method can estimate surface strain distributions during manipulation.

By Jiahe Chen, Naoki Tomii, Ichiro Sakuma, Etsuko Kobayashi