arXiv Computer Vision
Sep 3

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
arXiv Computer Vision
Sep 15

G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity

The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that remains consistent across different camera projections. By parameterizing rotary phases with camera-local ray angles, G-ray achieves projection-invariant positional consistency and can be integrated with existing encodings without extra learned parameters. Experiments on 3D reconstruction and novel-view synthesis benchmarks show that G-ray improves performance, notably reducing mean pointmap relative error by 45.8% over MapAnything.

By Shuo Zhang, Xin Su, Wei Wang, Jun Liu, Xinrui Zeng, Yongsen Chen, Chenjie Wang, Guibo Zhu, Jinqiao Wang, Bin Luo, Liangpei Zhang
Google AI Blog
Mar 18, 2024

MELON: Reconstructing 3D objects from images with unknown poses

Posted by Mark Matthews, Senior Software Engineer, and Dmitry Lagun, Research Scientist, Google Research A person's prior experience and understanding of the world generally enables them to easily infer what an object looks like in whole, even if only looking at a few 2D pictures of it. Yet the capacity for a computer to reconstruct the shape of an object in 3D given only a few images has remained a difficult algorithmic problem for years.

By Google AI
arXiv Machine Learning
Jul 8

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

arXiv:2601. 15275v3 Announce Type: replace-cross Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene.

By Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani
arXiv Computer Vision
Sep 4

Stable and Scalable Bundle Adjustment of Holistic 3D Structures

The paper introduces a unified bundle adjustment framework that jointly optimizes camera parameters, sparse 3D points, and richer geometric features such as lines, coplanarity, and parallelism. It classifies features into scalable ones with direct 2D measurements and higher‑order groups that can be treated as camera‑like entities, allowing group constraints and cross‑feature relations to be expressed via 2D reprojection errors. This approach preserves the sparsity of classical point‑based BA, maintains numerical stability, and achieves runtime comparable to point‑only BA while producing richer 3D structures and improved accuracy.

By Shaohui Liu, R\'emi Pautrat, Daniel Barath, Richard Hartley, Viktor Larsson, Marc Pollefeys