arXiv Computer Vision

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

arXiv Computer Vision
Aug 27

PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction

PIVOT is a new multi‑trajectory dataset and evaluation framework that captures real‑world scenes with diverse camera paths, preserving both sensor‑derived measured poses and COLMAP‑optimized poses along with calibrated and optimized intrinsics. It defines three benchmark families—seen vs. unseen trajectory generalization, measured vs. optimized pose sensitivity, and calibrated vs. optimized intrinsics sensitivity—and introduces a directed pose‑space Chamfer distance to assess pose coverage. The first version of PIVOT includes five scenes recorded with a DJI Mini 4 Pro and offers an open processing and Nerfstudio‑based evaluation toolchain, revealing a consistent quality gap between held‑out and unseen trajectories and significant sensitivity to pose source and camera intrinsics.

By Mary Raymond
Hugging Face Trending Papers
Aug 11

CasDeblurGS: Cascaded 2D-to-3D Multi-View Consistency for 3D Gaussian Splatting from Two Blurry Images

Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and motion blur. Although neural rendering has advanced sparse-view synthesis, existing blur-aware methods typically require substantial multi-view redundancy, accurate camera poses, or costly per-scene optimization.

arXiv Computer Vision
2d ago

Online camera-pose-free stereo endoscopic tissue deformation recovery with tissue-invariant vision-biomechanics consistency

The paper presents a camera‑pose‑free stereo endoscopic method for recovering tissue deformation by modeling geometry as a 3D point‑derivative map and deformation as a 3D displacement‑local deformation map. It optimizes inter‑frame deformation in a camera‑centric setting, eliminating the need for camera pose estimation, and introduces a canonical map for online geometry and deformation optimization. Experiments on in‑vivo and ex‑vivo laparoscopic data show accurate 3D reconstruction (≈0.37–0.39 mm surface distance) even under occlusion, and the method can estimate surface strain distributions during manipulation.

By Jiahe Chen, Naoki Tomii, Ichiro Sakuma, Etsuko Kobayashi
arXiv Computer Vision
1d ago

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv AI
Aug 28

Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from ambient cameras

EgoSurg is a framework that reconstructs dynamic operating room scenes from sparse wall‑mounted stereo video and renders arbitrary, role‑specific egocentric views without instrumenting personnel. It builds a per‑timestamp 3D Gaussian Splatting representation using scale‑aware stereo depth and refines it with an image‑conditioned diffusion model to correct artifacts from limited camera coverage, crowding, and occlusion. Evaluations on real robotic pulmonology procedures and simulated sessions show consistent near‑field reconstruction fidelity (PSNR 26.8 dB, SSIM 0.895) and synthesized egocentric view quality (PSNR 17.8 dB, SSIM 0.766) across workflow phases and hospital sites, with case studies demonstrating applications such as sterile field violation adjudication, role‑specific replay, and counterfactual personnel positioning.

By Han Zhang, Lalithkumar Seenivasan, Jose L. Porras, Roger D. Soberanis-Mukul, Hao Ding, Hongchao Shu, Benjamin D. Killeen, Ankita Ghosh, Lonny Yarmus, Jeffrey K. Jopling, Masaru Ishii, Angela C. Argento, Mathias Unberath