arXiv Computer Vision By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya

PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation

Read the original on arXiv Computer Vision →

The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 3

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
arXiv Computer Vision
Sep 15

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...

By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
arXiv Computer Vision
Aug 31

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

SUFLECA is a weakly supervised framework that improves zero‑shot CAD‑to‑image alignment by scaling geometry‑grounded feature learning using Normalized Object Coordinates across up to 12 real and synthetic datasets. It introduces a geometrically consistent matching algorithm that reliably establishes CAD‑to‑image correspondences, enabling accurate, sub‑second alignment without iterative pose refinement. On the ScanNet25k benchmark, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero‑shot baseline by 9.7/12.5 percentage points and surpassing existing pose‑supervised methods for the first time.

By Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez