arXiv Computer Vision

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

RAUL is a reference‑assisted reconstruction framework that recovers ureteroscope trajectories from endoscopic video alone, using a high‑quality reference exploration video for each phantom. It achieves a mean translation error of 0.5 mm and increases frame‑wise localization coverage from 50.5 % to 86.1 % compared to standard Structure‑from‑Motion. The reconstructed trajectories reveal significant differences in navigation metrics between high‑ and low‑experience trainees, enabling objective skill assessment without external tracking equipment.

arXiv Computer Vision
Sep 22

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.

By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv Computer Vision
Sep 17

Video-Based Markerless Motion Capture for Clinical and Rehabilitation Biomechanics: A PRISMA-ScR Scoping Review of Validated Architectures, Clinical Readiness, and Emerging Methods

This scoping review examined 117 studies on video-based markerless motion capture, most published from 2024 onward and focused on healthy adults walking in laboratories. The studies identified five main pipeline architectures, but most reported only raw joint angles without biomechanical refinement, achieving sagittal lower‑limb agreement of about 5–6°, which falls short of clinical acceptability. Validation of out‑of‑plane kinematics, kinetics, and performance in older or pathological populations was rare, and emerging computer‑vision techniques such as foundation‑model mesh recovery and differentiable inverse kinematics were largely absent from validated work.

By Florian Delaplace (LAMHESS, CHU), Elodie Piche (LAMHESS), Fr\'ed\'eric Chorin (IUF, LAMHESS), Raphael Zory (IUF, LAMHESS)
arXiv Computer Vision
Aug 25

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.

By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
arXiv Computer Vision
Sep 3

MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion

MuyBridge is an on-device system that estimates an athlete’s segmental center of mass (CoM) trajectory from a single phone camera video stream. It combines a compact 2D pose network with a distilled monocular depth network, fusing their outputs through anatomical and physical priors to produce metric CoM estimates without requiring 3D or task‑specific supervision. On the AthletePose3D dataset, MuyBridge achieves 33–41 mm vertical CoM error and 2.3–6.6 % absolute‑relative range error, delivering CoM estimates at 63 FPS on an iPhone 15 with asynchronous depth updates.

By Aidan Bradshaw, Marco Giordano, David Rode, Andreas Habersack, Elif Basokur, Annika Kruse, Markus Tilp, Michele Magno, Peter Wolf, Luca Benini, Christoph Leitner
arXiv AI
Aug 11

Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Improvement in Efficiency, Precision, and Depth Perception for Cardiac Interventions

arXiv:2608. 07606v1 Announce Type: cross Abstract: Despite advances in 3D ultrasound, most percutaneous cardiac interventions still rely on 2D visualization, limiting depth perception and spatial understanding.

By Mohsen Annabestani, Sandhya Sriram, Andrew Kuzemczak, S. Chiu Wong, Alexandros Sigaras, Bobak Mosadegh
arXiv Computer Vision
Aug 26

C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization

C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.

By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque