Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are time-consuming and difficult to scale. Automated feedback could significantly enhance surgical training by providing immediate, standardized metrics.
arXiv:2608. 07116v1 Announce Type: cross Abstract: Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limited training data.
By Lumin Chen, Qingyao Tian, Jinpeng Li, Haoyu Jiang, Huai Liao, Xinyan Huang, Hongbin Liu, Dong Yi
arXiv:2609.27227v1 Announce Type: new
Abstract: Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic...
By Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
This scoping review examined 117 studies on video-based markerless motion capture, most published from 2024 onward and focused on healthy adults walking in laboratories. The studies identified five main pipeline architectures, but most reported only raw joint angles without biomechanical refinement, achieving sagittal lower‑limb agreement of about 5–6°, which falls short of clinical acceptability. Validation of out‑of‑plane kinematics, kinetics, and performance in older or pathological populations was rare, and emerging computer‑vision techniques such as foundation‑model mesh recovery and differentiable inverse kinematics were largely absent from validated work.
By Florian Delaplace (LAMHESS, CHU), Elodie Piche (LAMHESS), Fr\'ed\'eric Chorin (IUF, LAMHESS), Raphael Zory (IUF, LAMHESS)
The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.
By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
MuyBridge is an on-device system that estimates an athlete’s segmental center of mass (CoM) trajectory from a single phone camera video stream. It combines a compact 2D pose network with a distilled monocular depth network, fusing their outputs through anatomical and physical priors to produce metric CoM estimates without requiring 3D or task‑specific supervision. On the AthletePose3D dataset, MuyBridge achieves 33–41 mm vertical CoM error and 2.3–6.6 % absolute‑relative range error, delivering CoM estimates at 63 FPS on an iPhone 15 with asynchronous depth updates.
By Aidan Bradshaw, Marco Giordano, David Rode, Andreas Habersack, Elif Basokur, Annika Kruse, Markus Tilp, Michele Magno, Peter Wolf, Luca Benini, Christoph Leitner
Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument...
arXiv:2606. 17615v1 Announce Type: cross Abstract: Estimating human proficiency from video is a key challenge for automated skill assessment, with applications in sports coaching, music pedagogy, surgical training, and workplace learning.
By Edoardo Bianchi, Antonio Liotta
arXiv:2609.22619v1 Announce Type: new
Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains co...
By Nethmi Jayasinghe, Mihir Parashar, Amit Ranjan Trivedi
arXiv:2608. 07606v1 Announce Type: cross Abstract: Despite advances in 3D ultrasound, most percutaneous cardiac interventions still rely on 2D visualization, limiting depth perception and spatial understanding.
By Mohsen Annabestani, Sandhya Sriram, Andrew Kuzemczak, S. Chiu Wong, Alexandros Sigaras, Bobak Mosadegh
C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.
By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque