arXiv Computer Vision

Wound3DAssist: A Practical Framework for 3D Wound Assessment

arXiv Computer Vision
2d ago

Online camera-pose-free stereo endoscopic tissue deformation recovery with tissue-invariant vision-biomechanics consistency

The paper presents a camera‑pose‑free stereo endoscopic method for recovering tissue deformation by modeling geometry as a 3D point‑derivative map and deformation as a 3D displacement‑local deformation map. It optimizes inter‑frame deformation in a camera‑centric setting, eliminating the need for camera pose estimation, and introduces a canonical map for online geometry and deformation optimization. Experiments on in‑vivo and ex‑vivo laparoscopic data show accurate 3D reconstruction (≈0.37–0.39 mm surface distance) even under occlusion, and the method can estimate surface strain distributions during manipulation.

By Jiahe Chen, Naoki Tomii, Ichiro Sakuma, Etsuko Kobayashi
arXiv Computer Vision
1d ago

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

arXiv:2609.02717v1 Announce Type: new Abstract: Large-scale training and refined optimization techniques have greatly improved sparse multi-view 3D reconstruction. Despite their relevance to surgery,...

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
arXiv Computer Vision
Aug 28

Surgical Video Generation From Diffusion to World Models: A Survey

This survey reviews recent advances in surgical video generation, categorizing methods into unconditional, conditional, and world modeling generation. It highlights a shift from creating visually plausible frames to modeling the causal dynamics of surgical scenes, and discusses challenges such as pixel-level fidelity versus clinical plausibility, generalization, physical realism, controllability, and interpretability. The paper also compiles experimental results from public datasets to serve as a quantitative benchmark for the field.

By Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang
arXiv AI
Aug 10

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

arXiv:2608. 06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions.

By Rulin Zhou, Wanhao Liu, Guoheng Ma, Liangjin Shao, Qiujie Song, Yidu Wang, Guankun Wang, Tong Chen, Long Bai, Luping Zhou, Hongliang Ren
arXiv AI
Jul 21

Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy

arXiv:2607. 17810v1 Announce Type: cross Abstract: Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision.

By Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali
arXiv AI
Aug 28

Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from ambient cameras

EgoSurg is a framework that reconstructs dynamic operating room scenes from sparse wall‑mounted stereo video and renders arbitrary, role‑specific egocentric views without instrumenting personnel. It builds a per‑timestamp 3D Gaussian Splatting representation using scale‑aware stereo depth and refines it with an image‑conditioned diffusion model to correct artifacts from limited camera coverage, crowding, and occlusion. Evaluations on real robotic pulmonology procedures and simulated sessions show consistent near‑field reconstruction fidelity (PSNR 26.8 dB, SSIM 0.895) and synthesized egocentric view quality (PSNR 17.8 dB, SSIM 0.766) across workflow phases and hospital sites, with case studies demonstrating applications such as sterile field violation adjudication, role‑specific replay, and counterfactual personnel positioning.

By Han Zhang, Lalithkumar Seenivasan, Jose L. Porras, Roger D. Soberanis-Mukul, Hao Ding, Hongchao Shu, Benjamin D. Killeen, Ankita Ghosh, Lonny Yarmus, Jeffrey K. Jopling, Masaru Ishii, Angela C. Argento, Mathias Unberath
Hugging Face Trending Papers
Jun 25

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.