arXiv AI

Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Improvement in Efficiency, Precision, and Depth Perception for Cardiac Interventions

arXiv:2608. 07606v1 Announce Type: cross Abstract: Despite advances in 3D ultrasound, most percutaneous cardiac interventions still rely on 2D visualization, limiting depth perception and spatial understanding.

arXiv AI
Jun 3

Echo-POSED: Geometric Self-Distillation for Echocardiography Guidance

arXiv:2606. 02634v1 Announce Type: cross Abstract: We introduce Echo-POSED, a self-supervised framework for real-time transthoracic echocardiography (TTE) guidance that recommends probe adjustments directly from 2D ultrasound images, without the need for expert-labelled views or tracked probe trajectories.

By Elias Stenhede, Edvart Gr\"uner Bjerke, Joanna Sulkowska, Eivind Bj{\o}rkan Orstad, Ole Jakob Elle, Ulysse C\^ot\'e-Allard, Arian Ranjbar
arXiv Computer Vision
Sep 18

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

RAUL is a reference‑assisted reconstruction framework that recovers ureteroscope trajectories from endoscopic video alone, using a high‑quality reference exploration video for each phantom. It achieves a mean translation error of 0.5 mm and increases frame‑wise localization coverage from 50.5 % to 86.1 % compared to standard Structure‑from‑Motion. The reconstructed trajectories reveal significant differences in navigation metrics between high‑ and low‑experience trainees, enabling objective skill assessment without external tracking equipment.

By Fangjie Li, Mai Bui, Charan Mohan, Michael Miga, Matthieu Chabanas, Nicholas Kavoussi, Jie Ying Wu
arXiv Computer Vision
Aug 28

DALE-CT: Depth-Aware 2D Slice Encoders Learn an Anatomical World Model of Chest CT

DALE-CT introduces depth‑aware 2D slice encoders that learn an anatomical world model of chest CT scans without 3D or positional supervision. By sampling self‑supervised views across a physical $z$‑axis slab, the encoder captures how anatomy changes between neighboring slices, enabling it to recover slice ordering and distinguish slices by anatomy alone. The model, trained on a large 287k‑scan corpus, achieves state‑of‑the‑art performance on CT‑RATE and is released with full code and evaluation tools.

By Evan W. Damron, Mahmut S. Gokmen, Mitchell A. Klusty, Caroline N. Leach, Emily B. Collier, V. K. Cody Bumgardner
arXiv Computer Vision
Sep 3

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
arXiv AI
Jun 12

Transformer-Guided Graph Attention for Direct Cardiac Mesh Reconstruction: A Structural Digital Twin Framework

arXiv:2606. 13188v1 Announce Type: cross Abstract: Building patient-specific cardiac models sits at the heart of precision cardiology, yet getting those models into clinical use keeps running into the same wall: mesh generation is slow, messy, and frustrating.

By Abhishek H S, Akash Ganamukhi, Abhimanyu Suresh, Aditya G Hiremath, Prasad B Honnavalli, Adithya Balasubramanyam
arXiv Machine Learning
Aug 5

Dual-domain U-Nets with embedded back projection operators for motion-resolved 4D CBCT reconstruction

arXiv:2608. 03430v1 Announce Type: cross Abstract: Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times, causing high patient dose and motion/sparse-sampling artifacts.

By Ivo Herzig, Pascal Paysan, Daniel Barco, Marc Andr\'e Stadelmann, Frank-Peter Schilling, Igor Peterlik, Michal Walczak, Lijin Aryananda, Woo Sang Ahn, Rudolf Marcel F\"uchslin, Lukas Lichtensteiger
arXiv Machine Learning
Jun 29

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

arXiv:2606. 28104v1 Announce Type: cross Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution.

By Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum
arXiv AI
Aug 19

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

The paper introduces LiftXR, a geometry‑guided framework that first reconstructs a 3D anatomical layout from bi‑planar X‑ray images and then uses this layout to guide CT volume reconstruction. An anatomical parser refines the layout by analyzing the reconstructed CT, enabling region‑specific intensity calibration. Experiments on two public datasets show LiftXR surpasses recent X‑ray‑to‑CT methods and improves downstream segmentation performance.

By Yifei Wu, Yicheng Wu, Qiang Ma, Qi Chen, Renyang Gu, Xinyu Liu, Yongsheng Pan, Yong Xia