arXiv AI By Jingxian Xu, Yuhao Huang, Rusi Chen, Yanfeng Zhou, Dong Ni

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

Read the original on arXiv AI →

arXiv:2608. 09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 24

Recursive Uncertainty-Gated Image Registration for Learning-based Algorithms

The paper introduces Recursive Uncertainty-Gated Image Registration (RUGI), an iterative refinement method that updates deformation fields predicted by learning-based registration models using a gating map. Two gating strategies are explored: an uncertainty-based approach and an image residual error approach, both concentrating updates on difficult regions. Experiments on cardiac MRI and echocardiography datasets show that RUGI consistently improves registration accuracy, with the error-gated variant reducing MSE by 27‑37% on pretrained models and lowering ejection fraction estimation errors.

By Clara Rodrigo Gonz\'alez, Oscar Bates, Fu Siong Ng, Meng-Xing Tang
arXiv Computer Vision
Sep 18

UniReg: Conditional Unified Model for Medical Image Registration

UniReg is a conditional unified model for medical image registration that adapts deformation field estimation based on anatomical priors, registration type constraints, and instance-specific features. It combines the precision of task‑specific learning with the generalization of traditional optimization, enabling effective alignment across diverse CT and MR scenarios within a single framework. Experiments show UniReg outperforms state‑of‑the‑art learning‑based methods in accuracy while providing strong cross‑scenario generalization and reducing training cost and model redundancy.

By Zi Li, Jianpeng Zhang, Tai Ma, Tony C. W. Mok, Yan-Jie Zhou, Zeli Chen, Xianghua Ye, Le Lu, Cheng Chen, Dakai Jin
arXiv Computer Vision
Aug 25

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.

By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
arXiv AI
Jun 17

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.

By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath