The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque classification tools, offering transparent, quantitative performance indicators that could complement or replace traditional scoring methods.
By Mohammad Javad Ahmadi, Hamid D. Taghirad
arXiv:2608. 15580v1 Announce Type: new Abstract: Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record.
By Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang
arXiv:2608.30872v1 Announce Type: new
Abstract: Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on...
By Chaohui Dang, Zheheng Jiang, James Glasbey, David Luke, Theodoros Arvanitis, Le Zhang
arXiv:2609.36407v1 Announce Type: new
Abstract: Cross-resolution knowledge distillation aims to improve low-magnification whole- slide analysis by transferring high-magnification representations, yet...
By Zhiyuan Yang, Jiahao Cheng, Mahdi S. Hosseini
The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque models, offering transparent, quantitative performance indicators that could complement or replace traditional subjective scoring.
arXiv:2609.18971v1 Announce Type: new
Abstract: Surgical phase recognition maps each video frame to a clinically meaningful workflow phase, supporting context-aware assistance, documentation, and pos...
By Ye Tao, Claudia Scherl, Sara Monji-Azad
Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are time-consuming and difficult to scale. Automated feedback could significantly enhance surgical training by providing immediate, standardized metrics.
arXiv:2609.38362v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining dis...
By Hung-Jen Chen, Yu-Heng Ho, Ting-Yao Huang, Po-Hsiang Hsu, Li-Yu Chen, Chun-Yi Lee, Min Sun
arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.
By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv:2603.29962v4 Announce Type: replace
Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....
By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv:2605. 23995v4 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data.
By Chathura Wimalasiri, Kishor Nandakishor, Marimuthu Palaniswami