The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque models, offering transparent, quantitative performance indicators that could complement or replace traditional subjective scoring.
The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque classification tools, offering transparent, quantitative performance indicators that could complement or replace traditional scoring methods.
By Mohammad Javad Ahmadi, Hamid D. Taghirad
arXiv:2608.30872v1 Announce Type: new
Abstract: Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on...
By Chaohui Dang, Zheheng Jiang, James Glasbey, David Luke, Theodoros Arvanitis, Le Zhang
Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are time-consuming and difficult to scale. Automated feedback could significantly enhance surgical training by providing immediate, standardized metrics.
arXiv:2603.29962v4 Announce Type: replace
Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....
By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
SurgAtlas is the largest surgical video‑language dataset, containing 15,291 videos (2,391 hours) across 18 specialties and over 5,000 procedure types, all sourced from public YouTube. It uniquely includes open‑surgery videos at scale (6,182) alongside more than 9,000 minimally invasive recordings, and introduces standardized benchmarks for open‑surgery video understanding. The dataset offers a rich, multi‑tier annotation schema—segment‑level captions, step/phase descriptions, video‑level surgical narratives, and reasoning‑oriented VQA pairs—validated by experts and built through an automated LLM‑enriched pipeline.
"whyItMatters":"SurgAtlas provides an unprecedentedly large, diverse, and clinically validated resource that can train and benchmark multimodal surgical AI models, advancing the development of next‑generation foundation models for surgery."
By Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin, Ahmad M. Hider, Li Yayuan, Jing Bi, Susan Liang, Chenliang Xu, Donald S. Likosky, Jason J. Corso