arXiv AI

AI for Quality Assurance in the Operating Room

arXiv:2606. 30657v1 Announce Type: cross Abstract: Surgical outcomes depend not only on patient factors and postoperative care but are also strongly influenced by the quality of the operation itself.

Hugging Face Trending Papers
Aug 18

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque models, offering transparent, quantitative performance indicators that could complement or replace traditional subjective scoring.

arXiv AI
Aug 19

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque classification tools, offering transparent, quantitative performance indicators that could complement or replace traditional scoring methods.

By Mohammad Javad Ahmadi, Hamid D. Taghirad
arXiv Computer Vision
Aug 25

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

arXiv:2603.29962v4 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....

By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv Computer Vision
Sep 4

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

SurgAtlas is the largest surgical video‑language dataset, containing 15,291 videos (2,391 hours) across 18 specialties and over 5,000 procedure types, all sourced from public YouTube. It uniquely includes open‑surgery videos at scale (6,182) alongside more than 9,000 minimally invasive recordings, and introduces standardized benchmarks for open‑surgery video understanding. The dataset offers a rich, multi‑tier annotation schema—segment‑level captions, step/phase descriptions, video‑level surgical narratives, and reasoning‑oriented VQA pairs—validated by experts and built through an automated LLM‑enriched pipeline. "whyItMatters":"SurgAtlas provides an unprecedentedly large, diverse, and clinically validated resource that can train and benchmark multimodal surgical AI models, advancing the development of next‑generation foundation models for surgery."

By Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin, Ahmad M. Hider, Li Yayuan, Jing Bi, Susan Liang, Chenliang Xu, Donald S. Likosky, Jason J. Corso
arXiv Computer Vision
Sep 18

State-Change Learning for Prediction of Future Events in Endoscopic Videos

The paper introduces SurgFUTR, a state‑change learning framework for predicting future events in endoscopic videos. Instead of forecasting raw observations, it classifies transitions between current and future states using a teacher‑student architecture and an Action Dynamics module. The authors also present SFPBench, a benchmark with five short‑ and long‑term prediction tasks, and demonstrate consistent improvements across multiple datasets and procedures, including cross‑procedure transfer.

By Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy
arXiv Machine Learning
Sep 11

Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge

The FedSurg Challenge is the first international effort to evaluate Federated Learning (FL) for surgical vision, using a multi‑center dataset of laparoscopic appendectomies. Three participant models were tested for generalization to an unseen clinical center and for center‑specific adaptation, compared against centralized, Swarm Learning, and parameter‑efficient fine‑tuning baselines. The study found that temporal modeling most consistently improves generalization, but overall performance remains low (26.31% F1‑score on the unseen center), highlighting the need for structured personalized FL and revealing limitations of current approaches.

By Max Kirchner, Hanna Hoffmann, Alexander C. Jenke, Oliver L. Saldanha, Kevin Pfeiffer, Weam Kanjo, Julia Alekseenko, Claas de Boer, Santhi Raj Kolamuri, Lorenzo Mazza, Nicolas Padoy, Sophia Bano, Annika Reinke, Lena Maier-Hein, Danail Stoyanov, Jakob N. Kather, Fiona R. Kolbinger, Sebastian Bodenstedt, Stefanie Speidel
arXiv AI
Aug 11

A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling

arXiv:2603. 27341v4 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites.

By Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
Hugging Face Trending Papers
Jun 24

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding.

arXiv Machine Learning
Aug 3

What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches

arXiv:2607. 29090v1 Announce Type: new Abstract: Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care.

By Yizhi Dong, Yuhe Ke, Hairil Rizal Abdullah, Yucheng Xing, Kevan Kai Bing Teo, Ling Huang, Mengling Feng