Hugging Face Trending Papers

Expert Consensus on Criteria for the Automated Assessment of Laparoscopic Camera Navigation

Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are time-consuming and difficult to scale. Automated feedback could significantly enhance surgical training by providing immediate, standardized metrics.

arXiv AI
Aug 19

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque classification tools, offering transparent, quantitative performance indicators that could complement or replace traditional scoring methods.

By Mohammad Javad Ahmadi, Hamid D. Taghirad
Hugging Face Trending Papers
Aug 18

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

The paper introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world’s largest dataset of 2,000 surgical videos. Using advanced computer vision and signal‑processing techniques, the system extracts ten objective motion‑based metrics that correlate strongly with expert subjective ratings, achieving up to 87% accuracy. The framework’s explainability distinguishes it from prior opaque models, offering transparent, quantitative performance indicators that could complement or replace traditional subjective scoring.

arXiv AI
Jul 1

AI for Quality Assurance in the Operating Room

arXiv:2606. 30657v1 Announce Type: cross Abstract: Surgical outcomes depend not only on patient factors and postoperative care but are also strongly influenced by the quality of the operation itself.

By Pietro Mascagni, Lalith Sharan, Deepak Alapatt, Nicolas Padoy
arXiv AI
Aug 24

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

arXiv:2608.02471v2 Announce Type: replace-cross Abstract: In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense label...

By Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Hongkuan Shi, Qiuyu Yu, Qiang Xie, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding
arXiv Computer Vision
Aug 25

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

arXiv:2603.29962v4 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....

By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv AI
Aug 11

A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling

arXiv:2603. 27341v4 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites.

By Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
arXiv Machine Learning
Aug 19

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

This study investigates whether vision‑based models for surgical skill assessment learn representations that transfer across different scoring rubrics (GOALS and OSATS) using the LASANA and JIGSAWS datasets. By evaluating end‑to‑end training, Adaptive Sharpness‑Aware Minimization, and self‑supervised/contrastive pretraining, the authors find that models pretrained on JIGSAWS can transfer reasonably well to LASANA, but transfer to JIGSAWS fails, likely due to annotation inconsistencies. Control experiments with a Kinetics‑pretrained backbone show that task‑specific heads carry most of the skill prediction load, while the backbone provides general spatiotemporal features.

By Hanna Hoffmann, Felix von Bechtolsheim, Stefanie Speidel, Rebecca Hisey
arXiv Computer Vision
Sep 18

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

RAUL is a reference‑assisted reconstruction framework that recovers ureteroscope trajectories from endoscopic video alone, using a high‑quality reference exploration video for each phantom. It achieves a mean translation error of 0.5 mm and increases frame‑wise localization coverage from 50.5 % to 86.1 % compared to standard Structure‑from‑Motion. The reconstructed trajectories reveal significant differences in navigation metrics between high‑ and low‑experience trainees, enabling objective skill assessment without external tracking equipment.

By Fangjie Li, Mai Bui, Charan Mohan, Michael Miga, Matthieu Chabanas, Nicholas Kavoussi, Jie Ying Wu
arXiv AI
Aug 11

SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control

arXiv:2608. 07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target operative region is not a stable physical object but a latent and temporally evolving attention state.

By Rulin Zhou, Qiujie Song, Yujie Ma, An Wang, Wanhao Liu, Guoheng Ma, Yidu Wang, Guankun Wang, Xingrong Diao, Jiankun Wang, Chaowei Zhu, Xianming Liu, Hongliang Ren
arXiv Computer Vision
Sep 24

CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

CasCVS‑Net is a staged multi‑task cascade that jointly performs object detection, semantic segmentation, and Critical View of Safety (CVS) assessment for laparoscopic cholecystectomy. The model couples tasks through predicted anatomy—boxes guide segmentation and masks provide region‑level features for CVS classification—allowing CVS assessment to rely solely on model predictions. Trained on the Endoscapes dataset, CasCVS‑Net outperforms state‑of‑the‑art methods, achieving higher mAP and mIoU scores across detection, segmentation, and CVS tasks, especially for rare hepatocystic structures.

By Bock-Zien Toh, Yuanchuan Ren, Tay Aw Yu, Ng Khee Ong, Zhehua Mao, Sophia Bano