arXiv:2606. 02156v1 Announce Type: cross Abstract: Anastomotic leak remains one of the most serious complications following colorectal cancer surgery, substantially affecting patient outcomes, recovery trajectories, and healthcare costs.
By Zahra Tabatabaei, Jon Sporring, Mark Bremholm Elleb{\ae}k, Alaa El-Hussuna
arXiv:2603.29962v4 Announce Type: replace
Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....
By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.
arXiv:2608.21441v1 Announce Type: cross
Abstract: Automated training of surgeons is one of the most crucial factors that significantly minimize surgical training risks and expenses. With recent advan...
By Mohammad Javad Ahmadi, Hamid D. Taghirad
arXiv:2603. 27341v4 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites.
By Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
arXiv:2606. 30657v1 Announce Type: cross Abstract: Surgical outcomes depend not only on patient factors and postoperative care but are also strongly influenced by the quality of the operation itself.
By Pietro Mascagni, Lalith Sharan, Deepak Alapatt, Nicolas Padoy
This study investigates whether vision‑based models for surgical skill assessment learn representations that transfer across different scoring rubrics (GOALS and OSATS) using the LASANA and JIGSAWS datasets. By evaluating end‑to‑end training, Adaptive Sharpness‑Aware Minimization, and self‑supervised/contrastive pretraining, the authors find that models pretrained on JIGSAWS can transfer reasonably well to LASANA, but transfer to JIGSAWS fails, likely due to annotation inconsistencies. Control experiments with a Kinetics‑pretrained backbone show that task‑specific heads carry most of the skill prediction load, while the backbone provides general spatiotemporal features.
By Hanna Hoffmann, Felix von Bechtolsheim, Stefanie Speidel, Rebecca Hisey
We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding.
arXiv:2602. 04819v5 Announce Type: replace-cross Abstract: Accurate risk stratification of precancerous polyps during routine colonoscopy screening is a key strategy to reduce the incidence of colorectal cancer (CRC).
By Aqsa Sultana, Rayan Afsar, Ahmed Rahu, Surendra P. Singh, Brian Shula, Brandon Combs, Derrick Forchetti, Vijayan K. Asari
arXiv:2512. 14732v3 Announce Type: replace-cross Abstract: Incidental findings in CT scans, though often benign, can have significant clinical implications and should be reported following established guidelines.
By Idan Tankel, Nir Mazor, Rafi Brada, Christina LeBedis, Guy ben-Yosef
Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are time-consuming and difficult to scale. Automated feedback could significantly enhance surgical training by providing immediate, standardized metrics.
arXiv:2606. 29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics.
By Jiashuo Sun, Yue He, Wenxuan Liu, Tao Mao, Jiazheng Wang, Xiang Chen, Min Liu