arXiv:2609.21402v1 Announce Type: new
Abstract: Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods f...
By Zhibo Zhang, Qijie Wang, Zengqiang Yan
arXiv:2603.29962v4 Announce Type: replace
Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....
By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding.
arXiv:2601.06847v2 Announce Type: replace-cross
Abstract: Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We po...
By Mengmeng Zhang, Xiaoping Wu, Hao Luo, Fan Wang, Yisheng Lv
SurgAtlas is the largest surgical video‑language dataset, containing 15,291 videos (2,391 hours) across 18 specialties and over 5,000 procedure types, all sourced from public YouTube. It uniquely includes open‑surgery videos at scale (6,182) alongside more than 9,000 minimally invasive recordings, and introduces standardized benchmarks for open‑surgery video understanding. The dataset offers a rich, multi‑tier annotation schema—segment‑level captions, step/phase descriptions, video‑level surgical narratives, and reasoning‑oriented VQA pairs—validated by experts and built through an automated LLM‑enriched pipeline.
"whyItMatters":"SurgAtlas provides an unprecedentedly large, diverse, and clinically validated resource that can train and benchmark multimodal surgical AI models, advancing the development of next‑generation foundation models for surgery."
By Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin, Ahmad M. Hider, Li Yayuan, Jing Bi, Susan Liang, Chenliang Xu, Donald S. Likosky, Jason J. Corso
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao