The paper introduces ReasonCVS, a structured reasoning framework for assessing the Critical View of Safety in laparoscopic cholecystectomy. It uses a Vision‑Language Model to build an Anatomical Scene Graph Abstraction and a Large Language Model–based Rationale‑Aware Reasoning Agent to verify sub‑criteria, producing a final verdict with traceable clinical rationale. Experiments on the Endoscapes‑CVS201 benchmark show ReasonCVS outperforms existing methods with a 68.1% mAP while offering interpretable, criterion‑level explanations.
By Qing Xu, Yuxiang Luo, Zhen Chen
arXiv:2608. 04575v1 Announce Type: cross Abstract: Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions.
By Chen Yang, Shenxiang Zeng, Haoyang Zhao, Zhouyuan Xu, Youquan He, Haoyu Li, Mingyi Deng, Jiansheng Fan, Chen Wang
SurgRAW introduces a multi‑agent, chain‑of‑thought workflow for robotic surgical video analysis, leveraging a new SurgCoTBench benchmark with 14,256 QA pairs across five surgical tasks. The system uses an orchestrator to split scene understanding into two reasoning streams, panel‑discussion‑style collaboration among task‑specific agents, and retrieval‑augmented generation to incorporate surgical knowledge. Experiments show SurgRAW outperforms mainstream vision‑language models and a supervised baseline by 14.61% accuracy.
By Chang Han Low, Ziyue Wang, Tianyi Zhang, Zhu Zhuo, Zhitao Zeng, Evangelos B. Mazomenos, Yueming Jin
arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
arXiv:2608.21864v1 Announce Type: cross
Abstract: The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often...
By Md Asaduzzaman Jabin, Zihao Wu, Tianming Liu
MedVA is an end‑to‑end neuro‑symbolic agentic system designed to streamline medical volume visualization. It combines a neuro‑symbolic intent formulation agent that refines natural‑language requests with symbolic reasoning, a multi‑model ROI identification agent that uses pretrained medical segmentation models to locate specified regions, and an objective‑driven visualization optimization agent that evaluates ROI visibility using a volume‑based objective. Extensive evaluations and a formative user study demonstrate the system’s effectiveness and high usability across users with varying expertise.
By Haill An, Suhyeon Kim, Minjun Kang, Eunwoo Lee, Bin Sheng, Lei Bi, Younhyun Jung