arXiv Computer Vision By Haill An, Suhyeon Kim, Minjun Kang, Eunwoo Lee, Bin Sheng, Lei Bi, Younhyun Jung

MedVA: An End-to-End Neuro-Symbolic Agentic System for Medical Volume Visualization

Read the original on arXiv Computer Vision →

MedVA is an end‑to‑end neuro‑symbolic agentic system designed to streamline medical volume visualization. It combines a neuro‑symbolic intent formulation agent that refines natural‑language requests with symbolic reasoning, a multi‑model ROI identification agent that uses pretrained medical segmentation models to locate specified regions, and an objective‑driven visualization optimization agent that evaluates ROI visibility using a volume‑based objective. Extensive evaluations and a formative user study demonstrate the system’s effectiveness and high usability across users with varying expertise.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 24

Volumetric Radiology AI in the Era of Multimodal Large Language Models

The article reviews how multimodal large language models (MLLMs) are expanding radiology AI beyond image‑specific tasks to multimodal reasoning, yet volumetric radiology poses a representational challenge because clinical interpretation needs full 3‑D spatial context and quantitative data. It surveys over 200 studies, categorizing advances in volumetric representation, multimodal understanding, and agentic orchestration, and introduces a Claim‑Design‑Validation framework to align technical, workflow, and clinical claims. The review emphasizes that native volumetric modeling and agentic capabilities must match spatial, quantitative, contextual, and workflow demands, and that clinical credibility hinges on faithful 3‑D representation, traceable behavior, proper validation, and defined human oversight.

By Zanting Ye, Shengyuan Liu, Xin Liu, Chenhui Wang, Zhisong Wang, Jiashuai Liu, Zipei Wang, Cheng Wang, Wentao Pan, Mengjie Fang, Di Dong, Mohammad Salmanpour, Arman Rahmim, Yu Gu, Yong Xia, Hongming Shan, Yixuan Yuan, Yefeng Zheng, Lijun Lu
arXiv AI
Sep 17

SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis

SurgRAW introduces a multi‑agent, chain‑of‑thought workflow for robotic surgical video analysis, leveraging a new SurgCoTBench benchmark with 14,256 QA pairs across five surgical tasks. The system uses an orchestrator to split scene understanding into two reasoning streams, panel‑discussion‑style collaboration among task‑specific agents, and retrieval‑augmented generation to incorporate surgical knowledge. Experiments show SurgRAW outperforms mainstream vision‑language models and a supervised baseline by 14.61% accuracy.

By Chang Han Low, Ziyue Wang, Tianyi Zhang, Zhu Zhuo, Zhitao Zeng, Evangelos B. Mazomenos, Yueming Jin
arXiv AI
Jul 28

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.

By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv AI
Jun 16

XMedFusion: A Knowledge-Guided Multimodal Perception and Reasoning Framework for Autonomous Medical Systems

arXiv:2606. 14766v1 Announce Type: cross Abstract: Autonomous medical and robotic systems increasingly rely on intelligent perception and reasoning capabilities to interpret visual data and support clinical decision making.

By Hamza Riaz, Arham Haroon, Maha Baig, Muhammad Dawood Rizwan, Muhammad Naseer Bajwa, Muhammad Moazam Fraz
arXiv Computer Vision
4d ago

Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment

The paper introduces ReasonCVS, a structured reasoning framework for assessing the Critical View of Safety in laparoscopic cholecystectomy. It uses a Vision‑Language Model to build an Anatomical Scene Graph Abstraction and a Large Language Model–based Rationale‑Aware Reasoning Agent to verify sub‑criteria, producing a final verdict with traceable clinical rationale. Experiments on the Endoscapes‑CVS201 benchmark show ReasonCVS outperforms existing methods with a 68.1% mAP while offering interpretable, criterion‑level explanations.

By Qing Xu, Yuxiang Luo, Zhen Chen