arXiv AI By Xiangyu Yin, Shiqi Wang, Abrar Alamri, Yasir Aljohani, Weichen Liu, Goeran Fiedler, Wei Gao

DrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis

Read the original on arXiv AI →

DrGait is a training‑free framework that transforms Vision‑Language Models into clinical planners for gait analysis. It separates semantic reasoning from geometric perception using a Triage‑Verification‑Synthesis workflow, where hypotheses are generated, verified with deterministic biomechanical tools, and refined in a closed‑loop. This approach reduces hallucinations and produces transparent, audit‑ready clinical reports with competitive diagnostic accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
4d ago

Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment

The paper introduces ReasonCVS, a structured reasoning framework for assessing the Critical View of Safety in laparoscopic cholecystectomy. It uses a Vision‑Language Model to build an Anatomical Scene Graph Abstraction and a Large Language Model–based Rationale‑Aware Reasoning Agent to verify sub‑criteria, producing a final verdict with traceable clinical rationale. Experiments on the Endoscapes‑CVS201 benchmark show ReasonCVS outperforms existing methods with a 68.1% mAP while offering interpretable, criterion‑level explanations.

By Qing Xu, Yuxiang Luo, Zhen Chen
arXiv AI
Sep 17

SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis

SurgRAW introduces a multi‑agent, chain‑of‑thought workflow for robotic surgical video analysis, leveraging a new SurgCoTBench benchmark with 14,256 QA pairs across five surgical tasks. The system uses an orchestrator to split scene understanding into two reasoning streams, panel‑discussion‑style collaboration among task‑specific agents, and retrieval‑augmented generation to incorporate surgical knowledge. Experiments show SurgRAW outperforms mainstream vision‑language models and a supervised baseline by 14.61% accuracy.

By Chang Han Low, Ziyue Wang, Tianyi Zhang, Zhu Zhuo, Zhitao Zeng, Evangelos B. Mazomenos, Yueming Jin
arXiv AI
Aug 17

MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.

By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
arXiv Computer Vision
Sep 15

MedVA: An End-to-End Neuro-Symbolic Agentic System for Medical Volume Visualization

MedVA is an end‑to‑end neuro‑symbolic agentic system designed to streamline medical volume visualization. It combines a neuro‑symbolic intent formulation agent that refines natural‑language requests with symbolic reasoning, a multi‑model ROI identification agent that uses pretrained medical segmentation models to locate specified regions, and an objective‑driven visualization optimization agent that evaluates ROI visibility using a volume‑based objective. Extensive evaluations and a formative user study demonstrate the system’s effectiveness and high usability across users with varying expertise.

By Haill An, Suhyeon Kim, Minjun Kang, Eunwoo Lee, Bin Sheng, Lei Bi, Younhyun Jung