arXiv Computer Vision

AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making

arXiv Machine Learning
1d ago

OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One

OmniMed-Jev is a new medical multimodal model that represents each decision as a Choice, Noul, or Score over a runtime-supplied candidate set, returning a full probability distribution for each decision. By unifying diverse imaging modalities and prediction tasks into a single candidate-conditioned probability model, it makes heterogeneous outputs comparable probabilities rather than task-specific strings. In controlled comparisons against a generative baseline, OmniMed-Jev’s reported probabilities align more closely with observed correctness, reducing calibration error by up to an order of magnitude and reliability error by up to two, while maintaining comparable point-prediction performance.

By Luyao Tang, Cheng Chen
arXiv AI
Jul 1

AI for Quality Assurance in the Operating Room

arXiv:2606. 30657v1 Announce Type: cross Abstract: Surgical outcomes depend not only on patient factors and postoperative care but are also strongly influenced by the quality of the operation itself.

By Pietro Mascagni, Lalith Sharan, Deepak Alapatt, Nicolas Padoy
arXiv AI
3d ago

From Compound Figures to Medical Multi-image Reasoning: Scaling Multimodal Large Language Models with Biomedical Literature

arXiv:2511.22232v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical inter...

By Zhen Chen, Yihang Fu, Rong Zhou, Serina Applebaum, Min Kyu Kim, Aidan Gilson, Morten Lee, Salahudeen Mirza, Gabriel Madera, Mauro Giuffre, Yuanting Pan, Roy Jiang, Hyunjae Kim, Hua Xu, Qingyu Chen
arXiv Computation and Language
Sep 4

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

The paper introduces MedQA-MM, a benchmark that exposes shortcut reasoning in medical multimodal multiple-choice questions. By auditing prompts, images, and modalities, the authors show that models often rely on textual cues rather than visual evidence, with full-input accuracy at 62.63% but only 5.21% when restricted to text. The study highlights the need for route-level evidence to validate true medical image reasoning.

By Benlu Wang, Yifan Zhang, Jiaqing Yu, Chin Siang Ong, Juncheng Huang, Zhuohao Li, Zhenyu Zhang, Arman Cohan, Hong Yu, Zonghai Yao
arXiv Computer Vision
Aug 25

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

arXiv:2603.29962v4 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....

By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv AI
Jun 11

Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark

arXiv:2602. 19502v2 Announce Type: replace Abstract: Agentic AI systems are increasingly capable of autonomous data science workflows, yet clinical prediction tasks demand domain expertise that purely automated approaches struggle to provide.

By Lalitha Pranathi Pulavarthy, Raajitha Muthyala, Aravind V Kuruvikkattil, Zhenan Yin, Rashmita Kudamala, Saptarshi Purkayastha
arXiv AI
Jun 30

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.

By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv Computer Vision
Sep 14

OphBiWSSD: Scaling Temporal Action Localization in Ophthalmic Surgeries with Bidirectional Weight-tied State Space Duality

OphBiWSSD is a new framework for temporal action localization in ophthalmic surgeries that uses Bidirectional State Space Duality to avoid the quadratic memory cost of Transformers. It employs a weight‑tied selective scan that incorporates both past and future surgical context, enabling linear‑time global synthesis of non‑causal temporal cues. On the OphNet benchmark, OphBiWSSD achieves state‑of‑the‑art mean Average Precisions of 44.42 % for phases and 43.08 % for operations, outperforming baselines by 6.80 % and 6.66 % respectively.

By Yang Liu, Qionghong Ma, Joongwon Chae, Lihui Luo, Yibing Shen, Yulin Zhuo, Yingting Zhu, Jiashu Chang, Xiaoyun Zhong, Dongmei Yu, Peter E. Lobie, Peiwu Qin, Chengming Yang
arXiv AI
Aug 24

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

arXiv:2608.02471v2 Announce Type: replace-cross Abstract: In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense label...

By Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Hongkuan Shi, Qiuyu Yu, Qiang Xie, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding