arXiv:2606.27264v3 Announce Type: replace
Abstract: Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form tex...
By Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed, Muzammal Naseer, Salman Khan, Christoph Lippert
arXiv:2608.21140v1 Announce Type: cross
Abstract: Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report genera...
By Simon Vincent Abel, Heiko Hillenhagen, Michael G\"otz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
The article reviews how multimodal large language models (MLLMs) are expanding radiology AI beyond image‑specific tasks to multimodal reasoning, yet volumetric radiology poses a representational challenge because clinical interpretation needs full 3‑D spatial context and quantitative data. It surveys over 200 studies, categorizing advances in volumetric representation, multimodal understanding, and agentic orchestration, and introduces a Claim‑Design‑Validation framework to align technical, workflow, and clinical claims. The review emphasizes that native volumetric modeling and agentic capabilities must match spatial, quantitative, contextual, and workflow demands, and that clinical credibility hinges on faithful 3‑D representation, traceable behavior, proper validation, and defined human oversight.
By Zanting Ye, Shengyuan Liu, Xin Liu, Chenhui Wang, Zhisong Wang, Jiashuai Liu, Zipei Wang, Cheng Wang, Wentao Pan, Mengjie Fang, Di Dong, Mohammad Salmanpour, Arman Rahmim, Yu Gu, Yong Xia, Hongming Shan, Yixuan Yuan, Yefeng Zheng, Lijun Lu
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.
arXiv:2409.16183v2 Announce Type: replace
Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...
By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.
By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.
By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
MultiViewDx is a physician‑validated multimodal instruction dataset that links medical imaging studies with patient context and normalizes heterogeneous reports into an evidence‑linked workflow (evidence → findings → differential discussion → diagnosis). The dataset covers a wide range of imaging modalities and uses a unified image‑text retriever to ensure that instruction synthesis is grounded in source‑supported evidence. Fine‑tuned models on MultiViewDx achieve the highest average accuracy on four MedVQA benchmarks and receive the strongest overall rating on JAMA Clinical Challenge cases, with ablations confirming the importance of case‑level multi‑view organization and evidence‑linked reasoning.
By Junda Wang, Zonghai Yao, Yujan Ting, Eric Z. Chen, Hieu Tran, Hong Yu, Weijing Huang, Terrence Chen
arXiv:2608.28455v1 Announce Type: new
Abstract: Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated lab...
By Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol, \c{S}eyda Ertekin
arXiv:2604. 15231v2 Announce Type: replace Abstract: Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT).
By M\'elanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor
arXiv:2606. 06407v1 Announce Type: cross Abstract: Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and follow-up rely on comparison across prior studies and analogous reference cases.
By Tengfei Zhang, Ziheng Zhao, Lisong Dai, Xiaoman Zhang, Pengcheng Qiu, Ya Zhang, Yanfeng Wang, Weidi Xie
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao