The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
arXiv:2609.37283v1 Announce Type: new
Abstract: Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evide...
By Xuyang Cao, Enyou Liu, Jun Zhao, Zhuoyun Liu, Jintao Fei, Leo
arXiv:2601.06847v2 Announce Type: replace-cross
Abstract: Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We po...
By Mengmeng Zhang, Xiaoping Wu, Hao Luo, Fan Wang, Yisheng Lv
arXiv:2601.09879v2 Announce Type: replace-cross
Abstract: Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report gen...
By Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
arXiv:2607.12896v3 Announce Type: replace
Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fr...
By Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye, Jianfeng Cao, Yixuan Yuan, Ting Ma
arXiv:2603. 14579v3 Announce Type: replace-cross Abstract: Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos.
By Andrew Seohwan Yu, Mohsen Hariri, Kunio Nakamura, Mingrui Yang, Xiaojuan Li, Vipin Chaudhary