MedGround: Bridging the Evidence Gap in Medical Vision-Language Models with Verified Grounding Data
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 20477v1 Announce Type: cross Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations.
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
arXiv:2603. 14579v3 Announce Type: replace-cross Abstract: Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos.
arXiv:2509.22404v2 Announce Type: replace Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
arXiv:2608. 09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding.
arXiv:2505.03380v2 Announce Type: replace Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains...