arXiv AI By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

Read the original on arXiv AI →

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 22

Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning

The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.

By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
arXiv AI
Aug 26

RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding

RefineRank introduces a lightweight module, RefineNet, that jointly refines bounding boxes and ranks them for surgical spatio‑temporal grounding. By combining frozen medical vision‑language features with proposals from a frozen open‑set detector, it predicts coordinate corrections and quality scores for each candidate box, selecting the best refined or original box. On MedVidBench, RefineRank achieves the highest reported STG mIoU of 0.421 and improves oracle bounds and overall mIoU in controlled evaluations.

By Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque
arXiv AI
Aug 28

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.

By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao