arXiv Computation and Language By Erfan Nourbakhsh, Ke Yang, Anthony Rios

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

Read the original on arXiv Computation and Language →

MedProb is a lightweight probing framework that predicts multiple-choice medical visual question answering (Med‑VQA) answers directly from frozen vision‑language model (VLM) representations, avoiding free‑text generation. On datasets such as PATH‑VQA, SLAKE, and VQA‑RAD, MedProb extracts more answer‑relevant signal than prompting and outperforms both medical VLMs and agentic systems. The approach also narrows the performance gap between small and large models, shows that medical adaptation does not consistently improve linear decodability, and reveals positional biases in both prompting and generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv AI
Aug 14

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

arXiv:2608. 12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification.

By Jakub Pokrywka, {\L}ukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
arXiv AI
Aug 28

MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

MedFG-VQA is a lightweight medical visual question answering framework that uses a memory bank to enhance low‑frequency DCT features and graph‑enhanced cross‑attention for visual‑textual alignment. It introduces Frequency‑Memory Fusion to retrieve and fuse low‑frequency information from a learnable memory bank, and Graph‑Aware Cross‑Attention to refine cross‑modal features via graph convolution. The authors also create SynMed‑VQA, a synthetic dataset of over 2 million QA pairs across nine imaging modalities, and show that MedFG‑VQA matches or outperforms larger models on several biomedical VQA benchmarks while keeping computational costs low.

By Haowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren, Xiangbo Shu, Yazhou Yao, Fumin Shen