arXiv Computation and Language

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

MedProb is a lightweight probing framework that predicts multiple-choice medical visual question answering (Med‑VQA) answers directly from frozen vision‑language model (VLM) representations, avoiding free‑text generation. On datasets such as PATH‑VQA, SLAKE, and VQA‑RAD, MedProb extracts more answer‑relevant signal than prompting and outperforms both medical VLMs and agentic systems. The approach also narrows the performance gap between small and large models, shows that medical adaptation does not consistently improve linear decodability, and reveals positional biases in both prompting and generation.

arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv AI
Aug 14

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

arXiv:2608. 12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification.

By Jakub Pokrywka, {\L}ukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
arXiv AI
Aug 28

MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

MedFG-VQA is a lightweight medical visual question answering framework that uses a memory bank to enhance low‑frequency DCT features and graph‑enhanced cross‑attention for visual‑textual alignment. It introduces Frequency‑Memory Fusion to retrieve and fuse low‑frequency information from a learnable memory bank, and Graph‑Aware Cross‑Attention to refine cross‑modal features via graph convolution. The authors also create SynMed‑VQA, a synthetic dataset of over 2 million QA pairs across nine imaging modalities, and show that MedFG‑VQA matches or outperforms larger models on several biomedical VQA benchmarks while keeping computational costs low.

By Haowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren, Xiangbo Shu, Yazhou Yao, Fumin Shen
arXiv AI
Jul 7

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

arXiv:2607. 04344v1 Announce Type: cross Abstract: While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular surface diseases (OSDs) is severely hindered by a paucity of high-fidelity, multimodal instruction-tuning data.

By Hao Wei, Wenjin Qi, Dasen Dai, Minqing Zhang, Wu Yuan