arXiv AI By Yunhang Qian, Jiaquan Yu, Jiawei Liu, Meng Wang, Hongwei Bran Li, Xiaobin Hu

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation

Read the original on arXiv AI →

arXiv:2607. 16303v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 21

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

LiteMedCoT-VL is a parameter‑efficient pipeline that transfers chain‑of‑thought reasoning from a 235B teacher model to a 2B student model using LoRA fine‑tuning on explanation‑enriched data. The approach enables a compact vision‑language model to perform medical visual question answering without relying on image captions, achieving 64.9% accuracy on the PMC‑VQA benchmark—an 11‑point improvement over the zero‑shot Qwen3‑VL‑4B baseline. Visual grounding analysis confirms that the model bases its predictions on image content rather than textual priors.

By Runze Ma, Shunbo Jia, Haonan Lyu, Guo Liu, Caizhi Liao
arXiv Computation and Language
Sep 23

MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis

MultiViewDx is a physician‑validated multimodal instruction dataset that links medical imaging studies with patient context and normalizes heterogeneous reports into an evidence‑linked workflow (evidence → findings → differential discussion → diagnosis). The dataset covers a wide range of imaging modalities and uses a unified image‑text retriever to ensure that instruction synthesis is grounded in source‑supported evidence. Fine‑tuned models on MultiViewDx achieve the highest average accuracy on four MedVQA benchmarks and receive the strongest overall rating on JAMA Clinical Challenge cases, with ablations confirming the importance of case‑level multi‑view organization and evidence‑linked reasoning.

By Junda Wang, Zonghai Yao, Yujan Ting, Eric Z. Chen, Hieu Tran, Hong Yu, Weijing Huang, Terrence Chen
arXiv AI
Jun 11

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.

By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv AI
Aug 28

From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.

By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao