arXiv Computer Vision By Yaojun Hu, Danyang Tu, Yang Liu, Jiajin Zhang, Wei Fang, Zhiqiang Liu, Chunlai Dong, Yingda Xia, Haochao Ying, Jian Wu, Ling Zhang

SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

Read the original on arXiv Computer Vision →

SeVeR is a selective visual exposure framework designed for volumetric medical visual question answering, particularly in multi-sequence MRI where redundant anatomical regions can overwhelm decoders. The authors introduce BreMRIs-VQA, a new breast MRI benchmark with 1.19 million QA pairs from 71 k sequences and 12.9 k patients, covering free-text and multiple-choice questions. SeVeR compresses dense volumes into modality-wise prototypes, retrieves complementary multi-level evidence using change-aware gated attention, and is trained with a marginal-utility self-consistency objective to suppress unhelpful retrieval, leading to improved discriminative and generative performance while exposing far fewer visual tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv AI
Jun 17

Enhancing Pathological VLMs with Cross-scale Reasoning

arXiv:2606. 17412v1 Announce Type: cross Abstract: Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis.

By Chi Phan, Tianyi Zhang, Qiaochu Xue, Yufeng Wu, Dan Hu, Zeyu Liu, Sudong Wang, Yueming Jin