SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
SeVeR is a selective visual exposure framework designed for volumetric medical visual question answering, particularly in multi-sequence MRI where redundant anatomical regions can overwhelm decoders. The authors introduce BreMRIs-VQA, a new breast MRI benchmark with 1.19 million QA pairs from 71 k sequences and 12.9 k patients, covering free-text and multiple-choice questions. SeVeR compresses dense volumes into modality-wise prototypes, retrieves complementary multi-level evidence using change-aware gated attention, and is trained with a marginal-utility self-consistency objective to suppress unhelpful retrieval, leading to improved discriminative and generative performance while exposing far fewer visual tokens.
arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.
arXiv:2608. 08307v1 Announce Type: cross Abstract: Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse density changes, with clinical language.
arXiv:2606. 06534v1 Announce Type: cross Abstract: Longitudinal medical visual question answering (VQA) requires reasoning about anatomical differences between an image of a current time point and an image of a referred time point.
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy.
arXiv:2606. 17412v1 Announce Type: cross Abstract: Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis.