SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering
Read the original on arXiv Computer Vision →SeVeR is a selective visual exposure framework designed for volumetric medical visual question answering, particularly in multi-sequence MRI where redundant anatomical regions can overwhelm decoders. The authors introduce BreMRIs-VQA, a new breast MRI benchmark with 1.19 million QA pairs from 71 k sequences and 12.9 k patients, covering free-text and multiple-choice questions. SeVeR compresses dense volumes into modality-wise prototypes, retrieves complementary multi-level evidence using change-aware gated attention, and is trained with a marginal-utility self-consistency objective to suppress unhelpful retrieval, leading to improved discriminative and generative performance while exposing far fewer visual tokens.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.