arXiv Computer Vision By Hai-Dang Nguyen, Huy-Hieu Pham

Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA

Read the original on arXiv Computer Vision →

The paper introduces a mixed-format medical visual question answering system that stabilizes both multiple-choice and free-text outputs. It employs an answer-text memory, a permutation-stabilized vision–language expert, and a sparse candidate-expanding router, extending the routable candidate set to include the expert’s top‑2 predictions. On a 1,403‑case retrospective analysis, this candidate expansion improves binary routing accuracy from 88.95% to 91.73%, rescues 56 errors, and achieves 92.23% overall performance, while ensuring all 475 open‑ended responses are schema‑valid without repair or retry.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Sep 7

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

MedProb is a lightweight probing framework that predicts multiple-choice medical visual question answering (Med‑VQA) answers directly from frozen vision‑language model (VLM) representations, avoiding free‑text generation. On datasets such as PATH‑VQA, SLAKE, and VQA‑RAD, MedProb extracts more answer‑relevant signal than prompting and outperforms both medical VLMs and agentic systems. The approach also narrows the performance gap between small and large models, shows that medical adaptation does not consistently improve linear decodability, and reveals positional biases in both prompting and generation.

By Erfan Nourbakhsh, Ke Yang, Anthony Rios
arXiv AI
Sep 1

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

Co-Annotator is a clinical AI system that distills expert gaze and dictation into two guidance components: a gaze‑aligned Vision Transformer that highlights fixation‑aligned areas of interest (AOIs) and an ontology‑bounded vision‑language model that pre‑fills editable biomarker summaries for retinal OCT. In controlled studies, each modality independently improved diagnostic accuracy and biomarker generation, and when combined across two academic institutions, the system increased correct diagnoses per minute by 40% and reduced comment editing time by 67% without compromising accuracy.

By Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor