Hugging Face Trending Papers

Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA

arXiv Computer Vision
Sep 2

Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA

The paper introduces a mixed-format medical visual question answering system that stabilizes both multiple-choice and free-text outputs. It employs an answer-text memory, a permutation-stabilized vision–language expert, and a sparse candidate-expanding router, extending the routable candidate set to include the expert’s top‑2 predictions. On a 1,403‑case retrospective analysis, this candidate expansion improves binary routing accuracy from 88.95% to 91.73%, rescues 56 errors, and achieves 92.23% overall performance, while ensuring all 475 open‑ended responses are schema‑valid without repair or retry.

By Hai-Dang Nguyen, Huy-Hieu Pham
arXiv Computation and Language
Sep 7

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

MedProb is a lightweight probing framework that predicts multiple-choice medical visual question answering (Med‑VQA) answers directly from frozen vision‑language model (VLM) representations, avoiding free‑text generation. On datasets such as PATH‑VQA, SLAKE, and VQA‑RAD, MedProb extracts more answer‑relevant signal than prompting and outperforms both medical VLMs and agentic systems. The approach also narrows the performance gap between small and large models, shows that medical adaptation does not consistently improve linear decodability, and reveals positional biases in both prompting and generation.

By Erfan Nourbakhsh, Ke Yang, Anthony Rios
arXiv AI
Sep 1

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

Co-Annotator is a clinical AI system that distills expert gaze and dictation into two guidance components: a gaze‑aligned Vision Transformer that highlights fixation‑aligned areas of interest (AOIs) and an ontology‑bounded vision‑language model that pre‑fills editable biomarker summaries for retinal OCT. In controlled studies, each modality independently improved diagnostic accuracy and biomarker generation, and when combined across two academic institutions, the system increased correct diagnoses per minute by 40% and reduced comment editing time by 67% without compromising accuracy.

By Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor
arXiv AI
Jun 11

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.

By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv AI
Aug 12

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

arXiv:2608. 10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust.

By Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu
arXiv Computation and Language
Sep 4

Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?

The paper evaluates uncertainty estimation (UE) methods for clinical vision‑language models (VLMs) on visual question answering (VQA). Across 8 UE techniques and 12 VLMs, UE quality tracks model accuracy, degrading where performance is weakest, and fails to signal uncertainty when models are stressed by hiding the correct answer (NOTA perturbations). However, UE on unperturbed inputs reliably predicts which predictions will collapse under NOTA, suggesting UE can diagnose model fragility.

By Arnisa Fazla, Alberto Testoni, Ameen Abu-Hanna, Barbara Plank, Iacer Calixto