Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be scored as correct even when the same model fails on associated atomic questions.
arXiv:2508.07292v4 Announce Type: replace-cross
Abstract: Endoscopic diagnosis is an iterative process in which clinicians acquire, compare, and verify local visual evidence before reaching a conclus...
By Yi Tang, Kai-Ni Wang, Liang-Peng Pu, Hui Tang, Xiaopu He, Guang-Quan Zhou
arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.
By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv:2609.24799v1 Announce Type: new
Abstract: Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generi...
By Yeji Kim, Mi-Young Kim, Randy Goebel
arXiv:2610.10156v1 Announce Type: new
Abstract: Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, exi...
By Leon Mayer, Lucas Luttner, Patrick Godau, Kai Fritzsche, Annika Reinke, Leonie Boland, Jule Brandt, Janne Heinecke, Chloe K. Nobuhara, Niklas Holzwarth, Evangelia Christodoulou, Marcel Knopp, Dominik Michael, Pascale Piermarco, Saliq Neyaz, Korhan Derin \"Ozarslan, Jakob Hennighausen, Carlos Aumente-Maestro, Tim R\"adsch, Dheeraj Baji, Peter Maximilian Full, Finn Aichholz, Justus Veit Erpenbeck, Linus Finn Schott, Bastian Winkelhausen, Claas de Boer, Bianca G\"uttner, Anneli Hummel, Gregor Just, Max Kirchner, Chenyang Li, Rozenn Raffaut, Ariel Rodriguez, Danush Kumar Venkatesh, Kevin Wang, Jinjing Xu, Mona Sheikh Zeinoddin, Salman Khan, Thomas M. Pausch, Stefanie Speidel, Danail Stoyanov, Daniel A. Hashimoto, Fiona R. Kolbinger, Thomas G. Weiser, Lena Maier-Hein
arXiv:2604.01569v2 Announce Type: replace
Abstract: Video multimodal large language models achieve strong results on existing benchmarks, but answer accuracy alone does not establish whether they can...
By Jiahao Meng, Yue Tan, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, Xiangtai Li, Haodong Duan, Yunhai Tong, Ming-Hsuan Yang
The paper introduces a semantic correctness taxonomy that categorizes open‑ended QA answers into eight ordered classes, distinguishing between correct, verbose, and hallucinated responses. It releases two datasets—CAP‑Correctness and CAP‑Statements—to support benchmark evaluation and NLI‑based training. The authors also propose CAP (Context‑Aware Precision), a reference‑based metric that scores question‑conditioned statements via bidirectional NLI and demonstrates superior performance under a monotonicity protocol.
By Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov
LiteMedCoT-VL is a parameter‑efficient pipeline that transfers chain‑of‑thought reasoning from a 235B teacher model to a 2B student model using LoRA fine‑tuning on explanation‑enriched data. The approach enables a compact vision‑language model to perform medical visual question answering without relying on image captions, achieving 64.9% accuracy on the PMC‑VQA benchmark—an 11‑point improvement over the zero‑shot Qwen3‑VL‑4B baseline. Visual grounding analysis confirms that the model bases its predictions on image content rather than textual priors.
By Runze Ma, Shunbo Jia, Haonan Lyu, Guo Liu, Caizhi Liao
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surfac...
arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.
By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existing multimodal RAG pipelines primarily face two critical challenges: first, standard semantic similarity retrievers frequently fetch topically overlapping yet answer-void distractor pages that mislead downstream generation; second, rigid single-pass pipelines heavily depend on initial retrieval success, where any omission of core evidence inevitably causes cascading errors.
The paper introduces SCoRE, an agentic framework for Visual Retrieval-Augmented Generation that explicitly selects and consolidates visual evidence before generating answers. It addresses two key challenges: sparse, scattered evidence and noisy exploration trajectories that obscure reasoning. By maintaining a textual ledger of relevant observations and reloading original images for a logical evidence sequence, SCoRE decouples reasoning from exploration and enforces strict visual grounding, with training that rewards evidence coverage, compactness, and answer correctness.
By Yucheng Shen, Lingyong Yan, Jiulong Wu, Shuaiqiang Wang, Jianmin WU, Dawei Yin, Min Cao