arXiv:2607. 17834v1 Announce Type: cross Abstract: Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries.
By Yuhao Liu, Cheng Zhao, Guanghui Yue
arXiv:2508.07292v4 Announce Type: replace-cross
Abstract: Endoscopic diagnosis is an iterative process in which clinicians acquire, compare, and verify local visual evidence before reaching a conclus...
By Yi Tang, Kai-Ni Wang, Liang-Peng Pu, Hui Tang, Xiaopu He, Guang-Quan Zhou
arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.
By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv:2609.24799v1 Announce Type: new
Abstract: Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generi...
By Yeji Kim, Mi-Young Kim, Randy Goebel
arXiv:2610.10156v1 Announce Type: new
Abstract: Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, exi...
By Leon Mayer, Lucas Luttner, Patrick Godau, Kai Fritzsche, Annika Reinke, Leonie Boland, Jule Brandt, Janne Heinecke, Chloe K. Nobuhara, Niklas Holzwarth, Evangelia Christodoulou, Marcel Knopp, Dominik Michael, Pascale Piermarco, Saliq Neyaz, Korhan Derin \"Ozarslan, Jakob Hennighausen, Carlos Aumente-Maestro, Tim R\"adsch, Dheeraj Baji, Peter Maximilian Full, Finn Aichholz, Justus Veit Erpenbeck, Linus Finn Schott, Bastian Winkelhausen, Claas de Boer, Bianca G\"uttner, Anneli Hummel, Gregor Just, Max Kirchner, Chenyang Li, Rozenn Raffaut, Ariel Rodriguez, Danush Kumar Venkatesh, Kevin Wang, Jinjing Xu, Mona Sheikh Zeinoddin, Salman Khan, Thomas M. Pausch, Stefanie Speidel, Danail Stoyanov, Daniel A. Hashimoto, Fiona R. Kolbinger, Thomas G. Weiser, Lena Maier-Hein
The paper introduces SCoRE, an agentic framework for Visual Retrieval-Augmented Generation that explicitly selects and consolidates visual evidence before generating answers. It addresses two key challenges: sparse, scattered evidence and noisy exploration trajectories that obscure reasoning. By maintaining a textual ledger of relevant observations and reloading original images for a logical evidence sequence, SCoRE decouples reasoning from exploration and enforces strict visual grounding, with training that rewards evidence coverage, compactness, and answer correctness.
By Yucheng Shen, Lingyong Yan, Jiulong Wu, Shuaiqiang Wang, Jianmin WU, Dawei Yin, Min Cao
LiteMedCoT-VL is a parameter‑efficient pipeline that transfers chain‑of‑thought reasoning from a 235B teacher model to a 2B student model using LoRA fine‑tuning on explanation‑enriched data. The approach enables a compact vision‑language model to perform medical visual question answering without relying on image captions, achieving 64.9% accuracy on the PMC‑VQA benchmark—an 11‑point improvement over the zero‑shot Qwen3‑VL‑4B baseline. Visual grounding analysis confirms that the model bases its predictions on image content rather than textual priors.
By Runze Ma, Shunbo Jia, Haonan Lyu, Guo Liu, Caizhi Liao
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surfac...
arXiv:2604.01569v2 Announce Type: replace
Abstract: Video multimodal large language models achieve strong results on existing benchmarks, but answer accuracy alone does not establish whether they can...
By Jiahao Meng, Yue Tan, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, Xiangtai Li, Haodong Duan, Yunhai Tong, Ming-Hsuan Yang
The paper introduces a semantic correctness taxonomy that categorizes open‑ended QA answers into eight ordered classes, distinguishing between correct, verbose, and hallucinated responses. It releases two datasets—CAP‑Correctness and CAP‑Statements—to support benchmark evaluation and NLI‑based training. The authors also propose CAP (Context‑Aware Precision), a reference‑based metric that scores question‑conditioned statements via bidirectional NLI and demonstrates superior performance under a monotonicity protocol.
By Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existing multimodal RAG pipelines primarily face two critical challenges: first, standard semantic similarity retrievers frequently fetch topically overlapping yet answer-void distractor pages that mislead downstream generation; second, rigid single-pass pipelines heavily depend on initial retrieval success, where any omission of core evidence inevitably causes cascading errors.
arXiv:2608. 12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification.
By Jakub Pokrywka, {\L}ukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa