arXiv:2608. 02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains.
By Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
arXiv:2608. 01664v1 Announce Type: cross Abstract: We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks.
By Mohamed Basem, Vincent Christlein
arXiv:2608.22429v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this relian...
By Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie
The paper introduces V‑Rubrics, a reinforcement‑learning framework that evaluates vision‑language model responses by breaking them into atomic propositions and scoring them on Visual Faithfulness, Reasoning Consistency, and Instruction Following. Using a fine‑tuned Qwen3‑VL‑8B‑Instruct model and a newly created 50K‑example V‑Rubrics dataset, the authors demonstrate that rubric‑based GRPO outperforms both a shared SFT baseline and an answer‑only GRPO, especially on knowledge‑oriented and visually grounded reasoning tasks.
By Shulin Tian, Minglun Li, Yuhao Dong, Hao Ding, Jiarui Yao, Haiwen Diao, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
LiteMedCoT-VL is a parameter‑efficient pipeline that transfers chain‑of‑thought reasoning from a 235B teacher model to a 2B student model using LoRA fine‑tuning on explanation‑enriched data. The approach enables a compact vision‑language model to perform medical visual question answering without relying on image captions, achieving 64.9% accuracy on the PMC‑VQA benchmark—an 11‑point improvement over the zero‑shot Qwen3‑VL‑4B baseline. Visual grounding analysis confirms that the model bases its predictions on image content rather than textual priors.
By Runze Ma, Shunbo Jia, Haonan Lyu, Guo Liu, Caizhi Liao
LEGO-OPD introduces a factorized teacher composition for multimodal on‑policy distillation, combining a Language Expert and a Grounding Expert into a single teacher distribution. By treating the language expert as a prior over tokens and the grounding expert as a visual likelihood that updates this prior, the method decouples language reasoning from visual grounding. Adaptive calibration further adjusts the influence of visual evidence at each decoding prefix, preventing over‑ or under‑supervision. Experiments with Qwen3 models demonstrate that LEGO‑OPD outperforms both single‑ and multi‑teacher baselines on multimodal and text‑only reasoning tasks, improving visual perception while preserving language reasoning.
By Jaeyun Shin, Hangeol Chang, Jong Chul Ye
arXiv:2604. 01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visual cues with retrieved textual evidence.
By Marco Morini, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise.
arXiv:2606. 05718v1 Announce Type: cross Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher.
By Kanghui Tian, Siyuan Liu, Ziang Yan, Sheng Xia, Shuai Dong, Yi Wang
arXiv:2609.36838v1 Announce Type: cross
Abstract: Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teache...
By Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li, Wen Luo, Yang Xu, Yufan Shen, Luke Mao, Yang Du, Asher Qin, Houfeng Wang
The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set into a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the system supports both VQA tasks with frozen inference, while also highlighting areas for improvement such as schema robustness and cross-family transfer.
By Xunlan Zhou, Xianliang Yang, Li Zhao
arXiv:2606. 19120v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target.
By Sihan Wang, Xiyao Liu, Lianqing Liu, Zhi Han