The paper investigates how evaluation protocols and ground‑truth quality affect the classification performance of Multimodal Large Language Models (MLLMs). It identifies and corrects key issues such as discarded out‑of‑list outputs, inflated distractor choices, and poor open‑world mapping, and shows that design choices like batch size, image ordering, and text encoder selection significantly influence accuracy. Using a multilabel reannotation of 625 ImageNet‑1k classes (ReGT), the study finds that corrected labels can boost MLLM performance by up to 10.8%, narrowing the gap with supervised models, and demonstrates that MLLMs can assist human annotators in about half of difficult cases.
By Nikita Kisel, Illia Volkov, Klara Janouskova, Jiri Matas
arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.
By Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, Tianyang Wang, Hao Xu
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that...
arXiv:2608. 19726v1 Announce Type: cross Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder.
By Nyx Iskandar, Saathvik Selvan, Slater Victoroff
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e. g.
SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.
By Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
The paper introduces Vision‑RL2, a region‑level reinforcement learning approach that optimizes a lightweight proposal network for fine‑grained multimodal large language model (MLLM) perception. By treating coherent image regions as actions and scoring them with a frozen MLLM reader, the method selectively focuses visual resolution on evidence, reducing token usage while improving accuracy across multiple benchmarks and backbones. The approach eliminates the need for region annotations, response sampling, or reasoning trajectories, and the refined proposals enable sparse encoding that magnifies relevant evidence.
By Yuheng Shi, Xiaohuan Pei, Minjing Dong, Chang Xu
arXiv:2607. 05978v1 Announce Type: cross Abstract: Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prolifically.
By Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen, Tal Remez
arXiv:2607. 09450v1 Announce Type: cross Abstract: Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations.
By Xingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu, Jiannan Ge, Jiaheng Zhang, Long Chen
arXiv:2608.21819v1 Announce Type: cross
Abstract: Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions wh...
By Jihyung Ko, Eunji Jung, Hyeongsub Kim, Ziseok Lee, Jae Won Cho, Sanghyun Jo, Kyungsu Kim
The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.
By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.
By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha