FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates how evaluation protocols and ground‑truth quality affect the classification performance of Multimodal Large Language Models (MLLMs). It identifies and corrects key issues such as discarded out‑of‑list outputs, inflated distractor choices, and poor open‑world mapping, and shows that design choices like batch size, image ordering, and text encoder selection significantly influence accuracy. Using a multilabel reannotation of 625 ImageNet‑1k classes (ReGT), the study finds that corrected labels can boost MLLM performance by up to 10.8%, narrowing the gap with supervised models, and demonstrates that MLLMs can assist human annotators in about half of difficult cases.
arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that...
arXiv:2608. 19726v1 Announce Type: cross Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder.
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e. g.
SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.