arXiv:2606. 26079v1 Announce Type: cross Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines.
By Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
The paper demonstrates that multiple‑choice visual question answering (MC‑VQA) benchmarks are unreliable because model performance is highly sensitive to semantically neutral prompt formatting choices—such as option ID sets, delimiters, and separators—despite protocols that mitigate option‑order effects. Across seven multimodal large language models and five datasets, the authors observed frequent rank reversals when systematically varying 48 equivalent prompt formats, attributing the instability to tokenizer‑induced token fusion or removal and to how option ID sets influence attention patterns. Consequently, MC‑VQA rankings correlate weakly with open‑ended evaluation, revealing that MC‑VQA reflects option‑selection dynamics as well as multimodal reasoning.
By Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagdonat, Stephan G\"unnemann, Leo Schwinn
arXiv:2608.28316v1 Announce Type: new
Abstract: Static importance scores compress visual evidence into a single ranking, but the value of remaining evidence can change after one cue has been observed...
By Yunxuan Fang, Xinhe Wang
We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure.
arXiv:2606. 16682v3 Announce Type: replace Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge.
By Zewen Liu
arXiv:2605. 18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy.
By Qinwu Xu, Zhuoheng Li, Jessie Salas
arXiv:2608. 11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge.
By Karl Hanna, Chen Feng
ReWEIGH the Evidence is a training‑free decoding technique that calibrates token‑level ordinal visual evidence to reduce hallucinations in large vision‑language models. It aggregates vocabulary ranks across visual positions, compares candidates to a token‑specific reference derived from unlabeled images, and applies a bounded penalty only when evidence falls below this reference. Experiments on four 7B backbones show up to a 21.3% reduction in hallucinated object mentions while largely preserving or improving descriptive and general performance, with minimal added latency.
By Jihae Jeong, Junha Choi, Hwanjo Yu
arXiv:2605.26380v2 Announce Type: replace-cross
Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
By Jingru Chen, Yiming Liu, Mingtao Chen, Sijie Chen, Richeng Xuan, Liang Yang, Zhichao Hu, Fanyang Lu
The paper introduces FlipDir, a training‑free inference‑time technique that mitigates answer flips in vision‑language models by steering hidden states along a low‑rank subspace derived from contrastive image pairs. It employs a margin‑based gate to attenuate steering only during uncertain decoding steps, thereby restoring original predictions while keeping stable ones unchanged. The authors also present VisFlip, a benchmark framework that evaluates models across nine dataset‑variation combinations in scientific reasoning, robot‑scene understanding, and medical VQA, showing that FlipDir consistently outperforms existing methods on recovery and preservation metrics.
By Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang
arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2608.20999v1 Announce Type: new
Abstract: Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, a...
By Haiming Li, Yingsheng Liu, Jingmin Zhu, Siyuan Yan, Xieji Li, Jiajun Sun, Zhen Yu, Zongyuan Ge