arXiv AI

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

arXiv:2608. 11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge.

arXiv Machine Learning
Sep 10

Unexplored flaws in multiple-choice VQA make benchmarking unreliable

The paper demonstrates that multiple‑choice visual question answering (MC‑VQA) benchmarks are unreliable because model performance is highly sensitive to semantically neutral prompt formatting choices—such as option ID sets, delimiters, and separators—despite protocols that mitigate option‑order effects. Across seven multimodal large language models and five datasets, the authors observed frequent rank reversals when systematically varying 48 equivalent prompt formats, attributing the instability to tokenizer‑induced token fusion or removal and to how option ID sets influence attention patterns. Consequently, MC‑VQA rankings correlate weakly with open‑ended evaluation, revealing that MC‑VQA reflects option‑selection dynamics as well as multimodal reasoning.

By Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagdonat, Stephan G\"unnemann, Leo Schwinn
arXiv Computation and Language
6d ago

Large Language Model Selection with Limited Annotations

arXiv:2605.24981v2 Announce Type: replace Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotati...

By Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel
arXiv Computation and Language
Aug 27

Localize-Then-Decide Guarantees for LLM Judgments

Large language models (LLMs) are increasingly used to evaluate output quality, but guaranteeing agreement with human judgments is difficult. The paper introduces a Localize-Then-Decide framework that first uses conformal prediction to narrow down a shortlist likely to contain the human-preferred response, then applies a calibrated confidence rule to select a single response or abstain. Experiments show this two-stage approach consistently yields higher guarantee success rates and greater coverage than single-stage baselines across various candidate sizes and datasets.

By Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin
arXiv Computation and Language
Sep 1

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

The paper argues that traditional global calibration metrics, such as Expected Calibration Error and Brier Score, are confounded by differences in model accuracy when comparing large language models. It introduces ACE, an accuracy‑controlled evaluation framework that offers Instance‑Aligned, Distribution‑Aligned, and Candidate‑Aligned views to provide fairer cross‑model comparisons. Experiments across various benchmarks reveal that many reported calibration advantages disappear after accuracy control and that model rankings often reverse, indicating that raw global metrics are unreliable for cross‑model calibration assessment.

By Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li, Nigel Collier, Deqing Yang
arXiv Computation and Language
Sep 10

Are LLMs Positionally Consistent Ordinal Classifiers? A Systematic Evaluation

Large language models (LLMs) used for ordinal classification exhibit positional bias, where changes in label order, demonstration order, and demonstration placement affect predictions. Systematic experiments across ten frontier LLMs, eight prompt/task/model factors, and five datasets reveal that all models are sensitive to these positional sources, and that accuracy and stability often diverge. Various correction methods, including pointwise, pairwise, and listwise inference, do not reliably mitigate the bias, though a comparison-based listwise approach shows the best overall balance yet varies across models and bias types.

By Yu Wang, Zhe Zhou, Menglin Liu, Ge Shi
arXiv AI
6d ago

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.

By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv AI
Sep 25

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

The paper shows that chain‑of‑thought (CoT) instructions can distort evaluation of vision‑language models (VLMs) when a scorer reads answer‑label logits before the model generates a rationale. On ScienceQA, Qwen2.5‑VL‑7B’s accuracy falls from 80.76% to 45.48% under this CoT‑prefix scoring, and most predictions incorrectly pick the first option. Linear probes and free generation recover most of the lost accuracy, indicating that the answer information remains in the hidden states but is missed by the early readout. The authors explain the mismatch with vocabulary and layer diagnostics, noting that probability mass shifts toward continuation tokens while answer information stays linearly accessible in later layers. The effect varies across datasets and models, but the study demonstrates that CoT‑prefix scoring can misrepresent model knowledge unless the requested and scored outputs are aligned.

By Zeyan Li, Siyuan Qiu, Jianfeng Xu