arXiv Computation and Language By Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci

Likelihood Ranking doesn't Scale Like Prompting in LLMs

Read the original on arXiv Computation and Language →

The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 10

Unexplored flaws in multiple-choice VQA make benchmarking unreliable

The paper demonstrates that multiple‑choice visual question answering (MC‑VQA) benchmarks are unreliable because model performance is highly sensitive to semantically neutral prompt formatting choices—such as option ID sets, delimiters, and separators—despite protocols that mitigate option‑order effects. Across seven multimodal large language models and five datasets, the authors observed frequent rank reversals when systematically varying 48 equivalent prompt formats, attributing the instability to tokenizer‑induced token fusion or removal and to how option ID sets influence attention patterns. Consequently, MC‑VQA rankings correlate weakly with open‑ended evaluation, revealing that MC‑VQA reflects option‑selection dynamics as well as multimodal reasoning.

By Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagdonat, Stephan G\"unnemann, Leo Schwinn
arXiv Computation and Language
Aug 28

Cascaded Batch Prompting

Cascaded Batch Prompting introduces a two‑stage method that separates complex reasoning from symbol grounding to address the unpredictability of conventional batch prompting. Experiments on multiple‑choice question answering and natural language inference show that this approach outperforms standard single prompting while maintaining a speedup proportional to batch size. The technique establishes a new state‑of‑the‑art position on the Pareto frontier for efficiency and performance.

By Sho Hoshino, Peinan Zhang
Hugging Face Trending Papers
Sep 24

Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

The paper shows that chain‑of‑thought (CoT) instructions can distort multiple‑choice vision‑language model evaluation when a scorer appends a reasoning cue but reads answer‑label logits before the model generates any rationale. This CoT‑prefix scoring causes significant drops in accuracy (e.g., Qwen2.5‑VL‑7B falls from 80.76% to 45.48% on ScienceQA) and leads most predictions to choose the first option. Analysis reveals that while answer information remains linearly accessible in late layers, the immediate readout is misled by probability mass shifting toward continuation tokens, and the issue varies across datasets and models.