arXiv Computation and Language

Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

The paper investigates how different scoring schemes affect the evaluation of multiple-choice question answering (MCQA) models. It introduces six education-inspired scoring methods that assess abilities such as distractor elimination, abstention, confidence calibration, and self-correction. Experiments on large language models show that these alternative schemes can change model rankings, better predict user preferences, and reveal distinct capabilities compared to traditional accuracy scoring.

arXiv Computation and Language
Sep 11

Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.

By Yu-Chung Hsiao
arXiv AI
Aug 26

Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring

The paper introduces a cost‑aware framework that treats each prompt type as an arm in a multi‑armed bandit controller, enabling adaptive selection of optimal prompting strategies during inference for automated essay scoring. Experiments on IELTS Writing Task 2 essays demonstrate that this bandit-driven approach achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls by 78.4%. The study also presents the first cost‑reliability learning curves for essay scoring, offering actionable insights for educational technology platforms balancing operational costs against assessment validity.

By Olga Manakina, Igor Bogdanov
arXiv Computation and Language
Sep 25

Likelihood Ranking doesn't Scale Like Prompting in LLMs

The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.

By Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci
arXiv AI
Aug 19

Grading Needs a Rubric, Not Intelligence

Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.

By Jhen-Ke Lin