Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation
Read the original on arXiv Computation and Language →The paper investigates how different scoring schemes affect the evaluation of multiple-choice question answering (MCQA) models. It introduces six education-inspired scoring methods that assess abilities such as distractor elimination, abstention, confidence calibration, and self-correction. Experiments on large language models show that these alternative schemes can change model rankings, better predict user preferences, and reveal distinct capabilities compared to traditional accuracy scoring.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.