arXiv Computation and Language By Nishant Balepur, Paiheng Xu, Wei Ai, Eunsol Choi, Rachel Rudinger, Jordan Boyd-Graber

Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

Read the original on arXiv Computation and Language →

The paper investigates how different scoring schemes affect the evaluation of multiple-choice question answering (MCQA) models. It introduces six education-inspired scoring methods that assess abilities such as distractor elimination, abstention, confidence calibration, and self-correction. Experiments on large language models show that these alternative schemes can change model rankings, better predict user preferences, and reveal distinct capabilities compared to traditional accuracy scoring.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.