arXiv AI

Grading Needs a Rubric, Not Intelligence

Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.

arXiv AI
Aug 20

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

The paper introduces THPT‑Ladder, a benchmark based on Vietnam’s 2025 National High School Graduation Examination’s convex grading scheme, which rewards partial credit non‑additively. It shows that standard accuracy metrics inflate model scores because they treat partial knowledge proportionally, whereas the official rubric penalizes incomplete correct sets. Using the benchmark, the authors demonstrate that this discrepancy can shift a model’s percentile ranking by up to 13 points among over 480,000 candidates.

By Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh
arXiv AI
Sep 25

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.

By Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan
arXiv Machine Learning
Sep 18

How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?

The study evaluates zero‑shot essay scoring using sub‑3B open language models that run locally on a single 8 GB consumer GPU. Four instruction‑tuned models (Qwen2.5‑0.5B, 1.5B, 3B and SmolLM2‑1.7B) were tested on all eight ASAP‑AES prompts, comparing rubric‑decomposed versus holistic prompting, different aggregation methods, and trait‑mapping strategies. Results show rubric‑decomposed prompting consistently outperforms holistic prompting, trait‑mapping is sensitive to calibration, and longer essays reduce error, yet the best local configuration (macro QWK 0.388) still falls short of human agreement and a length‑only baseline, suggesting these models are best suited for formative, human‑supervised feedback.

By Nguyen Dung Son, Dang Quang Minh, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
arXiv Computation and Language
3d ago

Three Ways Classical Test Theory Can Mislead About LLM Judges

The article examines how classical test theory statistics—Kuder‑Richardson coefficient, dependability index, and Livingston‑Lewis accuracy—can mislead when applied to large language model (LLM) judges that are evaluated with a single prompt and no gold labels. Using Claude Haiku 4.5 on 210 short‑answer items, the authors show that these metrics fail to isolate the judge’s performance because the judge’s single administration provides no variance component. They argue that reliable statements about an LLM judge require gold labels or varied scorer facets, and that bank design heavily influences reliability estimates.

By Louis Yiven Zhu
arXiv Computation and Language
Sep 25

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains. whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."

By Delip Rao, Chris Callison-Burch