Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.
By Jhen-Ke Lin
arXiv:2608. 07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment.
By Hongfei Yan, Jiangkai Xiong, Yiqing Li, Chong Chen
The paper introduces THPT‑Ladder, a benchmark based on Vietnam’s 2025 National High School Graduation Examination’s convex grading scheme, which rewards partial credit non‑additively. It shows that standard accuracy metrics inflate model scores because they treat partial knowledge proportionally, whereas the official rubric penalizes incomplete correct sets. Using the benchmark, the authors demonstrate that this discrepancy can shift a model’s percentile ranking by up to 13 points among over 480,000 candidates.
By Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh
The study evaluates zero‑shot essay scoring using sub‑3B open language models that run locally on a single 8 GB consumer GPU. Four instruction‑tuned models (Qwen2.5‑0.5B, 1.5B, 3B and SmolLM2‑1.7B) were tested on all eight ASAP‑AES prompts, comparing rubric‑decomposed versus holistic prompting, different aggregation methods, and trait‑mapping strategies. Results show rubric‑decomposed prompting consistently outperforms holistic prompting, trait‑mapping is sensitive to calibration, and longer essays reduce error, yet the best local configuration (macro QWK 0.388) still falls short of human agreement and a length‑only baseline, suggesting these models are best suited for formative, human‑supervised feedback.
By Nguyen Dung Son, Dang Quang Minh, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
arXiv:2606. 11477v1 Announce Type: cross Abstract: Correcting handwritten exams by hand is time-consuming and error-prone, particularly for large cohorts, while fully digital exams tend to force a didactic narrowing towards closed question formats.
By Hartwig Grabowski
arXiv:2609.01345v1 Announce Type: new
Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natura...
By Dushyant Rajput