Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.
By Jhen-Ke Lin
arXiv:2606. 05180v1 Announce Type: cross Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced.
By Ivo Bueno, Babette B\"uhler, Philipp Stark, Tim F\"utterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
The paper discusses how large language models (LLMs) are used in various evaluation roles—examining benchmarks, judging other models, and rating human content—and frames each as a measurement problem. It proposes using Rasch measurement theory (RMT) to decompose ordinal ratings into distinct facets on a common scale, offering diagnostics for miscalibration and rater bias. A case study applying RMT to the Measuring Hate Speech corpus reveals systematic differences between LLMs and human raters in severity, calibration, robustness, sensitivity, and scale use, suggesting RMT should be part of the evaluation toolkit for LLMs in all roles.
By Pratik S. Sachdeva, Nathan Boudol
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
By Zewei Tian, Alex Liu, Lief Esbenshade, Michael Xiao, Zachary Zhang, Yulia L\'apicus, Thomas Han, Kevin He, Min Sun
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
By Delip Rao, Chris Callison-Burch
arXiv:2609.13824v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faste...
By Aakash Kumar Tiwari
arXiv:2607. 19219v1 Announce Type: cross Abstract: Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG).
By Xuefeng Jin, Jiashuo Zhang, Teng Cao, Bin Yang
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited.
Edu-QuRating is a pipeline that scores and curates educational data across multiple dimensions—accuracy, engagement, structure, and audience appropriateness—using an LLM judge to label document pairs and distill these preferences into reusable Edu-QuRaters. The best Edu-QuRater achieves 91.7% accuracy against held‑out GPT‑4.1‑mini judgments and is applied to filter 322.25 M FineWeb‑Edu‑Fortified documents, improving small‑model pre‑training performance on nine benchmarks. Additionally, Edu-QuRater scores serve as reward signals in GRPO post‑training, yielding responses that are preferred for pedagogical quality and instruction following over the Qwen3‑4B base model.
By Oliver G. B. Garrod, Robin A. A. Ince, Meng Liu, Mohamed Huti, Moritz Boos, Amy Waldock, Dominic Andrews, Paul Atherton
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
arXiv:2601. 02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration.
By Christopher Ormerod