arXiv Computation and Language

Bridging Network Psychometrics and Artificial Intelligence: An Ising-Potts Model with LLM-Derived Weights

arXiv AI
Aug 19

Grading Needs a Rubric, Not Intelligence

Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.

By Jhen-Ke Lin
arXiv AI
Jun 6

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

arXiv:2606. 05180v1 Announce Type: cross Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced.

By Ivo Bueno, Babette B\"uhler, Philipp Stark, Tim F\"utterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
arXiv AI
Sep 3

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

The paper discusses how large language models (LLMs) are used in various evaluation roles—examining benchmarks, judging other models, and rating human content—and frames each as a measurement problem. It proposes using Rasch measurement theory (RMT) to decompose ordinal ratings into distinct facets on a common scale, offering diagnostics for miscalibration and rater bias. A case study applying RMT to the Measuring Hate Speech corpus reveals systematic differences between LLMs and human raters in severity, calibration, robustness, sensitivity, and scale use, suggesting RMT should be part of the evaluation toolkit for LLMs in all roles.

By Pratik S. Sachdeva, Nathan Boudol
Hugging Face Trending Papers
Jul 21

Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited.

arXiv Computation and Language
Sep 10

Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

Edu-QuRating is a pipeline that scores and curates educational data across multiple dimensions—accuracy, engagement, structure, and audience appropriateness—using an LLM judge to label document pairs and distill these preferences into reusable Edu-QuRaters. The best Edu-QuRater achieves 91.7% accuracy against held‑out GPT‑4.1‑mini judgments and is applied to filter 322.25 M FineWeb‑Edu‑Fortified documents, improving small‑model pre‑training performance on nine benchmarks. Additionally, Edu-QuRater scores serve as reward signals in GRPO post‑training, yielding responses that are preferred for pedagogical quality and instruction following over the Qwen3‑4B base model.

By Oliver G. B. Garrod, Robin A. A. Ince, Meng Liu, Mohamed Huti, Moritz Boos, Amy Waldock, Dominic Andrews, Paul Atherton