arXiv AI

Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment

arXiv AI
Aug 19

Grading Needs a Rubric, Not Intelligence

Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.

By Jhen-Ke Lin
arXiv AI
Jul 3

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.

By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
arXiv AI
Aug 24

Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes

The study evaluates GPT‑5.5’s ability to grade handwritten physics assessments, using 10,364 scanned pages from 520 submissions by 416 candidates across a national Olympiad theory exam, a final selection camp, and a university quantum‑mechanics exam. Each submission was graded twice, with the second round incorporating refined instructions after analyzing first‑round disagreements. The AI’s total‑score correlations with official marks ranged from 0.91 to 0.97, and it successfully identified the same five‑student team for the final Olympiad selection as human graders, though exact partial‑credit grading—especially in experimental work—remained challenging. "Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control."

By Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar