arXiv:2609.14284v1 Announce Type: new
Abstract: Criterion-level grading connects examination performance to learning outcomes, but manual marking introduces workload and variation between markers. Th...
By Md Khalid Syfullah, Asif Hasan Tonmoy, Saad Ahmed, S. M. Jahangir Alam
arXiv:2606. 08855v1 Announce Type: new Abstract: This paper examines the limitations of fully digital and partially digital e-assessment approaches in summative examinations in higher education.
By Hartwig Grabowski, Michael Canz
Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.
By Jhen-Ke Lin
arXiv:2606. 11477v1 Announce Type: cross Abstract: Correcting handwritten exams by hand is time-consuming and error-prone, particularly for large cohorts, while fully digital exams tend to force a didactic narrowing towards closed question formats.
By Hartwig Grabowski
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
The study evaluates GPT‑5.5’s ability to grade handwritten physics assessments, using 10,364 scanned pages from 520 submissions by 416 candidates across a national Olympiad theory exam, a final selection camp, and a university quantum‑mechanics exam. Each submission was graded twice, with the second round incorporating refined instructions after analyzing first‑round disagreements. The AI’s total‑score correlations with official marks ranged from 0.91 to 0.97, and it successfully identified the same five‑student team for the final Olympiad selection as human graders, though exact partial‑credit grading—especially in experimental work—remained challenging.
"Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control."
By Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar
arXiv:2607. 01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps.
By Aastha Sapkota, M. G. Sarwar Murshed
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
By Delip Rao, Chris Callison-Burch
arXiv:2608. 07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment.
By Hongfei Yan, Jiangkai Xiong, Yiqing Li, Chong Chen
arXiv:2606. 17507v1 Announce Type: new Abstract: Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment.
By Xiwei Xu, Chen Wang, Jacky Jiang, Phil Yang, Qian Fu, Mohan Dhall, Wenjie Zhang, Liming Zhu
arXiv:2603. 27223v2 Announce Type: replace-cross Abstract: We present EuraGovExam, a multilingual and multimodal benchmark sourced from real-world civil service examinations across five representative Eurasian regions: South Korea, Japan, Taiwan, India, and the European Union.
By Jaeseong Kim, Chaehwan Lim, Sang Hyun Gil, Suan Lee
arXiv:2607. 29624v1 Announce Type: cross Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback.
By Ilya Mikhelson