Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
Read the original on arXiv AI →The study evaluates GPT‑5.5’s ability to grade handwritten physics assessments, using 10,364 scanned pages from 520 submissions by 416 candidates across a national Olympiad theory exam, a final selection camp, and a university quantum‑mechanics exam. Each submission was graded twice, with the second round incorporating refined instructions after analyzing first‑round disagreements. The AI’s total‑score correlations with official marks ranged from 0.91 to 0.97, and it successfully identified the same five‑student team for the final Olympiad selection as human graders, though exact partial‑credit grading—especially in experimental work—remained challenging. "Reliable AI grading therefore depends on detailed rubrics and should be used as a second reader or audit tool under examiner control."
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.