arXiv AI

Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses

arXiv:2606. 08400v1 Announce Type: cross Abstract: Graduate-level research reading report assessment creates a substantial labor burden for educators.

Hugging Face Trending Papers
Sep 4

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 responses each, the study validates that AI-generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework also identifies when human review is most needed, allowing more efficient allocation of expert effort while maintaining assessment quality.

arXiv AI
Sep 7

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 student responses each, the authors validate that AI‑generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework includes a correction workflow that flags cases needing human review, thereby reducing manual workload while maintaining assessment quality.

By Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adri\'an Silveira, Andr\'es Peri
arXiv AI
Sep 25

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.

By Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan