arXiv:2504.02323v5 Announce Type: replace
Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored variou...
By Clayton Cohn, Ashwin T S, Naveeduddin Mohammed, Gautam Biswas
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 student responses each, the authors validate that AI‑generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework includes a correction workflow that flags cases needing human review, thereby reducing manual workload while maintaining assessment quality.
By Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adri\'an Silveira, Andr\'es Peri
The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 responses each, the study validates that AI-generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework also identifies when human review is most needed, allowing more efficient allocation of expert effort while maintaining assessment quality.
arXiv:2607. 01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps.
By Aastha Sapkota, M. G. Sarwar Murshed
arXiv:2606. 08400v1 Announce Type: cross Abstract: Graduate-level research reading report assessment creates a substantial labor burden for educators.
By Qilin Zhou, Zhuo Wang, Yue Li, W. K. Chan
arXiv:2606. 10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined.
By Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li, Jianlong Chen, Xiangfeng Wang
arXiv:2606. 05180v1 Announce Type: cross Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced.
By Ivo Bueno, Babette B\"uhler, Philipp Stark, Tim F\"utterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
The study examines how generative AI tools like ChatGPT perform on typical first‑year undergraduate mathematics assessment questions. By generating, transcribing, and blind‑marking AI responses to eight assessments covering the entire curriculum, the authors find that AI attains a first‑class level of performance, with consistency across modules that exceeds that of students in invigilated exams. The results suggest a need to redesign mathematics assessments to address the impact of generative AI.
By Benjamin J. Walker, Nikoleta Kalaydzhieva, Beatriz Navarro Lameda, Ruth A. Reynolds
arXiv:2608. 15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities.
By Alona Strugatski, Licol Zeinfeld, Giora Alexandron
The study evaluated whether a large language model (GPT‑5) could score teacher‑child interactions in early childhood classrooms using the Classroom Assessment Scoring System (CLASS) framework, comparing its results to human raters. Across 87 video‑recorded observations from 38 classrooms in Hong Kong, AI scores converged most closely with human ratings in the Emotional Support domain, especially the Quality of Feedback dimension, while diverging more in procedural or context‑dependent areas such as Classroom Organization and Instructional Support. The findings suggest that transcript‑based AI scoring can serve as a preliminary screening tool to aid teacher reflection, but it is not yet reliable enough to replace trained observers for full CLASS evaluations.
By Y. Fong, J. Xiang, T. Y. D. Chan, K. Lee, E. Y. H. Lau
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira