arXiv AI
Sep 17

I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

The study evaluated whether a large language model (GPT‑5) could score teacher‑child interactions in early childhood classrooms using the Classroom Assessment Scoring System (CLASS) framework, comparing its results to human raters. Across 87 video‑recorded observations from 38 classrooms in Hong Kong, AI scores converged most closely with human ratings in the Emotional Support domain, especially the Quality of Feedback dimension, while diverging more in procedural or context‑dependent areas such as Classroom Organization and Instructional Support. The findings suggest that transcript‑based AI scoring can serve as a preliminary screening tool to aid teacher reflection, but it is not yet reliable enough to replace trained observers for full CLASS evaluations.

By Y. Fong, J. Xiang, T. Y. D. Chan, K. Lee, E. Y. H. Lau
arXiv AI
Sep 18

Edustories: A Collection of Real-world Case Studies from Classroom Practices

Edustories is a dataset of 1,492 teacher‑written case studies that detail real elementary and high‑school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. The collection is designed to enable research on AI assistance in collective teaching contexts, such as evaluating large language models’ ability to predict the success of teacher interventions. Comparative tests show that current models achieve 58% accuracy, below the 64% accuracy of human experts, indicating a gap between AI and human expertise in predicting classroom outcomes.

By Michal \v{S}tef\'anik, Jan Nehyba, Jirina Karasova, Martin Fico, Lucie \v{S}karkov\'a, Mark\'eta Ko\v{s}atkov\'a, David Kosatka