I code or AI code: A comparative evaluation of AI-rated scores in classroom observations
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The study evaluated whether a large language model (GPT‑5) could score teacher‑child interactions in early childhood classrooms using the Classroom Assessment Scoring System (CLASS) framework, comparing its results to human raters. Across 87 video‑recorded observations from 38 classrooms in Hong Kong, AI scores converged most closely with human ratings in the Emotional Support domain, especially the Quality of Feedback dimension, while diverging more in procedural or context‑dependent areas such as Classroom Organization and Instructional Support. The findings suggest that transcript‑based AI scoring can serve as a preliminary screening tool to aid teacher reflection, but it is not yet reliable enough to replace trained observers for full CLASS evaluations.
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
Edustories is a dataset of 1,492 teacher‑written case studies that detail real elementary and high‑school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. The collection is designed to enable research on AI assistance in collective teaching contexts, such as evaluating large language models’ ability to predict the success of teacher interventions. Comparative tests show that current models achieve 58% accuracy, below the 64% accuracy of human experts, indicating a gap between AI and human expertise in predicting classroom outcomes.
arXiv:2606. 18617v1 Announce Type: cross Abstract: There exist numerous tutor training platforms.
arXiv:2608. 05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one.
arXiv:2509.05346v3 Announce Type: replace Abstract: While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how thei...