arXiv AI

Edustories: A Collection of Real-world Case Studies from Classroom Practices

Edustories is a dataset of 1,492 teacher‑written case studies that detail real elementary and high‑school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. The collection is designed to enable research on AI assistance in collective teaching contexts, such as evaluating large language models’ ability to predict the success of teacher interventions. Comparative tests show that current models achieve 58% accuracy, below the 64% accuracy of human experts, indicating a gap between AI and human expertise in predicting classroom outcomes.

arXiv AI
Jul 24

AI Assistants Overassist

arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.

By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv AI
Sep 17

I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

The study evaluated whether a large language model (GPT‑5) could score teacher‑child interactions in early childhood classrooms using the Classroom Assessment Scoring System (CLASS) framework, comparing its results to human raters. Across 87 video‑recorded observations from 38 classrooms in Hong Kong, AI scores converged most closely with human ratings in the Emotional Support domain, especially the Quality of Feedback dimension, while diverging more in procedural or context‑dependent areas such as Classroom Organization and Instructional Support. The findings suggest that transcript‑based AI scoring can serve as a preliminary screening tool to aid teacher reflection, but it is not yet reliable enough to replace trained observers for full CLASS evaluations.

By Y. Fong, J. Xiang, T. Y. D. Chan, K. Lee, E. Y. H. Lau
arXiv Computation and Language
Aug 25

LLM Pedagogical Behavior in AI Tutoring Interactions

arXiv:2608.22993v1 Announce Type: new Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...

By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
arXiv AI
Aug 25

Evaluating Large Language Models for automatic analysis of teacher simulations

The paper evaluates Large Language Models for automatically analyzing responses in digital teacher simulations. Experiments compare DeBERTaV3 and Llama 3 across zero‑shot, few‑shot, and fine‑tuning settings, revealing that performance varies by characteristic and that Llama 3 consistently outperforms DeBERTaV3, especially when new characteristics must be identified. The findings suggest Llama 3 is preferable for dynamic simulation environments where teacher educators introduce new evaluation criteria.

By David de-Fitero-Dominguez, Mariano Albaladejo-Gonz\'alez, Antonio Garcia-Cabot, Eva Garcia-Lopez, Antonio Moreno-Cediel, Erin Barno, Justin Reich
arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang