AI Assistants Overassist
arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
Edustories is a dataset of 1,492 teacher‑written case studies that detail real elementary and high‑school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. The collection is designed to enable research on AI assistance in collective teaching contexts, such as evaluating large language models’ ability to predict the success of teacher interventions. Comparative tests show that current models achieve 58% accuracy, below the 64% accuracy of human experts, indicating a gap between AI and human expertise in predicting classroom outcomes.
arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on tra...
The study evaluated whether a large language model (GPT‑5) could score teacher‑child interactions in early childhood classrooms using the Classroom Assessment Scoring System (CLASS) framework, comparing its results to human raters. Across 87 video‑recorded observations from 38 classrooms in Hong Kong, AI scores converged most closely with human ratings in the Emotional Support domain, especially the Quality of Feedback dimension, while diverging more in procedural or context‑dependent areas such as Classroom Organization and Instructional Support. The findings suggest that transcript‑based AI scoring can serve as a preliminary screening tool to aid teacher reflection, but it is not yet reliable enough to replace trained observers for full CLASS evaluations.
arXiv:2609.28470v1 Announce Type: new Abstract: Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advanc...
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
arXiv:2608. 11259v1 Announce Type: cross Abstract: Many AI tutors leverage large language models (LLMs) today.
arXiv:2608.22993v1 Announce Type: new Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...
The paper evaluates Large Language Models for automatically analyzing responses in digital teacher simulations. Experiments compare DeBERTaV3 and Llama 3 across zero‑shot, few‑shot, and fine‑tuning settings, revealing that performance varies by characteristic and that Llama 3 consistently outperforms DeBERTaV3, especially when new characteristics must be identified. The findings suggest Llama 3 is preferable for dynamic simulation environments where teacher educators introduce new evaluation criteria.
arXiv:2607. 18529v1 Announce Type: cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality.
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
arXiv:2509.05346v3 Announce Type: replace Abstract: While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how thei...
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.