arXiv:2607. 13041v1 Announce Type: cross Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them.
By Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird
arXiv:2607. 28647v1 Announce Type: cross Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design, interactive student learning, and feedback-driven refinement.
By Thang Doan Viet, Anh Nguyen Hoang, Tinh Luong Son, Anh Hoang Thi Ngoc, Huyen Giang Thi Thu, Tai Le Quy
The paper introduces a prompt‑engineering framework that personalizes large language model (LLM) teaching assistants across disciplines by tailoring responses to six learner‑specific dimensions, creating 96 distinct learner profiles. It also analyzes student queries through Bloom’s Taxonomy to gauge cognitive complexity, encoding both learner attributes and cognitive assessments into structured prompts that condition the LLM without retraining. Experiments using NLP metrics and a small human study demonstrate that this approach yields perceptible differences in response style and structure, with statistical evidence linking specific learner attributes to measurable changes.
By Saptarshi Basu, Sandeep Kakar, Ashok Goel
arXiv:2607. 11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address.
By Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran, Gabriel Stefan
arXiv:2607. 05403v1 Announce Type: cross Abstract: This paper aims to synthesize empirical research on AI tools used to support English as a second/foreign language (EL2) learners in Arab University classrooms (AUCs) between Jan 1st 2023 and Aug 31st 2025.
By Mohammed Q. Shormani, Muneef Yehia Alshawsh
arXiv:2608. 11259v1 Announce Type: cross Abstract: Many AI tutors leverage large language models (LLMs) today.
By Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski
PersonaPath is a new benchmark for knowledge‑centric personalized learning path planning, pairing 2,000 learner personas with a hierarchical knowledge graph of 347 textbooks, 1,751 units, and 4,092 concepts across 77 subjects. The study evaluates large language models on this benchmark, finding that even the best model achieves only a 29.5% final pass rate in Basic Education and fails to exceed 44.7% in tailoring paths to individual learners, highlighting a significant adaptivity gap. This work underscores the challenge of moving beyond exercise‑centric recommendation toward goal‑oriented, curriculum‑scale guidance.
By Yu Liu, Zeming Liu, Tianle Zhang, Zihao Cheng, Yuhang Guo, Kehai Chen, Min Zhang, Yunhong Wang, Haifeng Wang
arXiv:2608. 03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).
By Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon
arXiv:2604. 26962v3 Announce Type: replace-cross Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs).
By Bingxi Zhao, Jiahao Zhang, Xubin Ren, Zirui Guo, Tianzhe Chu, Yi Ma, Chao Huang
arXiv:2609. 29672v1 Announce Type: new Abstract: Artificial intelligence helps education most where an essential provision has been rationed by cost.
By Qiming Guo, Jinwen Tang, Xingran Huang, Hung-Yu Lin, Yafu Zhong, Xiatian Zhuang
Edustories is a dataset of 1,492 teacher‑written case studies that detail real elementary and high‑school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. The collection is designed to enable research on AI assistance in collective teaching contexts, such as evaluating large language models’ ability to predict the success of teacher interventions. Comparative tests show that current models achieve 58% accuracy, below the 64% accuracy of human experts, indicating a gap between AI and human expertise in predicting classroom outcomes.
By Michal \v{S}tef\'anik, Jan Nehyba, Jirina Karasova, Martin Fico, Lucie \v{S}karkov\'a, Mark\'eta Ko\v{s}atkov\'a, David Kosatka
Preply uses OpenAI to launch AI-generated lesson summaries, providing personalised feedback and language learning exercises.