The study examines the nature of questions students pose to generative AI during two CS2 programming tasks, classifying 830 interactions into 18 categories based on the Graesser taxonomy. Results reveal that a limited set of question types dominates student inquiries and that the distribution of question types shifts significantly as the task progresses.
By Matin Amoozadeh, Amin Alipour
arXiv:2606. 11744v1 Announce Type: cross Abstract: Large language models are now widely used for everyday learning, but the underlying interactions are typically unstructured chats rather than following a curriculum.
By Sidney Tio, Arunesh Sinha, Pradeep Varakantham
TutorTrace is a new dataset and behavioral abstraction pipeline that captures learners’ low‑level IDE telemetry to make their behavioral context visible and computable in real time. The dataset, collected across 480 students in two introductory Python courses, includes 180 K telemetry events, 13 633 behavioral segments, and 27 continuously computed metrics, and it underpins a taxonomy of learner activity before, between, and after AI queries. Preliminary classroom tests show that behavior‑aware prompts reduce the time between queries, and the system can predict upcoming queries with AUROC scores of .726 and .717 on two held‑out tasks.
By David Barron, Xiaohang Tang, Rezky Dwisantika, Minsun Kim, David H. Smith IV, Jiaming Cui, Yan Chen
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv:2606. 12767v1 Announce Type: new Abstract: Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use.
By Sarah Elshabrawy, Rahul K. Dass, Ashok K. Goel
arXiv:2606. 13684v1 Announce Type: cross Abstract: Automatic Bloom's taxonomy classification of assessment questions can substantially reduce instructor workload, but labeling is subjective and teacher-dependent.
By Abdolali Faraji, Mohammadreza Molavi, Zohreh Rasoulkhani, Mohammadreza Tavakoli, G\'abor Kismih\'ok
arXiv:2606. 29049v1 Announce Type: new Abstract: Knowledge Tracing (KT) is important for personalized education but traditionally suffers from two key limitations: a reliance on shallow ID-based representations that neglect semantic depth and a restriction to single-granularity mastery estimation that overlooks hierarchical knowledge dependencies.
By Xinjin Li, Mengyue Wang, Yuzhen Lin, Pengbin Feng, Ziqi Sha, Yeyang Zhou, Yu Ma
arXiv:2608.22993v1 Announce Type: new
Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...
By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
arXiv:2607. 22598v1 Announce Type: cross Abstract: Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisions such as content selection and didactic structuring implicitly to the LLM, making tutoring strategies difficult to trace, evaluate, and reproduce.
By Laurent Brisson (IMT Atlantique - DSD), Maria Segarra (IMT Atlantique - INFO, Lab-STICC\_MOTEL), Gr\'egory Smits (IMT Atlantique - INFO, Lab-STICC\_MOTEL)
The paper presents a tutoring platform that combines a generative AI chatbot with a reinforcement learning algorithm to adaptively sequence practice problems for students learning Python. In a five‑month field study across ten high schools, the adaptive sequencing improved unassisted final exam performance by 0.15 standard deviations, with mediation analysis indicating that higher engagement drove the gains. The study demonstrates that signals from student‑chatbot interactions can be leveraged to personalize and optimize learning at scale.
By Angel Tsai-Hsuan Chung, Botong Zhang, Ling-Chieh Kung, Hamsa Bastani, Osbert Bastani
The paper introduces Practical Integrated Cross-consistent Knowledge Tracing (PICKT), a model that incorporates multiple feature types to improve Knowledge Tracing robustness when new questions lack interaction history. It evaluates the impact of difficulty, textual, and knowledge‑map relational features, finding that difficulty is especially informative for hard questions, while fused text and map features help estimate unseen questions by leveraging similar ones seen during training. The study concludes that prioritizing feature annotation aligned with educational service characteristics is essential for maintaining robust diagnostics in Intelligent Tutoring Systems.
By Wonbeen Lee, Channyoung Lee, Junho Sohn, Hansam Cho
arXiv:2606. 15766v1 Announce Type: new Abstract: A central pedagogical value evaluated in AI tutor benchmarks is scaffolding: guiding students through graduated steps toward a solution.
By Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson