arXiv:2606. 28882v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown the potential to generate code explanations that surpass those of peers in quality, offering promising opportunities for computer science education.
By Seth Bernstein, Paul Denny, Juho Leinonen, Kush Patel, Rayhona Nasimova, Matt Littlefield, Stephen MacNeil
arXiv:2607. 04572v1 Announce Type: new Abstract: Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem.
By Bonan Shen, Dingyan Shang, Youting Wang, Tao Ning
arXiv:2609.06095v1 Announce Type: cross
Abstract: Motivation: Undergraduate computing students increasingly turn to generative AI (GenAI) tools to understand abstract concepts through analogies. Anal...
By Seth Bernstein, Naaz Sibia
The study analyzes 20,462 student turns from 1,260 sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns categorized as conceptual errors, expressed uncertainty, or help‑seeking. Three tutoring conditions—baseline, no‑direct‑answer, and guided—were simulated, revealing that the baseline tutor often gave direct answers, the no‑direct‑answer tutor always asked follow‑up questions, and the guided tutor varied its responses based on context. Impasse trajectories showed that each additional impasse turn reduced the likelihood of recovery, while addressing errors became increasingly beneficial compared to repeated scripted questioning.
By Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec
arXiv:2608.22993v1 Announce Type: new
Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...
By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
The study examined how different designs of AI teaching assistants (AI TAs) affect students in an introductory programming course. Four AI TAs were compared based on pedagogical style (Socratic vs. Direct instruction) and context awareness (no context vs. full context). Results showed that the Socratic AI TA with full context received the lowest favorability ratings, had the highest interaction stress, the most external LLM use, and the lowest comprehension outcomes, though differences were not statistically significant.
By Madeleine Eastwood, Harshith Narne, Joseph Hilby, Paul Denny, Ashish Aggarwal, Amanpreet Kapoor
arXiv:2606. 30774v1 Announce Type: new Abstract: We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone.
By Bart{\l}omiej Cupia{\l}, Jan {\L}ojek, Miko{\l}aj Garstecki, Szymon Pob{\l}ocki, Alicja Ziarko, Piotr Mi{\l}o\'s
The paper introduces Knowledge Tracing Leveraging Problem‑Solving Process (KT‑PSP), a method that incorporates students’ problem‑solving steps to model mathematical proficiency more comprehensively than traditional knowledge tracing. It presents the KT‑PSP‑25 dataset and a new framework, StatusKT, which uses a teacher‑student‑teacher LLM pipeline to extract proficiency indicators, generate responses, and evaluate mastery. Experiments show that StatusKT improves prediction accuracy and offers interpretable explanations by explicitly modeling proficiency.
By Jungyang Park, Suho Kang, Jaewoo Park, Jaehong Kim, Jaewoo Shin, Seonjoon Park, Youngjae Yu
arXiv:2607. 22629v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time.
By Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).
By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang
arXiv:2607. 03303v1 Announce Type: new Abstract: While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use in education.
By Jerome Brender, Laila El-Hamamsy, Kim Uittenhove, Aitor Perez, Patrick Jermann, Francesco Mondada, Engin Bumbacher
arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.
By Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath