arXiv AI

Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming

arXiv:2607. 03303v1 Announce Type: new Abstract: While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use in education.

arXiv AI
Sep 25

Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming

The study examined how different designs of AI teaching assistants (AI TAs) affect students in an introductory programming course. Four AI TAs were compared based on pedagogical style (Socratic vs. Direct instruction) and context awareness (no context vs. full context). Results showed that the Socratic AI TA with full context received the lowest favorability ratings, had the highest interaction stress, the most external LLM use, and the lowest comprehension outcomes, though differences were not statistically significant.

By Madeleine Eastwood, Harshith Narne, Joseph Hilby, Paul Denny, Ashish Aggarwal, Amanpreet Kapoor
arXiv Computation and Language
Aug 25

LLM Pedagogical Behavior in AI Tutoring Interactions

arXiv:2608.22993v1 Announce Type: new Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...

By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
arXiv AI
3d ago

Examining Variation in How Guided AI Tutors Resolve Student Impasses

The study analyzes 20,462 student turns from 1,260 sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns categorized as conceptual errors, expressed uncertainty, or help‑seeking. Three tutoring conditions—baseline, no‑direct‑answer, and guided—were simulated, revealing that the baseline tutor often gave direct answers, the no‑direct‑answer tutor always asked follow‑up questions, and the guided tutor varied its responses based on context. Impasse trajectories showed that each additional impasse turn reduced the likelihood of recovery, while addressing errors became increasingly beneficial compared to repeated scripted questioning.

By Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec
arXiv AI
Aug 28

TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education

TutorTrace is a new dataset and behavioral abstraction pipeline that captures learners’ low‑level IDE telemetry to make their behavioral context visible and computable in real time. The dataset, collected across 480 students in two introductory Python courses, includes 180 K telemetry events, 13 633 behavioral segments, and 27 continuously computed metrics, and it underpins a taxonomy of learner activity before, between, and after AI queries. Preliminary classroom tests show that behavior‑aware prompts reduce the time between queries, and the system can predict upcoming queries with AUROC scores of .726 and .717 on two held‑out tasks.

By David Barron, Xiaohang Tang, Rezky Dwisantika, Minsun Kim, David H. Smith IV, Jiaming Cui, Yan Chen
arXiv AI
Sep 21

Self-Explanation Tutor for Active Study of CS1 Worked Examples

The paper presents ESSE, a self‑explanation tutor that uses a large language model to give immediate feedback on students’ line‑by‑line explanations of introductory programming worked examples. It evaluates the LLM’s judgments against a domain expert and a crowd of non‑experts, finding that the model is reliable enough to serve as the tutor’s assessment engine. In an introductory Java course, the tutor’s feedback encourages students to persist, improves the completeness and conceptual depth of their explanations, and shows evidence of learning.

By Arun-Balajiee Lekshmi-Narayanan, Mohammad Hassany, Kamil Akhuseyinoglu, Rully Hendrawan, Peter Brusilovsky
arXiv AI
Sep 21

Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education

The paper introduces Beacon, a course‑specific Retrieval‑Augmented Generation (RAG) system that offers private, module‑aligned academic support to students. By grounding responses in approved teaching materials, Beacon aims to lower barriers to help‑seeking and encourage independent learning, especially in computing education where tasks are cumulative and demanding. Evaluation through questionnaires and interviews showed students found Beacon’s responses trustworthy and closely aligned with course content, viewing it as a useful first point of support before consulting lecturers or official resources.

By Andy Gray, Jake Hobbs
arXiv AI
Aug 19

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

The paper presents a tutoring platform that combines a generative AI chatbot with a reinforcement learning algorithm to adaptively sequence practice problems for students learning Python. In a five‑month field study across ten high schools, the adaptive sequencing improved unassisted final exam performance by 0.15 standard deviations, with mediation analysis indicating that higher engagement drove the gains. The study demonstrates that signals from student‑chatbot interactions can be leveraged to personalize and optimize learning at scale.

By Angel Tsai-Hsuan Chung, Botong Zhang, Ling-Chieh Kung, Hamsa Bastani, Osbert Bastani
arXiv AI
Jul 28

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

arXiv:2607. 22996v1 Announce Type: cross Abstract: Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes.

By Xiaokun Wang, Siyu Song, Wentao Liu, Xiaodong Zou