Hugging Face Trending Papers

Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing

arXiv AI
Sep 2

Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing

The study compares three AI interaction designs for learning nuclear safety protocols: an unrestricted conversational bot, a Socratic hint‑guided bot, and an adaptive tutoring system that uses brain‑derived cognitive engagement. Fifty participants, with no prior knowledge, completed a video lesson, pre‑test, AI‑driven assessment, and post‑test. Results showed the unrestricted bot yielded the highest learning gains, the adaptive system produced the greatest EEG engagement, and usage patterns revealed that unrestricted users mainly sought direct answers while Socratic users initially tried reasoning before disengaging.

By Alexandre Clin Deffarges, Nataliya Kosmyna, Pattie Maes
arXiv AI
3d ago

Examining Variation in How Guided AI Tutors Resolve Student Impasses

The study analyzes 20,462 student turns from 1,260 sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns categorized as conceptual errors, expressed uncertainty, or help‑seeking. Three tutoring conditions—baseline, no‑direct‑answer, and guided—were simulated, revealing that the baseline tutor often gave direct answers, the no‑direct‑answer tutor always asked follow‑up questions, and the guided tutor varied its responses based on context. Impasse trajectories showed that each additional impasse turn reduced the likelihood of recovery, while addressing errors became increasingly beneficial compared to repeated scripted questioning.

By Bakhtawar Ahtisham, Kirk Vanacore, Alessandra Napoli, Josh Arens, Ksenia Ionova, Clayton Cohn, Shima Salehi, Rene Kizilcec
arXiv Computation and Language
Sep 21

CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

CoLearn is an interactive, agentic tutoring system that learns about each learner through a persistent memory of mastery and misconceptions, updated with a Bayesian Knowledge Tracing model that uses a large language model as an observation function. It generates personalized questions targeting the learner’s weakest topics and recurring misconceptions, and provides a live evidence view for progress visualization and blind A/B comparison. In blind A/B tests, learners preferred questions conditioned on this memory 68‑69% of the time, and simulations show the agent’s belief converges toward the learner’s true mastery.

By Kailai He, Zhihao Wu, Linhai Zhang, Runcong Zhao, Yulan He, Jiazheng Li
arXiv AI
Sep 25

Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming

The study examined how different designs of AI teaching assistants (AI TAs) affect students in an introductory programming course. Four AI TAs were compared based on pedagogical style (Socratic vs. Direct instruction) and context awareness (no context vs. full context). Results showed that the Socratic AI TA with full context received the lowest favorability ratings, had the highest interaction stress, the most external LLM use, and the lowest comprehension outcomes, though differences were not statistically significant.

By Madeleine Eastwood, Harshith Narne, Joseph Hilby, Paul Denny, Ashish Aggarwal, Amanpreet Kapoor
arXiv AI
Aug 28

TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education

TutorTrace is a new dataset and behavioral abstraction pipeline that captures learners’ low‑level IDE telemetry to make their behavioral context visible and computable in real time. The dataset, collected across 480 students in two introductory Python courses, includes 180 K telemetry events, 13 633 behavioral segments, and 27 continuously computed metrics, and it underpins a taxonomy of learner activity before, between, and after AI queries. Preliminary classroom tests show that behavior‑aware prompts reduce the time between queries, and the system can predict upcoming queries with AUROC scores of .726 and .717 on two held‑out tasks.

By David Barron, Xiaohang Tang, Rezky Dwisantika, Minsun Kim, David H. Smith IV, Jiaming Cui, Yan Chen
arXiv Machine Learning
Sep 14

Simulating Disengaged Students to Evaluate LLM-based Tutors

The paper introduces Disengagement-Aware Student Simulators (DAS2), a protocol that models five learner-engagement states—engaged, gaming, wheel-spinning, off-task, and mixed—to evaluate AI tutor performance before deployment. Using annotated tutoring sessions from ASSISTments09, DAS2’s rule-based labels matched human consensus in 81% of cases, and conditioning simulations on intended states narrowed the correctness-rate gap between simulated and authentic sessions for gaming and wheel-spinning behaviors. The study also compares five AI tutors across these states, finding stable relative rankings but state-specific performance differences, and notes that automated evaluation does not fully align with human judgment.

By Xianghui Meng, Jionghao Lin