arXiv Machine Learning

One Mastery Threshold Does Not Fit All Knowledge Tracing Models

The study investigates how a single mastery threshold can produce divergent outcomes across different knowledge tracing (KT) models. By evaluating six KT models on four datasets with thresholds ranging from 0.50 to 0.99, the authors find that Bayesian Knowledge Tracing (BKT) is relatively insensitive to threshold changes, whereas neural models become increasingly selective as thresholds rise. The optimal threshold varies widely across models and instructional settings, and stricter thresholds can disproportionately limit advancement for weaker students.

arXiv Machine Learning
Aug 19

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

The study clusters 5,000 EdNet-KT3 learners into eight study‑strategy groups based on early‑session behaviors such as resource use, revision, video watching, and problem practice. These clusters predict later engagement metrics—like continued practice and session completion—but do not reliably forecast later unassisted accuracy or mastery. The findings suggest that behavioral clustering captures learning styles and engagement patterns rather than knowledge gains.

By Qingchuan Lyu, Yingxin Li, Albert Yang
arXiv Machine Learning
1d ago

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

The paper investigates how teacher signals influence parameter updates in Multi‑Teacher On‑Policy Distillation (MOPD) by analyzing Qwen3‑1.7B and SmolLM3‑3B. It shows that loss averaging, Adam’s first‑moment bias, BF16 rounding, and the choice of averaging rule all shape the gradients and ultimately affect task performance. The study quantifies these effects, revealing, for example, that token‑averaging favors longer responses and that BF16 rounding masks most weight changes.

By Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You
arXiv Machine Learning
Sep 25

Stable and Faithful Explanations for Knowledge Tracing

The paper introduces a validation protocol for knowledge‑tracing models that jointly assesses predictive performance, explanation stability, and faithfulness. Using engineered behavioral features from ASSISTments data, the authors compare an XGBoost model explained with TreeSHAP against four deep‑learning baselines, finding comparable predictive accuracy when information is matched and demonstrating that TreeSHAP rankings are stable and impactful. The study highlights how data preprocessing (e.g., rebuilding the 2009 dataset) can affect both model performance and explanation outcomes.

By Praveena Padi, Arun Morampudi, Ujval Sai Gopal Irrinki, Pradeep Kumar Dolabehera Kakitapelli
arXiv AI
Aug 5

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

arXiv:2608. 03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).

By Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon