arXiv AI

An Integrated Machine Learning and Hierarchical Variance Decomposition Pipeline for Student Performance Prediction and Metacognitive Calibration on Multi-Signal Telemetry

arXiv:2606. 28881v1 Announce Type: cross Abstract: Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring systems.

arXiv AI
Sep 21

Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis

The paper introduces an ability‑residual decoupled framework for affective cognitive diagnosis, which first isolates unmodeled cognitive residuals—such as item calibration bias, concept bias, and student‑concept deviations—using student, item, concept, student‑concept, and low‑rank student‑item components. It then applies an affective module that modulates guess/slip effects, with a Q‑matrix‑constrained concept residual attention mechanism to aggregate only item‑relevant concept residuals. Experiments on multiple datasets and backbones demonstrate improved response prediction and better affect alignment, while ablation and analysis studies show that the residual modeling reduces cognitive contamination in the affective branch and enhances robustness and accuracy.

By Boyuan Zhao, Meng Ye
arXiv Machine Learning
Aug 19

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.

By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
arXiv Machine Learning
Sep 14

Simulating Disengaged Students to Evaluate LLM-based Tutors

The paper introduces Disengagement-Aware Student Simulators (DAS2), a protocol that models five learner-engagement states—engaged, gaming, wheel-spinning, off-task, and mixed—to evaluate AI tutor performance before deployment. Using annotated tutoring sessions from ASSISTments09, DAS2’s rule-based labels matched human consensus in 81% of cases, and conditioning simulations on intended states narrowed the correctness-rate gap between simulated and authentic sessions for gaming and wheel-spinning behaviors. The study also compares five AI tutors across these states, finding stable relative rankings but state-specific performance differences, and notes that automated evaluation does not fully align with human judgment.

By Xianghui Meng, Jionghao Lin
arXiv Machine Learning
Sep 25

Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets

The study audited seven public educational prediction datasets using four pre‑modeling reliability checks—baseline gap, split instability, null separation, and metadata adequacy under group‑aware holdout. Only three datasets passed all checks; the others failed either group‑aware generalization tests or lacked necessary provenance metadata. The audit revealed that cross‑group fragility, rather than weak iid performance, was the dominant failure mode, and that increasing model complexity did not resolve these structural issues.

By Yan Ma, Lizhuo Zhang
arXiv Machine Learning
Sep 25

Stable and Faithful Explanations for Knowledge Tracing

The paper introduces a validation protocol for knowledge‑tracing models that jointly assesses predictive performance, explanation stability, and faithfulness. Using engineered behavioral features from ASSISTments data, the authors compare an XGBoost model explained with TreeSHAP against four deep‑learning baselines, finding comparable predictive accuracy when information is matched and demonstrating that TreeSHAP rankings are stable and impactful. The study highlights how data preprocessing (e.g., rebuilding the 2009 dataset) can affect both model performance and explanation outcomes.

By Praveena Padi, Arun Morampudi, Ujval Sai Gopal Irrinki, Pradeep Kumar Dolabehera Kakitapelli