arXiv:2603. 25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity).
By Jon-Paul Cacioli
arXiv:2608.03854v4 Announce Type: replace
Abstract: Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability...
By Anton Rasmussen, Hong Qin
arXiv:2607. 26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data.
By Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson
arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.
By Anton Rasmussen, Hong Qin
The paper introduces an ability‑residual decoupled framework for affective cognitive diagnosis, which first isolates unmodeled cognitive residuals—such as item calibration bias, concept bias, and student‑concept deviations—using student, item, concept, student‑concept, and low‑rank student‑item components. It then applies an affective module that modulates guess/slip effects, with a Q‑matrix‑constrained concept residual attention mechanism to aggregate only item‑relevant concept residuals. Experiments on multiple datasets and backbones demonstrate improved response prediction and better affect alignment, while ablation and analysis studies show that the residual modeling reduces cognitive contamination in the affective branch and enhances robustness and accuracy.
By Boyuan Zhao, Meng Ye
The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.
By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
arXiv:2607. 29093v1 Announce Type: cross Abstract: Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement.
By Saurabh Ranjan, Mukesh Makwana, Konstantina Sokratous, Brian Odegaard
arXiv:2606. 32032v1 Announce Type: cross Abstract: Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes.
By Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona, Idan Szpektor, Arman Cohan
The paper introduces Disengagement-Aware Student Simulators (DAS2), a protocol that models five learner-engagement states—engaged, gaming, wheel-spinning, off-task, and mixed—to evaluate AI tutor performance before deployment. Using annotated tutoring sessions from ASSISTments09, DAS2’s rule-based labels matched human consensus in 81% of cases, and conditioning simulations on intended states narrowed the correctness-rate gap between simulated and authentic sessions for gaming and wheel-spinning behaviors. The study also compares five AI tutors across these states, finding stable relative rankings but state-specific performance differences, and notes that automated evaluation does not fully align with human judgment.
By Xianghui Meng, Jionghao Lin
The study audited seven public educational prediction datasets using four pre‑modeling reliability checks—baseline gap, split instability, null separation, and metadata adequacy under group‑aware holdout. Only three datasets passed all checks; the others failed either group‑aware generalization tests or lacked necessary provenance metadata. The audit revealed that cross‑group fragility, rather than weak iid performance, was the dominant failure mode, and that increasing model complexity did not resolve these structural issues.
By Yan Ma, Lizhuo Zhang
arXiv:2606. 18617v1 Announce Type: cross Abstract: There exist numerous tutor training platforms.
By Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Clara Brandt, Conrad Borchers, Kenneth R. Koedinger
The paper introduces a validation protocol for knowledge‑tracing models that jointly assesses predictive performance, explanation stability, and faithfulness. Using engineered behavioral features from ASSISTments data, the authors compare an XGBoost model explained with TreeSHAP against four deep‑learning baselines, finding comparable predictive accuracy when information is matched and demonstrating that TreeSHAP rankings are stable and impactful. The study highlights how data preprocessing (e.g., rebuilding the 2009 dataset) can affect both model performance and explanation outcomes.
By Praveena Padi, Arun Morampudi, Ujval Sai Gopal Irrinki, Pradeep Kumar Dolabehera Kakitapelli