arXiv:2603. 25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity).
By Jon-Paul Cacioli
arXiv:2608.03854v4 Announce Type: replace
Abstract: Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability...
By Anton Rasmussen, Hong Qin
arXiv:2607. 26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data.
By Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson
arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.
By Anton Rasmussen, Hong Qin
The paper introduces an ability‑residual decoupled framework for affective cognitive diagnosis, which first isolates unmodeled cognitive residuals—such as item calibration bias, concept bias, and student‑concept deviations—using student, item, concept, student‑concept, and low‑rank student‑item components. It then applies an affective module that modulates guess/slip effects, with a Q‑matrix‑constrained concept residual attention mechanism to aggregate only item‑relevant concept residuals. Experiments on multiple datasets and backbones demonstrate improved response prediction and better affect alignment, while ablation and analysis studies show that the residual modeling reduces cognitive contamination in the affective branch and enhances robustness and accuracy.
By Boyuan Zhao, Meng Ye
The study explores whether combining traditional and digital learning analytics can predict failure in a first‑year CS1 course. Using data from 284 students across four cohorts, the authors identified ten candidate factors and built a logistic regression model that achieved 74.7% accuracy and 0.742 macro F1, with 87% recall for failing students. Weighted academic momentum, basic demographics, and LMS activity emerged as the most predictive features, suggesting that simple digital markers can enable early‑warning systems by week five.
By Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko