The study investigates how a single mastery threshold can produce divergent outcomes across different knowledge tracing (KT) models. By evaluating six KT models on four datasets with thresholds ranging from 0.50 to 0.99, the authors find that Bayesian Knowledge Tracing (BKT) is relatively insensitive to threshold changes, whereas neural models become increasingly selective as thresholds rise. The optimal threshold varies widely across models and instructional settings, and stricter thresholds can disproportionately limit advancement for weaker students.
By Xianghui Meng, Yujing Zhang, Jionghao Lin
arXiv:2507. 08150v4 Announce Type: replace-cross Abstract: Accurate uncertainty quantification is critical for reliable predictive modeling.
By Ilia Azizi, Juraj Bodik, Jakob Heiss, Bin Yu
The paper introduces XConf, an experiential confidence estimator that augments a language model’s current inference with a record of its past graded episodes. By recalling similar past tasks and reflecting on past outcomes, XConf generates confidence scores without accessing logits or updating weights, achieving superior discrimination and calibration across diverse benchmarks. The method demonstrates significant gains in selective prediction, improving success rates on agent tasks by up to 8.7 points.
By Caiqi Zhang, Xiaochen Zhu, Chengzu Li, Yulong Chen, Dharshan Kumaran, Nigel Collier
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2607. 28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias.
By Plawan Kumar Rath
arXiv:2603. 02830v2 Announce Type: replace-cross Abstract: Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective interventions.
By Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud, Simon Woodhead
arXiv:2604. 24827v2 Announce Type: replace-cross Abstract: Closed-source frontier labs do not disclose parameter counts.
By Bojie Li
arXiv:2608. 05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human.
By Jianru Shen
The study audits demographic bias across four deep knowledge tracing architectures—DKT, DKVMN, SAKT, and AKT—using two large public datasets (Eedi and OULAD). It finds that bias is context‑dependent: socioeconomic bias is significant on Eedi, while gender bias appears on OULAD for most models. The most accurate model, AKT, also exhibits the greatest bias, and standard mitigation techniques such as reweighting and adversarial debiasing fail to reduce bias without sacrificing accuracy.
By Dang Quang Minh, Nguyen Dung Son, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself.
whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."
By Rohit Patel, Alexandre Rezende, Steven McClain
arXiv:2607. 25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items.
By Juan Francisco, Mandujano Reyes
arXiv:2603. 25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity).
By Jon-Paul Cacioli