arXiv:2606. 06804v1 Announce Type: new Abstract: Digital learning environments record learners' responses to individual items, making it possible to study the development of specific skills rather than overall scores.
By Yawen Ma, Sahoko Ishida, Kate Cain, Gabriel Wallin
arXiv:2607. 26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data.
By Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson
arXiv:2606. 28186v1 Announce Type: cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction.
By Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou
arXiv:2606. 12945v1 Announce Type: new Abstract: Long-running LLM agents accumulate interaction histories far larger than any context window, forcing a standing decision: what to encode deeply, what to forget, and what to retrieve under a fixed memory budget.
By Zhibao Chen, Qian Cheng
arXiv:2607. 01278v1 Announce Type: new Abstract: The research proposes a multilayer Q-matrix-embedded neural network for cognitive diagnosis (M-QCDNet), which integrates the structural interpretability of cognitive diagnostic models (CDMs) with the deep learning neural network (NN).
By Yiyao Yang
arXiv:2601. 02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration.
By Christopher Ormerod
arXiv:2606. 28881v1 Announce Type: cross Abstract: Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring systems.
By Gurdeep Singh Virdee
arXiv:2607. 28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias.
By Plawan Kumar Rath
The paper introduces Process-aware Language Cognitive Diagnosis (PLCD), a framework that replaces traditional ID-based embeddings in Cognitive Diagnosis Models with language-derived structures and response records. PLCD employs large language models to build concept schemas and cognitive process graphs, and uses a Language-to-Cognition Mapper with DA-MoE experts and contrastive learning to map textual evidence into a unified cognitive space. Experiments demonstrate that PLCD outperforms conventional baselines in student performance prediction and shows strong cognitive transfer, improving cold-start robustness and cognitive grounding.
By Minghang Liu, Yuanzhuo Wang, Qiang Qiu, Huawei Shen, Xueqi Cheng
arXiv:2608. 15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities.
By Alona Strugatski, Licol Zeinfeld, Giora Alexandron
Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions.
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
By Delip Rao, Chris Callison-Burch