arXiv AI

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

arXiv:2607. 26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data.

arXiv AI
Jul 3

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.

By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
arXiv AI
Sep 21

Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis

The paper introduces an ability‑residual decoupled framework for affective cognitive diagnosis, which first isolates unmodeled cognitive residuals—such as item calibration bias, concept bias, and student‑concept deviations—using student, item, concept, student‑concept, and low‑rank student‑item components. It then applies an affective module that modulates guess/slip effects, with a Q‑matrix‑constrained concept residual attention mechanism to aggregate only item‑relevant concept residuals. Experiments on multiple datasets and backbones demonstrate improved response prediction and better affect alignment, while ablation and analysis studies show that the residual modeling reduces cognitive contamination in the affective branch and enhances robustness and accuracy.

By Boyuan Zhao, Meng Ye
Hugging Face Trending Papers
Jun 17

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While various existing works have explored whether large language models (LLMs) can estimate item difficulty, it remains unclear whether they can capture item discrimination.