arXiv AI

LearnOpt: Recovering the Latent Cognitive Structure of Standardized Examinations via Knowledge Graphs and Constrained Optimization

arXiv:2606. 15349v1 Announce Type: cross Abstract: Standardized examinations are typically treated as uniform syllabus coverage problems.

arXiv AI
Aug 25

GIM: Evaluating models via tasks that integrate multiple cognitive domains

The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."

By Rohit Patel, Alexandre Rezende, Steven McClain
arXiv Machine Learning
1d ago

One Mastery Threshold Does Not Fit All Knowledge Tracing Models

The study investigates how a single mastery threshold can produce divergent outcomes across different knowledge tracing (KT) models. By evaluating six KT models on four datasets with thresholds ranging from 0.50 to 0.99, the authors find that Bayesian Knowledge Tracing (BKT) is relatively insensitive to threshold changes, whereas neural models become increasingly selective as thresholds rise. The optimal threshold varies widely across models and instructional settings, and stricter thresholds can disproportionately limit advancement for weaker students.

By Xianghui Meng, Yujing Zhang, Jionghao Lin
arXiv Computation and Language
Aug 27

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

CaSKG introduces a counterfactual‑causal skill graph framework that calibrates procedural relations before retrieval, building a high‑recall directed candidate graph from semantic, lexical, input/output, and structural evidence and refining it with repair evidence and optional LLM judgment. The framework applies direction‑conditioned textual counterfactual probes—removing, substituting, and reordering skill pairs—to aggregate evidence with Bayesian smoothing, producing a state‑filtered weighted graph for task‑conditioned expansion. Evaluated across six LLM backbones on ALFWorld and ScienceWorld, CaSKG outperforms existing Graph‑of‑Skills methods, improving macro‑average scores and reducing mean environment steps while preserving essential skill dependencies.

By Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang
arXiv AI
Jun 19

Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

arXiv:2606. 19469v1 Announce Type: new Abstract: Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproducible way to measure how completely they cover the current guidelines and how that coverage shifts when the guidelines are restructured.

By Sherzod Turaev, Mary John, Saja Aldabet, Mamoun Awad, Nazar Zaki, Khaled Shuaib
arXiv AI
Aug 17

TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

arXiv:2608. 13708v1 Announce Type: cross Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula.

By Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha, Jubayer Al Mahmud
arXiv Computation and Language
Sep 17

PersonaPath: Towards Knowledge-Centric Personalized Learning Path Planning

PersonaPath is a new benchmark for knowledge‑centric personalized learning path planning, pairing 2,000 learner personas with a hierarchical knowledge graph of 347 textbooks, 1,751 units, and 4,092 concepts across 77 subjects. The study evaluates large language models on this benchmark, finding that even the best model achieves only a 29.5% final pass rate in Basic Education and fails to exceed 44.7% in tailoring paths to individual learners, highlighting a significant adaptivity gap. This work underscores the challenge of moving beyond exercise‑centric recommendation toward goal‑oriented, curriculum‑scale guidance.

By Yu Liu, Zeming Liu, Tianle Zhang, Zihao Cheng, Yuhang Guo, Kehai Chen, Min Zhang, Yunhong Wang, Haifeng Wang
arXiv AI
Aug 5

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

arXiv:2608. 03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).

By Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon
arXiv AI
Jul 3

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.

By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
arXiv Machine Learning
Aug 19

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

The study clusters 5,000 EdNet-KT3 learners into eight study‑strategy groups based on early‑session behaviors such as resource use, revision, video watching, and problem practice. These clusters predict later engagement metrics—like continued practice and session completion—but do not reliably forecast later unassisted accuracy or mastery. The findings suggest that behavioral clustering captures learning styles and engagement patterns rather than knowledge gains.

By Qingchuan Lyu, Yingxin Li, Albert Yang
arXiv Computation and Language
Sep 24

When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA

The study evaluates whether learned context planning can outperform strong retrieval methods in long-context multiple-choice question answering. Using 503 LongBench-v2 MCQ questions and a Qwen2.5-7B-Instruct model, the planner—trained on outcome-selected traces—achieves lower accuracy than anchored hybrid retrieval and BM25 across various character budgets. Even with tight budgets, the planner only marginally improves or matches retrieval, indicating that learned planning provides a weak relevance signal rather than a replacement for robust retrieval.

By Yingrui Li, Han Chen