The paper presents a risk‑adaptive, evidence‑constrained framework for providing feedback in introductory programming courses. Using data from 2,993 failed submissions by 215 students, the authors built models that predict persistent failure and generate four tailored feedback conditions for 136 cases. The framework employs calibrated risk to decide when to intervene, evidence gating to limit feedback content, and a progressive assistance strategy that moves from self‑checks to localized hints.
By Shihao Wang
The paper introduces CodeInsight, a large-scale dataset of over 3 million code submissions from 3,286 undergraduate students in two introductory C++ courses, capturing test‑case outcomes, timestamps, and source code. It presents a benchmark that evaluates various modeling approaches—including a Recurrent State Space Model and an LLM‑based predictor—on their ability to predict iterative problem‑solving dynamics such as performance changes and error persistence. The study finds that the RSSM outperforms other models on most courses, while the LLM generates full submissions but with lower predictive accuracy, suggesting it functions more as a generative solver than a behavior predictor.
By Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo, Nick Haber
arXiv:2606. 12425v1 Announce Type: cross Abstract: Active learning is widely recognized as an effective approach for improving learning outcomes in introductory programming courses.
By Muntasir Hoq, Griffin Pitts, Bradford Mott, Seung Lee, Jessica Vandenberg, Shuyin Jiao, Narges Norouzi, James Lester, Bita Akram
arXiv:2606. 18617v1 Announce Type: cross Abstract: There exist numerous tutor training platforms.
By Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Clara Brandt, Conrad Borchers, Kenneth R. Koedinger
arXiv:2609.39957v1 Announce Type: cross
Abstract: Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and incre...
By Jiangrui Zhao, Chenglong Li, Meng Zhang, Xiaoting Du
arXiv:2608. 12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort.
By Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom)
arXiv:2606. 07544v1 Announce Type: cross Abstract: Middle school is a key window for building core academic skills and the learning routines students carry into later grades, yet many students still fall behind because help is often limited and comes too late, after they have already been stuck for a while.
By Misan Paul Etchie, Taiwo Olutosin
Causal diagnostic models must explain how conclusions follow from evidence because diagnoses guide repairs and treatments. Yet serious cases are scarce, records rarely contain reasoning paths, and data transfer poorly across configurations, complicating local deployment.
arXiv:2607. 04412v1 Announce Type: new Abstract: Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals.
By Yujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song
The paper introduces Evidence‑Diagnosed Intervention Training (EDIT), a two‑phase framework designed to improve rubric‑faithful grading by large language models. EDIT‑SFT first identifies problematic reasoning steps using internal model signals—posterior belief over the final mark and input‑grounding scores—and revises only those steps with rubric checklists. EDIT‑RL then calibrates the grader with belief‑guided reward shaping, penalising harmful belief drifts while encouraging useful exploration. Experiments on two real‑world, multi‑subject grading benchmarks show that EDIT consistently outperforms strong supervised fine‑tuning and reinforcement learning baselines, with ablation studies confirming the importance of internal‑state diagnostics.
By Zhihao Wu, Linhai Zhang, Taiyi Wang, Runcong Zhao, Peter Andrews, Cesare Aloisi, Yulan He
arXiv:2608. 16831v1 Announce Type: new Abstract: Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations.
By Minh-Ha Nguyen, Cathy Shyr
arXiv:2509.05346v3 Announce Type: replace
Abstract: While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how thei...
By Bo Yuan, Jiazi Hu