Hugging Face Trending Papers

A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education

Read the original on Hugging Face Trending Papers →

The paper presents a risk‑adaptive, evidence‑constrained framework that uses learning analytics to provide personalized feedback in introductory programming. By training models on 2,993 failed‑submission states from 215 students, the authors predict persistent failure and generate four tailored feedback conditions for 136 cases. A calibrated risk policy selects interventions for 17.8% of eligible states, capturing 25.2% of persistent failures, and the framework ensures that generated messages contain all required components after evidence gating.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 25

A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education

The paper presents a risk‑adaptive, evidence‑constrained framework for providing feedback in introductory programming courses. Using data from 2,993 failed submissions by 215 students, the authors built models that predict persistent failure and generate four tailored feedback conditions for 136 cases. The framework employs calibrated risk to decide when to intervene, evidence gating to limit feedback content, and a progressive assistance strategy that moves from self‑checks to localized hints.

By Shihao Wang
arXiv Computation and Language
Sep 2

A Dataset for Modeling Iterative Problem-Solving

The paper introduces CodeInsight, a large-scale dataset of over 3 million code submissions from 3,286 undergraduate students in two introductory C++ courses, capturing test‑case outcomes, timestamps, and source code. It presents a benchmark that evaluates various modeling approaches—including a Recurrent State Space Model and an LLM‑based predictor—on their ability to predict iterative problem‑solving dynamics such as performance changes and error persistence. The study finds that the RSSM outperforms other models on most courses, while the LLM generates full submissions but with lower predictive accuracy, suggesting it functions more as a generative solver than a behavior predictor.

By Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo, Nick Haber
arXiv AI
Aug 14

Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection

arXiv:2608. 12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort.

By Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom)