arXiv:2608. 16318v1 Announce Type: cross Abstract: Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and explain source code.
By Marina Lepp, Joosep Kaimre
arXiv:2606. 30655v1 Announce Type: cross Abstract: AI-native course assessments in senior computer science courses and related fields should grade students by \emph{AI-resilient skill}: the ability to achieve outcomes beyond a strong AI baseline.
By Anshumali Shrivastava
arXiv:2608. 12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort.
By Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom)
arXiv:2606. 12425v1 Announce Type: cross Abstract: Active learning is widely recognized as an effective approach for improving learning outcomes in introductory programming courses.
By Muntasir Hoq, Griffin Pitts, Bradford Mott, Seung Lee, Jessica Vandenberg, Shuyin Jiao, Narges Norouzi, James Lester, Bita Akram
The study examines the nature of questions students pose to generative AI during two CS2 programming tasks, classifying 830 interactions into 18 categories based on the Graesser taxonomy. Results reveal that a limited set of question types dominates student inquiries and that the distribution of question types shifts significantly as the task progresses.
By Matin Amoozadeh, Amin Alipour
arXiv:2507.12674v3 Announce Type: replace-cross
Abstract: Evaluating Artificial Intelligence (AI) tutor feedback before deployment requires anticipating student engagement, typically assessed through...
By Rose Niousha, Mihran Miroyan, Abigail O'Neill, Joseph E. Gonzalez, Gireeja Ranade, John DeNero, Narges Norouzi
The study examines how generative AI tools like ChatGPT perform on typical first‑year undergraduate mathematics assessment questions. By generating, transcribing, and blind‑marking AI responses to eight assessments covering the entire curriculum, the authors find that AI attains a first‑class level of performance, with consistency across modules that exceeds that of students in invigilated exams. The results suggest a need to redesign mathematics assessments to address the impact of generative AI.
By Benjamin J. Walker, Nikoleta Kalaydzhieva, Beatriz Navarro Lameda, Ruth A. Reynolds
The paper introduces a cost‑aware framework that treats each prompt type as an arm in a multi‑armed bandit controller, enabling adaptive selection of optimal prompting strategies during inference for automated essay scoring. Experiments on IELTS Writing Task 2 essays demonstrate that this bandit-driven approach achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls by 78.4%. The study also presents the first cost‑reliability learning curves for essay scoring, offering actionable insights for educational technology platforms balancing operational costs against assessment validity.
By Olga Manakina, Igor Bogdanov
arXiv:2607. 13041v1 Announce Type: cross Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them.
By Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird
arXiv:2411.02455v3 Announce Type: replace
Abstract: The rapid adoption of generative AI has created new opportunities for teaching, learning, and quality assurance. Existing applications, however, re...
By Bo Yuan, Jiazi Hu, Haimei Zhao
arXiv:2608.21391v1 Announce Type: cross
Abstract: In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intelligence (AI),...
By Aditya Johri, Cory Brozina, Akriti Bagale
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
By Zewei Tian, Alex Liu, Lief Esbenshade, Michael Xiao, Zachary Zhang, Yulia L\'apicus, Thomas Han, Kevin He, Min Sun