arXiv AI

Dynamic Evidence Collection Ecosystem for Assessment Integrity and Authentic Competence

arXiv:2608. 16016v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GenAI) can produce high-quality essays, code, and design artefacts, challenging the validity of conventional assessments that rely on single-point submissions and product-only grading.

arXiv AI
Aug 14

Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection

arXiv:2608. 12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort.

By Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom)
arXiv AI
Sep 25

Beyond Simple Input-Output Assessment Tasks: Leveraging Automated Programming Assessment for Non-Trivial Courses

The article discusses how machine learning exercises can be designed for automated assessment tools, framing them as deterministic input-output tasks. It emphasizes that this approach does not create a new grading system but enables existing platforms (e.g., VPL for Moodle, Codeforces, MOJ) to support AI education more effectively. The authors argue that integrating theory with practice through such exercises can foster dynamic, interactive AI courses.

By Artur Jordao
arXiv AI
Jun 9

AI-Integrated Learning Management System for Middle School: A Longitudinal Study of Learning Outcomes Through High School and Beyond

arXiv:2606. 07544v1 Announce Type: cross Abstract: Middle school is a key window for building core academic skills and the learning routines students carry into later grades, yet many students still fall behind because help is often limited and comes too late, after they have already been stuck for a while.

By Misan Paul Etchie, Taiwo Olutosin
arXiv AI
Sep 12

Generative AI performance in core undergraduate mathematics: a curriculum-level case study

The study examines how generative AI tools like ChatGPT perform on typical first‑year undergraduate mathematics assessment questions. By generating, transcribing, and blind‑marking AI responses to eight assessments covering the entire curriculum, the authors find that AI attains a first‑class level of performance, with consistency across modules that exceeds that of students in invigilated exams. The results suggest a need to redesign mathematics assessments to address the impact of generative AI.

By Benjamin J. Walker, Nikoleta Kalaydzhieva, Beatriz Navarro Lameda, Ruth A. Reynolds
arXiv AI
Sep 18

greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI

The paper introduces greCAPTCHA, a proctored assessment designed to gauge authors’ understanding of their own research manuscripts by measuring their capacity to verify content. It generates questions at multiple levels of comprehension and produces an evaluative report based on responses. A prototype study with 31 researchers showed that automated scores could predict authorship with an AUC of 0.90, and participants reported positive experiences and constructive feedback for future deployment.

By Justin Payan, B\'alint Gyevn\'ar, Atoosa Kasirzadeh, Nihar B. Shah