arXiv AI

Argus: Academic Integrity in the Era of Generative AI

arXiv Machine Learning
Sep 14

Limits of LLM Text Detectors in Education

The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.

By Lukas Gehring, Benjamin Paa{\ss}en
arXiv Computation and Language
Sep 2

A Dataset for Modeling Iterative Problem-Solving

The paper introduces CodeInsight, a large-scale dataset of over 3 million code submissions from 3,286 undergraduate students in two introductory C++ courses, capturing test‑case outcomes, timestamps, and source code. It presents a benchmark that evaluates various modeling approaches—including a Recurrent State Space Model and an LLM‑based predictor—on their ability to predict iterative problem‑solving dynamics such as performance changes and error persistence. The study finds that the RSSM outperforms other models on most courses, while the LLM generates full submissions but with lower predictive accuracy, suggesting it functions more as a generative solver than a behavior predictor.

By Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo, Nick Haber
arXiv AI
Sep 21

Self-Explanation Tutor for Active Study of CS1 Worked Examples

The paper presents ESSE, a self‑explanation tutor that uses a large language model to give immediate feedback on students’ line‑by‑line explanations of introductory programming worked examples. It evaluates the LLM’s judgments against a domain expert and a crowd of non‑experts, finding that the model is reliable enough to serve as the tutor’s assessment engine. In an introductory Java course, the tutor’s feedback encourages students to persist, improves the completeness and conceptual depth of their explanations, and shows evidence of learning.

By Arun-Balajiee Lekshmi-Narayanan, Mohammad Hassany, Kamil Akhuseyinoglu, Rully Hendrawan, Peter Brusilovsky