The study examines how generative AI tools like ChatGPT perform on typical first‑year undergraduate mathematics assessment questions. By generating, transcribing, and blind‑marking AI responses to eight assessments covering the entire curriculum, the authors find that AI attains a first‑class level of performance, with consistency across modules that exceeds that of students in invigilated exams. The results suggest a need to redesign mathematics assessments to address the impact of generative AI.
By Benjamin J. Walker, Nikoleta Kalaydzhieva, Beatriz Navarro Lameda, Ruth A. Reynolds
arXiv:2606. 10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined.
By Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li, Jianlong Chen, Xiangfeng Wang
arXiv:2605. 21629v2 Announce Type: replace-cross Abstract: How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning outcomes?
By Sina Rismanchian, Hasan Uzun, Jeffrey Matayoshi, Eric Cosyn, Eyad Kurd-Misto
arXiv:2509. 13570v2 Announce Type: replace Abstract: With the rapid rise of generative AI in higher education, understanding how students use AI is increasingly important.
By Hannah Klawa, Shraddha Rajpal, Cigole Thomas
arXiv:2608.21391v1 Announce Type: cross
Abstract: In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intelligence (AI),...
By Aditya Johri, Cory Brozina, Akriti Bagale
arXiv:2607. 17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks.
By Luyu Qiu, Jianing Li, Hwanhee Kim, Xiaoyong Wei, Yueyuan Zheng, Janet Hsiao, Lei Chen