arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
arXiv:2607. 01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps.
By Aastha Sapkota, M. G. Sarwar Murshed
arXiv:2606. 08400v1 Announce Type: cross Abstract: Graduate-level research reading report assessment creates a substantial labor burden for educators.
By Qilin Zhou, Zhuo Wang, Yue Li, W. K. Chan
arXiv:2606. 10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined.
By Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li, Jianlong Chen, Xiangfeng Wang
arXiv:2606. 05180v1 Announce Type: cross Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced.
By Ivo Bueno, Babette B\"uhler, Philipp Stark, Tim F\"utterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
arXiv:2608. 15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities.
By Alona Strugatski, Licol Zeinfeld, Giora Alexandron
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
arXiv:2608. 12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort.
By Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom)
arXiv:2606. 05983v1 Announce Type: new Abstract: Generative AI makes answers easy and understanding hard, and uncritical use invites cognitive offloading.
By Alexander Apartsin, Yehudit Aperstein
arXiv:2606. 17507v1 Announce Type: new Abstract: Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment.
By Xiwei Xu, Chen Wang, Jacky Jiang, Phil Yang, Qian Fu, Mohan Dhall, Wenjie Zhang, Liming Zhu
arXiv:2606. 16206v1 Announce Type: new Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support.
By Junyi Yao, Zihao Zheng, Baichuan Li
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited.