The paper introduces a human‑in‑the‑loop framework for AI‑assisted scoring of short written responses in a large‑scale national assessment. Using data from two recent test editions with about 5,000 student responses each, the authors validate that AI‑generated scores align moderately to highly with human raters across multiple rubric dimensions. The framework includes a correction workflow that flags cases needing human review, thereby reducing manual workload while maintaining assessment quality.
By Mar\'ia Eugenia Curi, Germ\'an Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adri\'an Silveira, Andr\'es Peri
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
By Zewei Tian, Alex Liu, Lief Esbenshade, Michael Xiao, Zachary Zhang, Yulia L\'apicus, Thomas Han, Kevin He, Min Sun
arXiv:2607. 02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation.
By Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
arXiv:2601.08654v3 Announce Type: replace
Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the...
By Yihan Hong, Huaiyuan Yao, Bolin Shen, Wanpeng Xu, Hua Wei, Yushun Dong
arXiv:2504.02323v5 Announce Type: replace
Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored variou...
By Clayton Cohn, Ashwin T S, Naveeduddin Mohammed, Gautam Biswas
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
arXiv:2606. 24839v1 Announce Type: new Abstract: Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics.
By Tian Zheng, Kai-Tai Hsu
arXiv:2607. 15829v1 Announce Type: cross Abstract: Automated essay scoring (AES) enables scalable assessment and timely feedback but remains challenged by transformer input-length limitations, which can cause information loss when processing long essays.
By Haowei Hua
arXiv:2606. 08400v1 Announce Type: cross Abstract: Graduate-level research reading report assessment creates a substantial labor burden for educators.
By Qilin Zhou, Zhuo Wang, Yue Li, W. K. Chan
The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.
By Lukas Gehring, Benjamin Paa{\ss}en
arXiv:2607. 14524v1 Announce Type: new Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays.
By Adnan Labib, Yixuan Huang, Jiahui Wu, John Maurice Gayed, Zheng Yuan, Qiao Wang
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
By Ali Keramati, Mark Warschauer