Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.
By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych
The paper examines whether state‑of‑the‑art large language models (LLMs) produce feedback that aligns with expert teachers’ pedagogical practices, focusing on feedback type and adaptivity. Using a refined taxonomy of seven feedback focus types, the authors annotate and compare teacher and LLM‑generated feedback from three university writing courses, creating the FeedType benchmark. Their analysis shows that while most LLMs cover many feedback types, they do not match teachers’ distribution patterns or adaptive behavior across draft stages and student performance levels.
By Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman
arXiv:2607. 14524v1 Announce Type: new Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays.
By Adnan Labib, Yixuan Huang, Jiahui Wu, John Maurice Gayed, Zheng Yuan, Qiao Wang
arXiv:2609.36544v1 Announce Type: cross
Abstract: Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through whic...
By Divyansh Chandarana, Sandipan De, Vivek Gupta
The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.
By Lukas Gehring, Benjamin Paa{\ss}en
arXiv:2601. 11541v2 Announce Type: replace-cross Abstract: To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language Models (LLMs), this study evaluates a locally hosted Small Language Model (SLM).
By Suqing Liu, Runlong Ye, Christopher Eaton, Bogdan Simion, Michael Liut