arXiv Computation and Language By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych

Expos\'ia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Read the original on arXiv Computation and Language →

Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 9

A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

arXiv:2601. 11541v2 Announce Type: replace-cross Abstract: To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language Models (LLMs), this study evaluates a locally hosted Small Language Model (SLM).

By Suqing Liu, Runlong Ye, Christopher Eaton, Bogdan Simion, Michael Liut
arXiv Machine Learning
4d ago

Limits of LLM Text Detectors in Education

The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.

By Lukas Gehring, Benjamin Paa{\ss}en