arXiv AI By Suqing Liu, Runlong Ye, Christopher Eaton, Bogdan Simion, Michael Liut

A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

Read the original on arXiv AI →

arXiv:2601. 11541v2 Announce Type: replace-cross Abstract: To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language Models (LLMs), this study evaluates a locally hosted Small Language Model (SLM).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 3

Expos\'ia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.

By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych
arXiv AI
Sep 24

Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing

The paper examines whether state‑of‑the‑art large language models (LLMs) produce feedback that aligns with expert teachers’ pedagogical practices, focusing on feedback type and adaptivity. Using a refined taxonomy of seven feedback focus types, the authors annotate and compare teacher and LLM‑generated feedback from three university writing courses, creating the FeedType benchmark. Their analysis shows that while most LLMs cover many feedback types, they do not match teachers’ distribution patterns or adaptive behavior across draft stages and student performance levels.

By Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman
arXiv AI
Sep 7

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

The study explores how undergraduate computing students in Saudi Arabia perceive AI‑generated writing feedback when they are explicitly told that ChatGPT, not a human instructor, produced the score and comments. Through qualitative reflections, four themes emerged: students found the feedback useful for surface‑level revisions, recognized AI’s contextual and pedagogical limits, trusted the feedback conditionally—separating its utility from its authority—and reaffirmed the human instructor’s role as the ultimate grading authority. The findings highlight a clear distinction students make between feedback usefulness and evaluative authority, treating them as separate judgments rather than opposing ends of a single approval scale.

By Rayed AlGhamdi