arXiv AI

Effects of Varying LLM Access on Essay Writing Behavior

arXiv:2606. 00250v1 Announce Type: cross Abstract: Investigating the degree to which large language models (LLMs) affect teaching and learning in universities can help identify strategies for integrating LLMs in a way that supports, rather than undermines, student learning outcomes.

arXiv AI
Aug 28

How LLMs Distort Our Written Language

Large language models (LLMs) are widely used to assist writing, but this study shows they alter both tone and meaning of human text. A user study found that heavy LLM use increased neutral essays by nearly 70% and made writers feel less creative and less in their voice. Even when prompted to make only grammar edits, LLMs changed the semantic content of essays and produced AI-generated scientific reviews that were less focused on clarity and significance and scored higher on average.

By Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
arXiv AI
Sep 17

Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming

The study investigates whether AI assistance leaves a temporal fingerprint in writing and programming tasks. By analyzing keystroke-level data from three corpora, the authors find that AI contributions appear in distinct bursts and that temporal patterns can almost perfectly distinguish wholesale delegation from authentic work, though ordinary collaboration remains hard to detect. The research suggests that process visibility could serve as a basis for academic integrity checks.

By Eduardo Davalos, Yike Zhang
arXiv Machine Learning
Sep 4

SWIM: Student Writing Simulation via Proficiency-Conditioned Generation

The paper introduces SWIM, a task that frames student writing simulation as proficiency‑conditioned essay generation. It evaluates prompting, supervised fine‑tuning, and reinforcement learning for aligning generated essays with student proficiency profiles, using automated essay scoring as a metric. Results show that prompting alone offers limited control, while supervised fine‑tuning and reinforcement learning significantly improve alignment across content, lexical, grammatical, and organizational traits, though low‑proficiency writing remains difficult to replicate.

By Heejin Do, Jakub Kontak, Mrinmaya Sachan
arXiv AI
Sep 7

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

The study explores how undergraduate computing students in Saudi Arabia perceive AI‑generated writing feedback when they are explicitly told that ChatGPT, not a human instructor, produced the score and comments. Through qualitative reflections, four themes emerged: students found the feedback useful for surface‑level revisions, recognized AI’s contextual and pedagogical limits, trusted the feedback conditionally—separating its utility from its authority—and reaffirmed the human instructor’s role as the ultimate grading authority. The findings highlight a clear distinction students make between feedback usefulness and evaluative authority, treating them as separate judgments rather than opposing ends of a single approval scale.

By Rayed AlGhamdi
arXiv AI
Sep 17

The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses

The study investigates how generative AI (GenAI) affects student learning in AI-related courses, using survey data from 118 students across 12 courses. Four distinct user clusters were identified—high-use, light-use, and two moderate-use groups—each showing varying benefits and reliance patterns. The research highlights that early reliance, evaluation literacy, and instructor policies significantly influence perceived academic benefits and negative impacts, underscoring the need for institutional policies to address inequities in AI use.

By Lydia Manikonda, Mei Si, Sirajam Munira, Oshani Seneviratne, Kristin Bennett
arXiv AI
Sep 24

Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing

The paper examines whether state‑of‑the‑art large language models (LLMs) produce feedback that aligns with expert teachers’ pedagogical practices, focusing on feedback type and adaptivity. Using a refined taxonomy of seven feedback focus types, the authors annotate and compare teacher and LLM‑generated feedback from three university writing courses, creating the FeedType benchmark. Their analysis shows that while most LLMs cover many feedback types, they do not match teachers’ distribution patterns or adaptive behavior across draft stages and student performance levels.

By Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman