arXiv:2607. 00274v1 Announce Type: cross Abstract: Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive.
By Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho, Diane Litman, Xiang Lorraine Li
Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.
By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych
arXiv:2606. 30774v1 Announce Type: new Abstract: We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone.
By Bart{\l}omiej Cupia{\l}, Jan {\L}ojek, Miko{\l}aj Garstecki, Szymon Pob{\l}ocki, Alicja Ziarko, Piotr Mi{\l}o\'s
arXiv:2608. 03952v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners.
By Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, Huamin Qu, Zixin Chen
The paper introduces SWIM, a task that frames student writing simulation as proficiency‑conditioned essay generation. It evaluates prompting, supervised fine‑tuning, and reinforcement learning for aligning generated essays with student proficiency profiles, using automated essay scoring as a metric. Results show that prompting alone offers limited control, while supervised fine‑tuning and reinforcement learning significantly improve alignment across content, lexical, grammatical, and organizational traits, though low‑proficiency writing remains difficult to replicate.
By Heejin Do, Jakub Kontak, Mrinmaya Sachan
arXiv:2606. 05180v1 Announce Type: cross Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced.
By Ivo Bueno, Babette B\"uhler, Philipp Stark, Tim F\"utterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
arXiv:2504.02323v5 Announce Type: replace
Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored variou...
By Clayton Cohn, Ashwin T S, Naveeduddin Mohammed, Gautam Biswas
arXiv:2608. 11625v2 Announce Type: replace Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedback productively.
By Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Ga\v{s}evi'c, Naomi Winstone, Hassan Khosravi
arXiv:2608. 07509v1 Announce Type: cross Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers.
By Fares Fawzi, Jiaxu Zhao, Tanya Nazaretsky, Tanja K\"aser
arXiv:2608. 11625v1 Announce Type: new Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedback productively.
By Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Ga\v{s}evi'c, Naomi Winstone, Hassan Khosravi
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv:2605. 05598v2 Announce Type: replace Abstract: The proliferation of large language models (LLMs) in educational settings has paradoxically undermined the cognitive processes they purport to support.
By Ran Bi, Shiyao Wei, Yuanyiyi Zhou