Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.
By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych
arXiv:2609.36544v1 Announce Type: cross
Abstract: Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through whic...
By Divyansh Chandarana, Sandipan De, Vivek Gupta
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv:2608. 07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc.
By Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung
arXiv:2606. 01020v1 Announce Type: new Abstract: Identifying logical fallacies in everyday discourse is challenging for many people.
By Minjing Shi, Junling Wang, Jingwei Ni, Sankalan Pal Chowdhury, Mrinmaya Sachan
arXiv:2504.02323v5 Announce Type: replace
Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored variou...
By Clayton Cohn, Ashwin T S, Naveeduddin Mohammed, Gautam Biswas
arXiv:2606. 17507v1 Announce Type: new Abstract: Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment.
By Xiwei Xu, Chen Wang, Jacky Jiang, Phil Yang, Qian Fu, Mohan Dhall, Wenjie Zhang, Liming Zhu
arXiv:2607. 03303v1 Announce Type: new Abstract: While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use in education.
By Jerome Brender, Laila El-Hamamsy, Kim Uittenhove, Aitor Perez, Patrick Jermann, Francesco Mondada, Engin Bumbacher
arXiv:2608. 03952v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners.
By Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, Huamin Qu, Zixin Chen
The paper introduces a prompt‑engineering framework that personalizes large language model (LLM) teaching assistants across disciplines by tailoring responses to six learner‑specific dimensions, creating 96 distinct learner profiles. It also analyzes student queries through Bloom’s Taxonomy to gauge cognitive complexity, encoding both learner attributes and cognitive assessments into structured prompts that condition the LLM without retraining. Experiments using NLP metrics and a small human study demonstrate that this approach yields perceptible differences in response style and structure, with statistical evidence linking specific learner attributes to measurable changes.
By Saptarshi Basu, Sandeep Kakar, Ashok Goel
arXiv:2608. 10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them.
By Rose Niousha, Minwoo Kang, Narges Norouzi
The paper examines whether state‑of‑the‑art large language models (LLMs) produce feedback that aligns with expert teachers’ pedagogical practices, focusing on feedback type and adaptivity. Using a refined taxonomy of seven feedback focus types, the authors annotate and compare teacher and LLM‑generated feedback from three university writing courses, creating the FeedType benchmark. Their analysis shows that while most LLMs cover many feedback types, they do not match teachers’ distribution patterns or adaptive behavior across draft stages and student performance levels.
By Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman