The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
The paper introduces SWIM, a task that frames student writing simulation as proficiency‑conditioned essay generation. It evaluates prompting, supervised fine‑tuning, and reinforcement learning for aligning generated essays with student proficiency profiles, using automated essay scoring as a metric. Results show that prompting alone offers limited control, while supervised fine‑tuning and reinforcement learning significantly improve alignment across content, lexical, grammatical, and organizational traits, though low‑proficiency writing remains difficult to replicate.
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
arXiv:2605.30051v2 Announce Type: replace Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as stude...
The paper introduces a cost‑aware framework that treats each prompt type as an arm in a multi‑armed bandit controller, enabling adaptive selection of optimal prompting strategies during inference for automated essay scoring. Experiments on IELTS Writing Task 2 essays demonstrate that this bandit-driven approach achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls by 78.4%. The study also presents the first cost‑reliability learning curves for essay scoring, offering actionable insights for educational technology platforms balancing operational costs against assessment validity.
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited.
arXiv:2607. 14524v1 Announce Type: new Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays.
arXiv:2607. 19219v1 Announce Type: cross Abstract: Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG).
arXiv:2608. 10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them.
The paper introduces Trait-Aware Policy Optimization (TAPO), a post‑training framework for autoregressive models that score essays across multiple traits. TAPO decomposes rewards by sample and trait, integrating global consistency, trait accuracy, format validity, and inter‑trait dependencies, while enriching prompts with trait descriptions. Experiments on various backbone models show TAPO consistently outperforms supervised fine‑tuning and scalar‑reward baselines, proving its effectiveness and transferability for multi‑trait essay scoring.
arXiv:2606. 20152v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly understood.
Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.
arXiv:2608.30033v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human behavior but frequently fail to exhibit realistic cognitive constraints, sufferi...
The study examines whether domain‑adaptive continued pretraining (DAPT) on a learner‑writing corpus (EFCAMDAT) can enhance transformer‑based automated essay scoring (AES) for English proficiency tests. Researchers applied DAPT to BERT, RoBERTa, and DistilBERT and compared the adapted models with their original checkpoints on the FCE and IELTS datasets, evaluating both in‑domain scoring and few‑shot cross‑dataset transfer. Results show that full‑corpus DAPT yields mixed effects, while proficiency‑specific DAPT often outperforms full‑corpus DAPT and sometimes even the non‑adapted baseline, though benefits vary by proficiency composition and encoder architecture and do not consistently transfer across tests.