Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited.
The paper introduces HiFTS, a unified autoregressive framework that generates hierarchical chain-of-thought (CoT) feedback before predicting trait-level and holistic scores for multi-trait automated essay scoring. HiFTS distills rubric-grounded CoT feedback from a teacher large language model and trains student models to jointly produce feedback and scores, employing Group Relative Policy Optimization to balance score agreement, calibration, feedback quality, and structural validity. The authors also present CFMS-34, a new Chinese multi-trait AES dataset, and demonstrate that HiFTS achieves strong scoring performance while producing coherent, rubric-aligned feedback on CFMS-34 and ASAP++.
By Shihang Yang, Sanwoo Lee, Ningning Zhao, Yunfang Wu
CriticGen introduces a generation‑aware evaluation framework that generates sample‑specific evaluation dimensions and scoring criteria across categories such as subjective, objective, and self‑derived constraints. These dynamic rubrics produce a score, reason, executable refinement suggestion, and a refined answer, enabling models to diagnose and target flaws in their responses. Experiments show significant gains in rubric quality, score correlation, and actionable feedback, with 73.17% of answers improved and a 93.28% non‑degradation rate.
By Huifang Du, Zecheng Zuo, Sen Wang, Chenghao Fan, Haofen Wang, Yehui Yang
The paper introduces a cost‑aware framework that treats each prompt type as an arm in a multi‑armed bandit controller, enabling adaptive selection of optimal prompting strategies during inference for automated essay scoring. Experiments on IELTS Writing Task 2 essays demonstrate that this bandit-driven approach achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls by 78.4%. The study also presents the first cost‑reliability learning curves for essay scoring, offering actionable insights for educational technology platforms balancing operational costs against assessment validity.
By Olga Manakina, Igor Bogdanov
arXiv:2608.30005v1 Announce Type: new
Abstract: Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific c...
By Fengyu Xie, Yilun Zhao, Bingsen Chen, Arman Cohan, Chen Zhao
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillation often relies on chain-of-thought annotations that are expensive to obtain and may themselves be noisy, incomplete, or partially incorrect; even when the final solution is correct, an imperfect rationale can interfere with learning.
arXiv:2606. 19327v1 Announce Type: new Abstract: Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards.
By Siyi Gu, Jialin Chen, Sophia Zhou, Arman Cohan, Rex Ying
arXiv:2607. 14524v1 Announce Type: new Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays.
By Adnan Labib, Yixuan Huang, Jiahui Wu, John Maurice Gayed, Zheng Yuan, Qiao Wang
arXiv:2607. 25675v1 Announce Type: new Abstract: Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box.
By Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
By Ali Keramati, Mark Warschauer
arXiv:2608. 16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult.
By Huan Zhang, Mingju Chen, Dongxu Zhou, Can Lv, Heng Chang, Sen Cui, Faguo Wu, Shiji Zhou
arXiv:2602. 01747v2 Announce Type: replace-cross Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools.
By Hongseok Choi, Serynn Kim, Wencke Liermann, Jin Seong, Jin-Xia Huang