arXiv:2511. 21692v3 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation.
By Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach
arXiv:2606. 28186v1 Announce Type: cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction.
By Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou
arXiv:2608. 06933v1 Announce Type: cross Abstract: Today, we improve models by training and evaluating them on problems at the frontier of their abilities.
By Sarah Pratt, Jae Sung Park, Scott Geng, Ali Farhadi
arXiv:2601. 18778v3 Announce Type: replace Abstract: RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal.
By Shobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja, Yann Ollivier, Julia Kempe
arXiv:2601. 09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models.
By Jiali Cheng, Ziheng Chen, Chirag Agarwal, Hadi Amiri
The paper investigates why curriculum learning—ordering training data from easy to hard—varies in effectiveness across reasoning tasks. By studying optimization dynamics, the authors introduce Relative Transfer, a measure of cross‑difficulty knowledge transfer, and use it to create Transfer‑aware Dynamic Curriculum Sampling (TDCS). Experiments show TDCS outperforms existing scheduling strategies on multiple reasoning benchmarks, offering a unified optimization‑based explanation for curriculum learning.
By Zhikai Ding, Ziyi Ye
arXiv:2607. 28634v1 Announce Type: cross Abstract: The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments.
By Xinyi Wang, Hong Jiao, Ming Li, Sydney Peters, Hanna Choi, Tianyi Zhou, Qingshu Xu
The paper introduces CodeInsight, a large-scale dataset of over 3 million code submissions from 3,286 undergraduate students in two introductory C++ courses, capturing test‑case outcomes, timestamps, and source code. It presents a benchmark that evaluates various modeling approaches—including a Recurrent State Space Model and an LLM‑based predictor—on their ability to predict iterative problem‑solving dynamics such as performance changes and error persistence. The study finds that the RSSM outperforms other models on most courses, while the LLM generates full submissions but with lower predictive accuracy, suggesting it functions more as a generative solver than a behavior predictor.
By Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo, Nick Haber
arXiv:2606. 17706v1 Announce Type: cross Abstract: Curriculum learning couples two design choices, how samples are scored by difficulty and how harder samples are paced into training, making it difficult to attribute observed gains to either component.
By Savini Kommalage, Sanka Mohottala, Asiri Gawesha, Dulara Madhusanka, Menan Velayuthan, Dharshana Kasthurirathna, Mahima Milinda Alwis Weerasinghe, Charith Abhayaratne
arXiv:2602. 10014v3 Announce Type: replace Abstract: Iterative self-improvement fine-tunes an autoregressive large language model (LLM) on reward-verified outputs generated by the LLM itself.
By Chenruo Liu, Yijun Dong, Yiqiu Shen, Qi Lei
arXiv:2608. 01522v1 Announce Type: new Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling beyond which additional data yields no further improvement.
By Longtian Bao, Jianyou Wang, Yang Zhang, Youze Zheng, Ramamohan Paturi
arXiv:2510. 05969v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an essential capability for adaptive reasoning and efficient resource allocation.
By Sunbowen Lee, Qingyu Yin, Chak Tou Leong, Jialiang Zhang, Yicheng Gong, Shiwen Ni, Min Yang, Xiaoyu Shen