arXiv Machine Learning

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

arXiv:2607. 28634v1 Announce Type: cross Abstract: The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments.

arXiv AI
Sep 3

Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning

The paper proposes a response‑free method for estimating difficulty of reading‑comprehension multiple‑choice items by fine‑tuning a transformer on item wording. It introduces two extensions to a baseline joint‑encoding model: a component‑wise variant that encodes passage, question, and options separately, and a multi‑task variant that adds a question‑answering auxiliary task. Experiments on a corpus of nearly 30,000 items show that both extensions outperform the baseline, especially the multi‑task variant across all metrics and the component‑wise variant in rank ordering, even with limited training data.

By Jan Net\'ik, Patr\'icia Martinkov\'a
Hugging Face Trending Papers
Jul 8

From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings

Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibration. Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automate the design matrices traditionally specified by hand.

arXiv AI
Jul 21

Probing the Difficulty Perception Mechanism of Large Language Models

arXiv:2510. 05969v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an essential capability for adaptive reasoning and efficient resource allocation.

By Sunbowen Lee, Qingyu Yin, Chak Tou Leong, Jialiang Zhang, Yicheng Gong, Shiwen Ni, Min Yang, Xiaoyu Shen
arXiv AI
2d ago

Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.

By Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li
arXiv AI
Jun 10

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.

By Ziqian Zhang, Xingjian Hu, Yue Huang, Kai Zhang, Ruoxi Chen, Yixin Liu, Qingsong Wen, Kaidi Xu, Xiangliang Zhang, Neil Zhenqiang Gong, Lichao Sun