The paper proposes a model-based evaluation framework that merges multidimensional item response theory (IRT) with question context embeddings to predict large language model (LLM) performance on unseen questions. By representing LLMs with latent capability profiles and incorporating question content to inform item characteristics, the approach improves prediction accuracy over model-free baselines in within-scenario settings and offers a richer description of capability variation than unidimensional models. However, the study also finds that this generalizability does not reliably extend to cross-scenario shifts, indicating a key limitation for broader application.
By Ergan Shang, Weijing Tang, Yinqiu He
arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.
By Ziqian Zhang, Xingjian Hu, Yue Huang, Kai Zhang, Ruoxi Chen, Yixin Liu, Qingsong Wen, Kaidi Xu, Xiangliang Zhang, Neil Zhenqiang Gong, Lichao Sun
arXiv:2511. 21692v3 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation.
By Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach
arXiv:2610.01627v1 Announce Type: cross
Abstract: Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models o...
By Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan
arXiv:2608. 07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization.
By Karim Radouane, Jose G Moreno, Lynda Tamine
arXiv:2510.01030v2 Announce Type: replace
Abstract: The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust repres...
By Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh