Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions.
arXiv:2601. 02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration.
By Christopher Ormerod
arXiv:2512. 07019v3 Announce Type: replace-cross Abstract: The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements.
By Zhiyu Xu, Jia Liu, Yixin Wang, Yuqi Gu
arXiv:2608. 15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities.
By Alona Strugatski, Licol Zeinfeld, Giora Alexandron
arXiv:2312. 07762v3 Announce Type: replace Abstract: Psychiatry research seeks to understand the manifestations of psychopathology in behavior, as measured in questionnaire data, by identifying a small number of latent factors that explain them.
By Ka Chun Lam, Bridget W Mahony, Armin Raznahan, Francisco Pereira
arXiv:2607. 26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data.
By Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson
arXiv:2608. 14606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level.
By Mantas Lukauskas, Viktorija \v{S}arkauskait\.e
arXiv:2607. 25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items.
By Juan Francisco, Mandujano Reyes
Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibration. Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automate the design matrices traditionally specified by hand.
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.
By Yifan Wang
arXiv:2607. 27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive.
By Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet, Virginia Aglietti, Silvia Chiappa
arXiv:2606. 18535v1 Announce Type: cross Abstract: Multi-cause observational studies contain information about unmeasured confounding through the dependence structure among causes.
By Yordan P. Raykov, Hengrui Luo, Justin D. Strait, Wasiur R. KhudaBukhsh