arXiv AI By Louie Hong Yao, Nicholas Jarvis, Tiffany Zhan, Saptarshi Ghosh, Linfeng Liu, Tianyu Jiang

JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory

Read the original on arXiv AI →

arXiv:2509. 22888v2 Announce Type: replace Abstract: Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

The paper proposes a model-based evaluation framework that merges multidimensional item response theory (IRT) with question context embeddings to predict large language model (LLM) performance on unseen questions. By representing LLMs with latent capability profiles and incorporating question content to inform item characteristics, the approach improves prediction accuracy over model-free baselines in within-scenario settings and offers a richer description of capability variation than unidimensional models. However, the study also finds that this generalizability does not reliably extend to cross-scenario shifts, indicating a key limitation for broader application.

By Ergan Shang, Weijing Tang, Yinqiu He
arXiv AI
Jun 10

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.

By Ziqian Zhang, Xingjian Hu, Yue Huang, Kai Zhang, Ruoxi Chen, Yixin Liu, Qingsong Wen, Kaidi Xu, Xiangliang Zhang, Neil Zhenqiang Gong, Lichao Sun