arXiv AI By Ergan Shang, Weijing Tang, Yinqiu He

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

Read the original on arXiv AI →

The paper proposes a model-based evaluation framework that merges multidimensional item response theory (IRT) with question context embeddings to predict large language model (LLM) performance on unseen questions. By representing LLMs with latent capability profiles and incorporating question content to inform item characteristics, the approach improves prediction accuracy over model-free baselines in within-scenario settings and offers a richer description of capability variation than unidimensional models. However, the study also finds that this generalizability does not reliably extend to cross-scenario shifts, indicating a key limitation for broader application.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

LLM-Specific Utility for Retrieval-Augmented Generation

The paper introduces the concept of LLM‑specific utility, defining it as the performance gain a target large language model (LLM) achieves when provided with a passage compared to answering without evidence. A benchmark of utilitarian passages is built for four LLMs (Qwen3‑8B/14B/32B and Llama 3.1‑8B) across three QA datasets, revealing that each model benefits most from its own tailored evidence and that evidence optimized for other models is consistently suboptimal. The authors also create SpecUBench, a benchmark for LLM‑specific utility judgment, and show that current utility‑aware retrieval methods largely capture model‑agnostic usefulness, struggling to estimate LLM‑specific utility. "whyItMatters":"The study demonstrates that retrieval‑augmented generation must consider model‑specific evidence selection to truly improve LLM performance, highlighting a gap in existing utility‑aware methods."

By Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng
arXiv AI
2d ago

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

The paper introduces a dual‑dimensional framework called Automated Item Similarity Analysis (AISA) that uses Large Language Models to assess incidental content similarity in large‑scale assessments. It combines Structured Decomposition and Semantic Relatedness to capture both structural and semantic nuances that traditional metrics miss. Psychometric validation shows that LLM‑derived metrics better align with construct‑irrelevant local dependence and produce more coherent item groupings, and simulations in Computerized Adaptive Testing demonstrate improved estimation stability and reduced bias with minimal efficiency loss.

By Jing Huang, Jihong Zhang, Hua-Hua Chang
arXiv AI
Jun 10

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.

By Ziqian Zhang, Xingjian Hu, Yue Huang, Kai Zhang, Ruoxi Chen, Yixin Liu, Qingsong Wen, Kaidi Xu, Xiangliang Zhang, Neil Zhenqiang Gong, Lichao Sun