arXiv AI

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

The paper introduces a dual‑dimensional framework called Automated Item Similarity Analysis (AISA) that uses Large Language Models to assess incidental content similarity in large‑scale assessments. It combines Structured Decomposition and Semantic Relatedness to capture both structural and semantic nuances that traditional metrics miss. Psychometric validation shows that LLM‑derived metrics better align with construct‑irrelevant local dependence and produce more coherent item groupings, and simulations in Computerized Adaptive Testing demonstrate improved estimation stability and reduced bias with minimal efficiency loss.

arXiv AI
3d ago

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

The paper proposes a model-based evaluation framework that merges multidimensional item response theory (IRT) with question context embeddings to predict large language model (LLM) performance on unseen questions. By representing LLMs with latent capability profiles and incorporating question content to inform item characteristics, the approach improves prediction accuracy over model-free baselines in within-scenario settings and offers a richer description of capability variation than unidimensional models. However, the study also finds that this generalizability does not reliably extend to cross-scenario shifts, indicating a key limitation for broader application.

By Ergan Shang, Weijing Tang, Yinqiu He
arXiv AI
2d ago

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

The article examines how AI‑assisted item generation is filtered by a computational evaluator before expert review, focusing on representation, structural screening, and candidate‑form dependence. Through two in‑silico studies of 32,000 Big Five items, the authors show that subtle differences in semantic representation and structural evaluation lead to divergent item selections, even when overall content coverage appears stable. The findings reveal that the evaluator, often treated as a neutral technical step, actually shapes the evidence and wording that psychometricians ultimately review, highlighting its role as a revisable component of measurement design.

By Christopher Brooks (School of Information, University of Michigan)
arXiv AI
3d ago

LLM-Specific Utility for Retrieval-Augmented Generation

The paper introduces the concept of LLM‑specific utility, defining it as the performance gain a target large language model (LLM) achieves when provided with a passage compared to answering without evidence. A benchmark of utilitarian passages is built for four LLMs (Qwen3‑8B/14B/32B and Llama 3.1‑8B) across three QA datasets, revealing that each model benefits most from its own tailored evidence and that evidence optimized for other models is consistently suboptimal. The authors also create SpecUBench, a benchmark for LLM‑specific utility judgment, and show that current utility‑aware retrieval methods largely capture model‑agnostic usefulness, struggling to estimate LLM‑specific utility. "whyItMatters":"The study demonstrates that retrieval‑augmented generation must consider model‑specific evidence selection to truly improve LLM performance, highlighting a gap in existing utility‑aware methods."

By Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng
arXiv AI
Aug 6

Item Response Theory for AI Safety

arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.

By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)