arXiv AI By Jing Huang, Jihong Zhang, Hua-Hua Chang

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

Read the original on arXiv AI →

The paper introduces a dual‑dimensional framework called Automated Item Similarity Analysis (AISA) that uses Large Language Models to assess incidental content similarity in large‑scale assessments. It combines Structured Decomposition and Semantic Relatedness to capture both structural and semantic nuances that traditional metrics miss. Psychometric validation shows that LLM‑derived metrics better align with construct‑irrelevant local dependence and produce more coherent item groupings, and simulations in Computerized Adaptive Testing demonstrate improved estimation stability and reduced bias with minimal efficiency loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

The paper proposes a model-based evaluation framework that merges multidimensional item response theory (IRT) with question context embeddings to predict large language model (LLM) performance on unseen questions. By representing LLMs with latent capability profiles and incorporating question content to inform item characteristics, the approach improves prediction accuracy over model-free baselines in within-scenario settings and offers a richer description of capability variation than unidimensional models. However, the study also finds that this generalizability does not reliably extend to cross-scenario shifts, indicating a key limitation for broader application.

By Ergan Shang, Weijing Tang, Yinqiu He
arXiv AI
2d ago

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

The article examines how AI‑assisted item generation is filtered by a computational evaluator before expert review, focusing on representation, structural screening, and candidate‑form dependence. Through two in‑silico studies of 32,000 Big Five items, the authors show that subtle differences in semantic representation and structural evaluation lead to divergent item selections, even when overall content coverage appears stable. The findings reveal that the evaluator, often treated as a neutral technical step, actually shapes the evidence and wording that psychometricians ultimately review, highlighting its role as a revisable component of measurement design.

By Christopher Brooks (School of Information, University of Michigan)