arXiv AI By Juan Francisco, Mandujano Reyes

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

Read the original on arXiv AI →

arXiv:2607. 25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.