JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2609.24574v1 Announce Type: new Abstract: Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the...
arXiv:2609.22133v1 Announce Type: new Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at...
arXiv:2609.37647v1 Announce Type: cross Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
arXiv:2605. 11954v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables that can enter standard empirical designs.
The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains. whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."
arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.