Sebastian Raschka By Sebastian Raschka, PhD

Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)

Read the original on Sebastian Raschka →

Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Sebastian Raschka.

Hugging Face Trending Papers
Sep 24

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study compares Jev, a typed classifier that outputs probabilities over allowed answers, with three flash-tier LLM rubric judges across nine panels from seven benchmarks. Jev’s accuracy differs significantly from an LLM judge in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons are inconclusive. In terms of cost and speed, Jev is 29 to 325 times cheaper and 30 to 220 times faster than the LLM judges, and a cascade that defers uncertain Jev verdicts to an LLM yields only modest gains of up to 2.0 points over the best single judge.

arXiv Computation and Language
Sep 25

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains. whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."

By Delip Rao, Chris Callison-Burch