arXiv Computation and Language

HakemBench: A Turkish Benchmark of Typed Decisions

HakemBench is a Turkish benchmark for typed decision-making tasks, comprising 2,346 items and 4,275 choice, yes/no, and score questions across seven tracks (fact‑check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). The benchmark evaluates models on decision quality (macro F1), calibration (normalised Brier score), and selective automation (normalised area under the generalised risk‑coverage curve), combining these metrics by a geometric mean and reporting confidence intervals from 2,000 bootstrap draws. Gold labels are generated from blind passes of a single AI model family compared with votes from other large language model families, and the benchmark’s leader achieved a composite score of 0.888 while the lab’s own model ranked 7th with 0.660.

arXiv AI
4d ago

Benchmarking Candidate Coverage in Typed Decision Models

The paper introduces a paired candidate‑coverage benchmark protocol for typed decision models, evaluating two models—Laya and Jev—on datasets such as AG News, DBpedia, Emotion, and TREC. It reports that Laya detects a high percentage of missing-answer cases but also falsely rejects many valid candidates, whereas Jev shows lower false rejection rates but also lower detection of missing answers. The study highlights the need for separate measurements of classification, score ranking, and rejection policies, noting that the benchmark is descriptive and limited to reference‑label omission.

By Jiawen Lu, Tongtong Wu
arXiv Computation and Language
Oct 1

Three Ways Classical Test Theory Can Mislead About LLM Judges

The article examines how classical test theory statistics—Kuder‑Richardson coefficient, dependability index, and Livingston‑Lewis accuracy—can mislead when applied to large language model (LLM) judges that are evaluated with a single prompt and no gold labels. Using Claude Haiku 4.5 on 210 short‑answer items, the authors show that these metrics fail to isolate the judge’s performance because the judge’s single administration provides no variance component. They argue that reliable statements about an LLM judge require gold labels or varied scorer facets, and that bank design heavily influences reliability estimates.

By Louis Yiven Zhu
arXiv AI
Aug 20

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

The paper introduces THPT‑Ladder, a benchmark based on Vietnam’s 2025 National High School Graduation Examination’s convex grading scheme, which rewards partial credit non‑additively. It shows that standard accuracy metrics inflate model scores because they treat partial knowledge proportionally, whereas the official rubric penalizes incomplete correct sets. Using the benchmark, the authors demonstrate that this discrepancy can shift a model’s percentile ranking by up to 13 points among over 480,000 candidates.

By Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh