arXiv Computation and Language By Sait Furkan Teke (ufak AI)

HakemBench: A Turkish Benchmark of Typed Decisions

Read the original on arXiv Computation and Language →

HakemBench is a Turkish benchmark for typed decision-making tasks, comprising 2,346 items and 4,275 choice, yes/no, and score questions across seven tracks (fact‑check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). The benchmark evaluates models on decision quality (macro F1), calibration (normalised Brier score), and selective automation (normalised area under the generalised risk‑coverage curve), combining these metrics by a geometric mean and reporting confidence intervals from 2,000 bootstrap draws. Gold labels are generated from blind passes of a single AI model family compared with votes from other large language model families, and the benchmark’s leader achieved a composite score of 0.888 while the lab’s own model ranked 7th with 0.660.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
4d ago

Benchmarking Candidate Coverage in Typed Decision Models

The paper introduces a paired candidate‑coverage benchmark protocol for typed decision models, evaluating two models—Laya and Jev—on datasets such as AG News, DBpedia, Emotion, and TREC. It reports that Laya detects a high percentage of missing-answer cases but also falsely rejects many valid candidates, whereas Jev shows lower false rejection rates but also lower detection of missing answers. The study highlights the need for separate measurements of classification, score ranking, and rejection policies, noting that the benchmark is descriptive and limited to reference‑label omission.

By Jiawen Lu, Tongtong Wu
arXiv Computation and Language
Oct 1

Three Ways Classical Test Theory Can Mislead About LLM Judges

The article examines how classical test theory statistics—Kuder‑Richardson coefficient, dependability index, and Livingston‑Lewis accuracy—can mislead when applied to large language model (LLM) judges that are evaluated with a single prompt and no gold labels. Using Claude Haiku 4.5 on 210 short‑answer items, the authors show that these metrics fail to isolate the judge’s performance because the judge’s single administration provides no variance component. They argue that reliable statements about an LLM judge require gold labels or varied scorer facets, and that bank design heavily influences reliability estimates.

By Louis Yiven Zhu