HakemBench: A Turkish Benchmark of Typed Decisions
Read the original on arXiv Computation and Language →HakemBench is a Turkish benchmark for typed decision-making tasks, comprising 2,346 items and 4,275 choice, yes/no, and score questions across seven tracks (fact‑check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). The benchmark evaluates models on decision quality (macro F1), calibration (normalised Brier score), and selective automation (normalised area under the generalised risk‑coverage curve), combining these metrics by a geometric mean and reporting confidence intervals from 2,000 bootstrap draws. Gold labels are generated from blind passes of a single AI model family compared with votes from other large language model families, and the benchmark’s leader achieved a composite score of 0.888 while the lab’s own model ranked 7th with 0.660.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.