arXiv Computation and Language
HakemBench is a Turkish benchmark for typed decision-making tasks, comprising 2,346 items and 4,275 choice, yes/no, and score questions across seven tracks (fact‑check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). The benchmark evaluates models on decision quality (macro F1), calibration (normalised Brier score), and selective automation (normalised area under the generalised risk‑coverage curve), combining these metrics by a geometric mean and reporting confidence intervals from 2,000 bootstrap draws. Gold labels are generated from blind passes of a single AI model family compared with votes from other large language model families, and the benchmark’s leader achieved a composite score of 0.888 while the lab’s own model ranked 7th with 0.660.
arXiv:2609.15467v1 Announce Type: cross
Abstract: Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro examines this distinction using 12,000...
By M. Ali Bayram
The paper introduces a paired candidate‑coverage benchmark protocol for typed decision models, evaluating two models—Laya and Jev—on datasets such as AG News, DBpedia, Emotion, and TREC. It reports that Laya detects a high percentage of missing-answer cases but also falsely rejects many valid candidates, whereas Jev shows lower false rejection rates but also lower detection of missing answers. The study highlights the need for separate measurements of classification, score ranking, and rejection policies, noting that the benchmark is descriptive and limited to reference‑label omission.
By Jiawen Lu, Tongtong Wu
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
The article examines how classical test theory statistics—Kuder‑Richardson coefficient, dependability index, and Livingston‑Lewis accuracy—can mislead when applied to large language model (LLM) judges that are evaluated with a single prompt and no gold labels. Using Claude Haiku 4.5 on 210 short‑answer items, the authors show that these metrics fail to isolate the judge’s performance because the judge’s single administration provides no variance component. They argue that reliable statements about an LLM judge require gold labels or varied scorer facets, and that bank design heavily influences reliability estimates.
By Louis Yiven Zhu
arXiv:2606. 25984v2 Announce Type: replace Abstract: Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and apply the specific procedural decision frameworks of expert investors.
By Mingguang Chen, Bo Qu
arXiv:2609.38827v1 Announce Type: new
Abstract: Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet...
By Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang, Yuan Wu
arXiv:2608. 15428v1 Announce Type: cross Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question.
By Volodymyr Ovcharov
arXiv:2608.21601v1 Announce Type: new
Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference soluti...
By Aubrey Brueckner, Darshil Patel, Yuhuan He, Timothy Kassis
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model...
arXiv:2609.37493v1 Announce Type: cross
Abstract: Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error...
By Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim
arXiv:2608.21382v1 Announce Type: new
Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whe...
By V. S. Raghu Parupudi
The paper introduces THPT‑Ladder, a benchmark based on Vietnam’s 2025 National High School Graduation Examination’s convex grading scheme, which rewards partial credit non‑additively. It shows that standard accuracy metrics inflate model scores because they treat partial knowledge proportionally, whereas the official rubric penalizes incomplete correct sets. Using the benchmark, the authors demonstrate that this discrepancy can shift a model’s percentile ranking by up to 13 points among over 480,000 candidates.
By Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh