arXiv:2610.03387v2 Announce Type: replace
Abstract: Rejection policies must remain useful as candidate sets and tasks change. We compare Laya, Jev and Qwen2.5-7B-Instruct using public reference label...
By Jiawen Lu, Tongtong Wu
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
The paper evaluates System One decision models—typed models that output probabilities for branching decisions—against supervised classifiers and generative language models on automated decision gate tasks. Eight checkpoints from six families, including the hosted model Jev, were benchmarked on workflow, intent, and social‑science items, showing that small trained classifiers match or slightly outperform decision models on intent and workflow when labels are available, while decision models outperform zero‑shot classifiers when labels are absent. The study also explores calibration, risk thresholds, and cost‑efficiency trade‑offs, providing condition‑dependent design guidelines for automated decision gates.
By Amir Rafe, Subasish Das
arXiv:2610.02267v1 Announce Type: new
Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input...
By Jiawei Li
arXiv:2609.39496v1 Announce Type: cross
Abstract: Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefine...
By Jike Zhong, Ming Li, Yuxiang Lai
arXiv:2609.38827v1 Announce Type: new
Abstract: Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet...
By Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang, Yuan Wu
arXiv:2606. 12702v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems.
By Alyssa Unell, Miguel Fuentes, Brenna Li, Bridget Lin, Meena Jagadeesan, Sanmi Koyejo, Nigam Shah
arXiv:2609.24574v1 Announce Type: new
Abstract: Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the...
By Hazem Ibrahim, Yasir Zaki
The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains.
whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."
By Delip Rao, Chris Callison-Burch
PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.
By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model...
arXiv:2607. 20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly.
By Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang, Sanyam Kapoor