arXiv AI By Jiawen Lu, Tongtong Wu

Benchmarking Candidate Coverage in Typed Decision Models

Read the original on arXiv AI →

The paper introduces a paired candidate‑coverage benchmark protocol for typed decision models, evaluating two models—Laya and Jev—on datasets such as AG News, DBpedia, Emotion, and TREC. It reports that Laya detects a high percentage of missing-answer cases but also falsely rejects many valid candidates, whereas Jev shows lower false rejection rates but also lower detection of missing answers. The study highlights the need for separate measurements of classification, score ranking, and rejection policies, noting that the benchmark is descriptive and limited to reference‑label omission.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
6d ago

Benchmarking System One decision models against trained classifiers and language models for automated decision gates

The paper evaluates System One decision models—typed models that output probabilities for branching decisions—against supervised classifiers and generative language models on automated decision gate tasks. Eight checkpoints from six families, including the hosted model Jev, were benchmarked on workflow, intent, and social‑science items, showing that small trained classifiers match or slightly outperform decision models on intent and workflow when labels are available, while decision models outperform zero‑shot classifiers when labels are absent. The study also explores calibration, risk thresholds, and cost‑efficiency trade‑offs, providing condition‑dependent design guidelines for automated decision gates.

By Amir Rafe, Subasish Das