arXiv Machine Learning By Kimon Antonios Provatas, Ilias Georgakopoulos-Soares

Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model

Read the original on arXiv Machine Learning →

The paper audits a commercial non‑generative System‑1 model on biosecurity‑relevant benchmarks, evaluating accuracy, calibration, error detection, selective prediction, and sensitivity to answer‑option order. It finds that the model’s accuracy varies strongly by task, is reasonably well calibrated when the vendor’s uncertainty field is interpreted correctly, and that option order can cause significant prediction changes—averaging across rotations improves accuracy. The study also shows that applying averaging only to low‑confidence items recovers most of the gain at a lower cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Benchmarking System One decision models against trained classifiers and language models for automated decision gates

The paper evaluates System One decision models—typed models that output probabilities for branching decisions—against supervised classifiers and generative language models on automated decision gate tasks. Eight checkpoints from six families, including the hosted model Jev, were benchmarked on workflow, intent, and social‑science items, showing that small trained classifiers match or slightly outperform decision models on intent and workflow when labels are available, while decision models outperform zero‑shot classifiers when labels are absent. The study also explores calibration, risk thresholds, and cost‑efficiency trade‑offs, providing condition‑dependent design guidelines for automated decision gates.

By Amir Rafe, Subasish Das
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
Aug 28

Invocation-Level Reliability of Tool-Using Agents

The paper investigates the reliability of tool‑using agents, focusing on two failure modes: selecting the wrong tool and constructing incorrect arguments. It introduces a correct‑invocation rate metric to distinguish these errors and evaluates five open‑weight models on multi‑step tasks up to depth 8, finding that by depth 6 about 70% of a model’s clean‑context capability is lost due to earlier mistakes. The study reveals that exact‑match scoring against a fixed gold trajectory forces severity and recovery parameters to extreme values, and proposes a conditional‑on‑state scoring remedy that yields more realistic severity estimates.

By Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta