arXiv Computation and Language

StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection

StepCOPS is a new method for selecting a language‑model policy from many checkpoints, prompts, and decoding rules by providing closed‑testing lower‑tail certificates. It uses an independent proposal split to nominate a lower‑tail floor for each candidate, applies exact binomial tests on a fresh certification split, and employs Holm’s step‑down procedure to certify a set of floors. In experiments across 24 configurations and 11 benchmarks, StepCOPS achieves 96.4% selected‑policy coverage, raises the certified floor by 1.5 points over prior methods, stays 0.6 points below a large‑reference jury oracle, and abstains in 2.4% of trials.

arXiv AI
1d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu
arXiv Machine Learning
Aug 11

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb