arXiv Computation and Language By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection

Read the original on arXiv Computation and Language →

StepCOPS is a new method for selecting a language‑model policy from many checkpoints, prompts, and decoding rules by providing closed‑testing lower‑tail certificates. It uses an independent proposal split to nominate a lower‑tail floor for each candidate, applies exact binomial tests on a fresh certification split, and employs Holm’s step‑down procedure to certify a set of floors. In experiments across 24 configurations and 11 benchmarks, StepCOPS achieves 96.4% selected‑policy coverage, raises the certified floor by 1.5 points over prior methods, stays 0.6 points below a large‑reference jury oracle, and abstains in 2.4% of trials.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
1d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu