Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area und...
arXiv:2609.37493v1 Announce Type: cross
Abstract: Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error...
By Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2606. 15153v1 Announce Type: new Abstract: Selective prediction with distribution-free risk control promises that, with confidence 1-delta over the calibration draw, the error rate of accepted inputs stays below a user budget alpha.
By Jingwen Zhou, Mingzhe Wang
StepCOPS is a new method for selecting a language‑model policy from many checkpoints, prompts, and decoding rules by providing closed‑testing lower‑tail certificates. It uses an independent proposal split to nominate a lower‑tail floor for each candidate, applies exact binomial tests on a fresh certification split, and employs Holm’s step‑down procedure to certify a set of floors. In experiments across 24 configurations and 11 benchmarks, StepCOPS achieves 96.4% selected‑policy coverage, raises the certified floor by 1.5 points over prior methods, stays 0.6 points below a large‑reference jury oracle, and abstains in 2.4% of trials.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv:2609.06873v1 Announce Type: cross
Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...
By F. Setoudehtanzangi, Geoffrey J. McLachlan