arXiv Machine Learning By Tetsuji Kuboyama

How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction

Read the original on arXiv Machine Learning →

arXiv:2609. 18622v1 Announce Type: new Abstract: Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 25

StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection

StepCOPS is a new method for selecting a language‑model policy from many checkpoints, prompts, and decoding rules by providing closed‑testing lower‑tail certificates. It uses an independent proposal split to nominate a lower‑tail floor for each candidate, applies exact binomial tests on a fresh certification split, and employs Holm’s step‑down procedure to certify a set of floors. In experiments across 24 configurations and 11 benchmarks, StepCOPS achieves 96.4% selected‑policy coverage, raises the certified floor by 1.5 points over prior methods, stays 0.6 points below a large‑reference jury oracle, and abstains in 2.4% of trials.

By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv Machine Learning
Sep 10

Large Classification-Risk-Optional Label Acquisition

arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...

By F. Setoudehtanzangi, Geoffrey J. McLachlan