Hugging Face Trending Papers

How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction

arXiv Machine Learning
Sep 10

Large Classification-Risk-Optional Label Acquisition

arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...

By F. Setoudehtanzangi, Geoffrey J. McLachlan
arXiv Computation and Language
Sep 25

StepCOPS: Closed-Testing Lower-Tail Certificates for Language-Model Policy Selection

StepCOPS is a new method for selecting a language‑model policy from many checkpoints, prompts, and decoding rules by providing closed‑testing lower‑tail certificates. It uses an independent proposal split to nominate a lower‑tail floor for each candidate, applies exact binomial tests on a fresh certification split, and employs Holm’s step‑down procedure to certify a set of floors. In experiments across 24 configurations and 11 benchmarks, StepCOPS achieves 96.4% selected‑policy coverage, raises the certified floor by 1.5 points over prior methods, stays 0.6 points below a large‑reference jury oracle, and abstains in 2.4% of trials.

By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv Computer Vision
Aug 27

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.

By Juan I\~naki Larrea, Lucas Mansilla, Enzo Ferrante