Prediction-Powered Active Testing
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled. However, existing estimators fail to exploit the informative predictions of powerful black--box models, even though such predictions are increasingly available in settings where labels remain expensive.
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
arXiv:2602. 02229v2 Announce Type: replace Abstract: We study the problem of monitoring model performance in dynamic environments where labeled data are limited.
arXiv:2505. 20178v2 Announce Type: replace-cross Abstract: Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation.
arXiv:2606. 14506v1 Announce Type: cross Abstract: Understanding how a prediction model will perform in a new environment before deployment is essential to preventing harm when algorithms inform decision-making.
arXiv:2609.37687v1 Announce Type: new Abstract: Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively queryin...
arXiv:2608.27704v1 Announce Type: new Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version,...
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
The paper introduces prediction‑powered smoothing (PP‑S) and its taxonomy‑aware extension (PP‑TS) to improve point and interval estimates of domain‑specific AI performance when only a limited sample of labeled units is available. It also proposes a new design‑based cross‑validation score that is approximately unbiased for selecting between direct and smoothed estimators. Experiments on a curated benchmark and real‑world agent traffic show that the proposed methods outperform direct estimators in both accuracy and coverage, and that the new score matches the performance of an independent validation sample while providing more precise error estimates.
arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
arXiv:2606. 14909v1 Announce Type: cross Abstract: We consider the problem of uncertainty quantification for a pretrained classification model deployed under unknown distribution shift.