arXiv Machine Learning

Ensuring Trustworthy Online A/B Testing: Addressing Five Key Questions on CUPED

arXiv:2606. 18750v1 Announce Type: cross Abstract: A/B testing has become the gold standard for data-driven decision-making in large-scale online experimentation, providing critical guidance for feature launch, pricing optimization, and user experience enhancement.

arXiv Machine Learning
Jul 13

Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation

arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.

By Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu
arXiv Machine Learning
Jun 4

Validity Threats for Foundation Model Research

arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.

By Gunnar K\"onig, Martin Pawelczyk, Ulrike von Luxburg, Sebastian Bordt
arXiv Machine Learning
Jul 10

Prediction-Powered Active Testing

arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.

By Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, Fran\c{c}ois Caron