arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
By Olivier Jeunen
arXiv:2506. 10677v3 Announce Type: replace-cross Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline.
By Otmane Sakhi, Alexandre Gilotte, David Rohde
arXiv:2607. 01958v1 Announce Type: new Abstract: A/B testing is the gold standard for selecting the better algorithm in online services.
By Koki Konishi, Masataka Ushiku, Yuta Saito
arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.
By Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren
arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.
By Gunnar K\"onig, Martin Pawelczyk, Ulrike von Luxburg, Sebastian Bordt