Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
arXiv:2606. 18750v1 Announce Type: cross Abstract: A/B testing has become the gold standard for data-driven decision-making in large-scale online experimentation, providing critical guidance for feature launch, pricing optimization, and user experience enhancement.
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
arXiv:2506. 10677v3 Announce Type: replace-cross Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline.
arXiv:2607. 01958v1 Announce Type: new Abstract: A/B testing is the gold standard for selecting the better algorithm in online services.
arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.
arXiv:2608. 02345v2 Announce Type: replace-cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry.
arXiv:2602. 16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments.
arXiv:2603. 20775v2 Announce Type: replace Abstract: In personalized marketing, uplift models estimate the incremental effect of an intervention by modeling how customer behavior would change under alternative treatments using counterfactual analysis.
arXiv:2606. 04110v1 Announce Type: new Abstract: Online evaluation of ranking and retrieval systems often relies on downstream monetization metrics such as app revenue or creator earnings.
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
arXiv:2606. 30932v1 Announce Type: new Abstract: Two-sided marketplaces connect distinct user groups whose interests often conflict -- improving outcomes on one side could degrade the other side's experience.