Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
arXiv:2606. 18750v1 Announce Type: cross Abstract: A/B testing has become the gold standard for data-driven decision-making in large-scale online experimentation, providing critical guidance for feature launch, pricing optimization, and user experience enhancement.
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
arXiv:2506. 10677v3 Announce Type: replace-cross Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline.
The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.
arXiv:2607. 01958v1 Announce Type: new Abstract: A/B testing is the gold standard for selecting the better algorithm in online services.
arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.
arXiv:2608. 02345v2 Announce Type: replace-cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry.
arXiv:2602. 16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments.
The paper introduces BAFF, a Bid‑Aware Filter Family that mitigates training data interference in real‑time bidding (RTB) A/B tests by applying (k,l)-parameterized hard filters to control bias from ad‑ranking and bid‑pricing disagreements. It proposes a three‑stage online measurement protocol to evaluate data‑sharing strategies against an interference‑free reference model. Experiments show that BAFF variants outperform both log‑sharing and log‑splitting in offline simulations and live DSP deployments, preserving key business metrics more closely.
Information systems researchers increasingly rely on quasi‑experimental methods such as difference‑in‑differences and instrumental variables to infer causal effects from observational panel data. A large Monte Carlo study of 9,837 parameter settings (≈9.8 million simulated datasets) shows that the gap between planned and achieved power is largely driven by serial correlation, panel attrition, staggered adoption bias, and parallel‑trend pre‑testing—factors that no closed‑form power calculator can fully capture. For IV designs, increasing sample size does not improve power or reduce exclusion bias unless instrument strength is enhanced, underscoring that identification hinges on the instrument rather than on larger N.
arXiv:2603. 20775v2 Announce Type: replace Abstract: In personalized marketing, uplift models estimate the incremental effect of an intervention by modeling how customer behavior would change under alternative treatments using counterfactual analysis.