arXiv Machine Learning

Ensuring Trustworthy Online A/B Testing: Addressing Five Key Questions on CUPED

arXiv:2606. 18750v1 Announce Type: cross Abstract: A/B testing has become the gold standard for data-driven decision-making in large-scale online experimentation, providing critical guidance for feature launch, pricing optimization, and user experience enhancement.

arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv Machine Learning
Jul 13

Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation

arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.

By Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu
arXiv Machine Learning
Jun 4

Validity Threats for Foundation Model Research

arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.

By Gunnar K\"onig, Martin Pawelczyk, Ulrike von Luxburg, Sebastian Bordt
arXiv Machine Learning
Sep 10

BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests

The paper introduces BAFF, a Bid‑Aware Filter Family that mitigates training data interference in real‑time bidding (RTB) A/B tests by applying (k,l)-parameterized hard filters to control bias from ad‑ranking and bid‑pricing disagreements. It proposes a three‑stage online measurement protocol to evaluate data‑sharing strategies against an interference‑free reference model. Experiments show that BAFF variants outperform both log‑sharing and log‑splitting in offline simulations and live DSP deployments, preserving key business metrics more closely.

By Jeonglyul Oh, Ikkyu Choi, Inseop Youn, Youngjae Kim
arXiv AI
Sep 24

Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS

Information systems researchers increasingly rely on quasi‑experimental methods such as difference‑in‑differences and instrumental variables to infer causal effects from observational panel data. A large Monte Carlo study of 9,837 parameter settings (≈9.8 million simulated datasets) shows that the gap between planned and achieved power is largely driven by serial correlation, panel attrition, staggered adoption bias, and parallel‑trend pre‑testing—factors that no closed‑form power calculator can fully capture. For IV designs, increasing sample size does not improve power or reduce exclusion bias unless instrument strength is enhanced, underscoring that identification hinges on the instrument rather than on larger N.

By Spandan Ghose Chowdhury