The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.
By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv:2506. 10677v3 Announce Type: replace-cross Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline.
By Otmane Sakhi, Alexandre Gilotte, David Rohde
arXiv:2606. 18750v1 Announce Type: cross Abstract: A/B testing has become the gold standard for data-driven decision-making in large-scale online experimentation, providing critical guidance for feature launch, pricing optimization, and user experience enhancement.
By Yu Zhang, Bokui Wan, Yongli Qin, Jinyong Ma, Yifan Guo
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren
arXiv:2604. 02458v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible.
By Zonghan Li, Feng Ji
The paper presents a layered framework for evaluating conversational AI by aligning offline proxy signals with online A/B experiment outcomes. It introduces a three‑step alignment chain—behavioral label to product outcome, classifier to candidate behavior, and offline signal to experiment effect—alongside an audit protocol that compares confidence intervals and rankings. In a real‑world deployment, the composite proxy achieved 81.1% F1 versus 34.3% for the raw classifier, correctly predicting direction on all 113 contrasts and enabling efficient prioritization of candidate models before costly online testing.
By Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng