arXiv AI

Always-On Experimentation

arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv AI
Jun 15

CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation

arXiv:2606. 14581v1 Announce Type: cross Abstract: Granting LLMs direct control over costly, irreversible scientific experiments leads to unsafe exploration and unstable performance, but discarding LLM creativity entirely sacrifices significant optimization potential.

By Guanyu Liu, Weiyi Kong, Zeyu Wang, Boer Zhang, Baiqing Li, Peiyu Zhang, Tianyu Shi
arXiv Machine Learning
Jul 13

Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation

arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.

By Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu
arXiv Machine Learning
Jul 13

Global Sequential Testing for Multi-Stream Auditing

arXiv:2602. 21479v3 Announce Type: replace-cross Abstract: Across many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed.

By Beepul Bharti, Ambar Pal, Jeremias Sulam
arXiv Machine Learning
Sep 3

On Cost-Aware Designs for Sequential Hypothesis Testing

The paper introduces Cost-Aware Sequential Hypothesis Testing (CASHT), where a decision-maker selects sensing actions with varying random costs to identify the true hypothesis under an average-error constraint while minimizing expected total cost. For fixed costs, the optimal expected total cost scales as Θ(log(1/δ)) and can be achieved by Multihypothesis Sequential Probability Ratio Test-based procedures. The authors extend the framework to random costs under ex-post and ex-ante revelation models, analyze when action cancellation reduces cost, and demonstrate through simulations that CA variants consistently lower total cost compared to classical methods.

By George Vershinin, Asaf Cohen, Omer Gurewitz
arXiv AI
Aug 25

The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

The article argues that agentic auto‑research should be guided by dense, intermediate signals of epistemic progress rather than by sparse final benchmarks. It compares this approach to fuzz testing, where coverage provides continuous feedback that directs input mutation. The authors propose controlled experiments to test whether such signals improve discovery efficiency and reduce false positives, and demonstrate in a simulated physics setting that an AI agent using feedback‑driven search uncovers a hidden law while optimization‑driven baselines fail.

By Yifeng He, Jicheng Wang, Yinzhe Zhao, Chengyang Shi, Jiachen Liu, Hao Chen
arXiv Machine Learning
Jun 2

Bandit Simulation for Average Reward Inference

arXiv:2606. 00913v1 Announce Type: cross Abstract: Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge.

By Samya Praharaj, Chih-Yu Chang, Koulik Khamaru, Kelly W. Zhang
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv Statistics ML
Sep 25

Confidence Horizons

The paper introduces "confidence horizons", a new class of statistical tools that provide sharper large‑sample anytime‑valid inference when a finite time horizon is imposed. These objects function as large‑sample confidence sequences limited to a bounded number of interim looks, analogous to group sequential repeated confidence intervals. The authors connect confidence horizons to classic group sequential boundaries (Pocock, O’Brien–Fleming, Wang–Tsiatis), derive closed‑form distribution functions for certain statistics, and demonstrate their application to treatment effect estimation in sequentially randomized experiments with adaptive Neyman allocation.

By Chase Mathis, Ian Waudby-Smith