arXiv Machine Learning

Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation

arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.

arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv Machine Learning
Aug 28

Sequential Additivity in Distributionally Robust Ranking and Selection

The paper studies distributionally robust ranking and selection (DRR&S), where the goal is to identify the best alternative under input uncertainty by considering multiple plausible input distributions. It introduces the concept of sequential additivity, showing that efficient sampling should focus on a small, additive set of critical scenarios rather than a multiplicative number. The authors prove an algorithm‑independent lower bound on sampling, design an additive allocation (AA) procedure that meets this bound and achieves exponentially decreasing error probability, and extend the approach to a general additive allocation (GAA) framework that incorporates traditional R&S sampling rules.

By Zaile Li, Yuchen Wan, L. Jeff Hong
arXiv Machine Learning
Sep 3

On Cost-Aware Designs for Sequential Hypothesis Testing

The paper introduces Cost-Aware Sequential Hypothesis Testing (CASHT), where a decision-maker selects sensing actions with varying random costs to identify the true hypothesis under an average-error constraint while minimizing expected total cost. For fixed costs, the optimal expected total cost scales as Θ(log(1/δ)) and can be achieved by Multihypothesis Sequential Probability Ratio Test-based procedures. The authors extend the framework to random costs under ex-post and ex-ante revelation models, analyze when action cancellation reduces cost, and demonstrate through simulations that CA variants consistently lower total cost compared to classical methods.

By George Vershinin, Asaf Cohen, Omer Gurewitz
arXiv AI
3d ago

Always-On Experimentation

arXiv:2609.38695v1 Announce Type: cross Abstract: Generative AI has dramatically accelerated the rate at which new treatments---from novel pharmaceuticals to online marketing campaigns---can be conce...

By Ricardo J. Sandoval, David Arbour, Avi Feller, Michael I. Jordan
arXiv Computation and Language
Sep 15

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...

By Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung