The paper introduces a new approach to safety in contextual bandits with continuous actions, focusing on high‑probability constraints on the realized cost rather than expected cost. It presents the High‑Probability Constrained UCB algorithm, which balances reward exploration with conservative safety estimation, and provides theoretical regret guarantees for linear models and extensions to general function classes. Experiments demonstrate that this realized‑cost safety framework significantly reduces safety violations compared to expected‑cost constrained methods.
By Spyros Dragazis, Aldo Pacchiano
arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.
By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
The paper studies distributionally robust ranking and selection (DRR&S), where the goal is to identify the best alternative under input uncertainty by considering multiple plausible input distributions. It introduces the concept of sequential additivity, showing that efficient sampling should focus on a small, additive set of critical scenarios rather than a multiplicative number. The authors prove an algorithm‑independent lower bound on sampling, design an additive allocation (AA) procedure that meets this bound and achieves exponentially decreasing error probability, and extend the approach to a general additive allocation (GAA) framework that incorporates traditional R&S sampling rules.
By Zaile Li, Yuchen Wan, L. Jeff Hong
arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.
By Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu
arXiv:2608.28599v1 Announce Type: new
Abstract: Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosi...
By Qi Peng, Yi Cai, Changmeng Zheng, Xin Wu, Jiayuan Xie, Qing Li
The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.
By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha