arXiv Machine Learning

Sequential Additivity in Distributionally Robust Ranking and Selection

The paper studies distributionally robust ranking and selection (DRR&S), where the goal is to identify the best alternative under input uncertainty by considering multiple plausible input distributions. It introduces the concept of sequential additivity, showing that efficient sampling should focus on a small, additive set of critical scenarios rather than a multiplicative number. The authors prove an algorithm‑independent lower bound on sampling, design an additive allocation (AA) procedure that meets this bound and achieves exponentially decreasing error probability, and extend the approach to a general additive allocation (GAA) framework that incorporates traditional R&S sampling rules.

arXiv Machine Learning
Jul 13

Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation

arXiv:2506. 03062v2 Announce Type: replace Abstract: A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average treatment effect, especially with many metrics (e.

By Qining Zhang, Tanner Fiez, Yi Liu, Wenyang Liu
arXiv Machine Learning
Jul 30

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.

By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv AI
Jul 23

Long-Term Sequential Decision Making under Risk

arXiv:2607. 19914v1 Announce Type: new Abstract: We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns.

By Irmaan (Mohammad), Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib
arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv Machine Learning
Sep 22

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng
arXiv Machine Learning
Sep 3

On Cost-Aware Designs for Sequential Hypothesis Testing

The paper introduces Cost-Aware Sequential Hypothesis Testing (CASHT), where a decision-maker selects sensing actions with varying random costs to identify the true hypothesis under an average-error constraint while minimizing expected total cost. For fixed costs, the optimal expected total cost scales as Θ(log(1/δ)) and can be achieved by Multihypothesis Sequential Probability Ratio Test-based procedures. The authors extend the framework to random costs under ex-post and ex-ante revelation models, analyze when action cancellation reduces cost, and demonstrate through simulations that CA variants consistently lower total cost compared to classical methods.

By George Vershinin, Asaf Cohen, Omer Gurewitz
arXiv Machine Learning
Aug 28

UCB for Large-Scale Pure Exploration: Beyond Sub-Gaussianity

The paper studies upper confidence bound (UCB) algorithms for large‑scale pure exploration problems where the performance distributions may be heavy‑tailed and not sub‑Gaussian. It introduces a meta‑UCB framework that selects the alternative with the largest sample size as the best upon stopping, and derives a distribution‑free lower bound on the probability of correct selection. Using this bound, the authors show that the meta‑UCB algorithm achieves sample optimality in both indifference‑zone and non‑indifference‑zone settings under a uniform variance bound, and provide numerical experiments that illustrate the behavior of UCB algorithms beyond the meta‑UCB framework.

By Zaile Li, Weiwei Fan, L. Jeff Hong