arXiv Statistics ML

Confidence Horizons

The paper introduces "confidence horizons", a new class of statistical tools that provide sharper large‑sample anytime‑valid inference when a finite time horizon is imposed. These objects function as large‑sample confidence sequences limited to a bounded number of interim looks, analogous to group sequential repeated confidence intervals. The authors connect confidence horizons to classic group sequential boundaries (Pocock, O’Brien–Fleming, Wang–Tsiatis), derive closed‑form distribution functions for certain statistics, and demonstrate their application to treatment effect estimation in sequentially randomized experiments with adaptive Neyman allocation.

arXiv Machine Learning
Jun 2

Bandit Simulation for Average Reward Inference

arXiv:2606. 00913v1 Announce Type: cross Abstract: Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge.

By Samya Praharaj, Chih-Yu Chang, Koulik Khamaru, Kelly W. Zhang
arXiv Machine Learning
Jun 5

Multi-Armed Sequential Hypothesis Testing by Betting

arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.

By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
Hugging Face Trending Papers
Jun 8

Asymptotic Optimality of Thompson Sampling for Risk-Averse Bandits with Sub-Gaussian Rewards

We prove that $ρ\text{-}\mathrm{NPTS}_{\mathrm{SG}}$, an anchor-free nonparametric Thompson Sampling algorithm for risk-averse bandits, achieves regret matching the instance-dependent lower bound to leading order in $\log n$, establishing it as asymptotically optimal for any continuous risk functional $ρ$ (CVaR, mean-variance, Sharpe ratio, distortion risk measures, and more) on the class of distributions with bounded density and sub-Gaussian tails, including Gaussian arms. Both this result and its bounded-support counterpart require only continuity of $ρ$: strictly weaker than the dominance condition of prior parametric Thompson Sampling results, and strictly weaker than the Lipschitz condition of UCB-type algorithms, yielding the first instance-optimal guarantees for non-Lipschitz functionals such as the Sharpe ratio without parametric reward assumptions.

arXiv Machine Learning
Aug 28

Sequential Additivity in Distributionally Robust Ranking and Selection

The paper studies distributionally robust ranking and selection (DRR&S), where the goal is to identify the best alternative under input uncertainty by considering multiple plausible input distributions. It introduces the concept of sequential additivity, showing that efficient sampling should focus on a small, additive set of critical scenarios rather than a multiplicative number. The authors prove an algorithm‑independent lower bound on sampling, design an additive allocation (AA) procedure that meets this bound and achieves exponentially decreasing error probability, and extend the approach to a general additive allocation (GAA) framework that incorporates traditional R&S sampling rules.

By Zaile Li, Yuchen Wan, L. Jeff Hong
arXiv Machine Learning
Jun 9

Asymptotic Optimality of Thompson Sampling for Risk-Averse Bandits with Sub-Gaussian Rewards

arXiv:2606. 09191v1 Announce Type: new Abstract: We prove that $\rho\text{-}\mathrm{NPTS}_{\mathrm{SG}}$, an anchor-free nonparametric Thompson Sampling algorithm for risk-averse bandits, achieves regret matching the instance-dependent lower bound to leading order in $\log n$, establishing it as asymptotically optimal for any continuous risk functional $\rho$ (CVaR, mean-variance, Sharpe ratio, distortion risk measures, and more) on the class of distributions with bounded density and sub-Gaussian tails, including Gaussian arms.

By Joel Q. L. Chang