arXiv Machine Learning

Online Conformal Prediction Beyond Feedback

arXiv:2608. 07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications.

arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv Machine Learning
5d ago

Counterfactual Online Conformal Prediction Under Adaptive Logging

The paper addresses the failure of online conformal prediction when predictions influence actions that determine which outcomes are used for calibration. It introduces Propensity-Weighted Online Conformal Prediction (PW‑OCP), an inverse‑propensity‑weighted recursion that debiases calibration, and a doubly robust variant (DR‑OCP) that further reduces bias. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing demonstrate that PW‑OCP and DR‑OCP improve counterfactual coverage and downstream regret while preserving prediction‑set sharpness.

By Xinyu Qiao, Yichen Lin, Kaihong Ji, Xue Wang, Tao Yao
arXiv Machine Learning
Jun 26

Blackwell Approachability and Gradient Equilibrium are Equivalent

arXiv:2606. 27315v1 Announce Type: new Abstract: Gradient equilibrium (GEQ) is a recently introduced online optimization framework that generalizes first-order stationarity from offline optimization and abstracts problems like online conformal prediction.

By Brian W. Lee, Nika Haghtalab, Michael I. Jordan, Ryan J. Tibshirani
arXiv AI
Sep 3

Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor

The paper presents an online algorithm that achieves the same $0.401$ approximation factor for maximizing nonnegative, non-monotone DR-submodular functions over compact convex down-closed subsets of the $d$-dimensional unit cube as the best known offline construction. In the full-information value-oracle model, the algorithm attains this factor with sublinear regret, using $O(dT^{1/4})$ oracle calls per round and $O(T^{3/4})$ regret, and offers flexible batching trade-offs. Under a positive-anchor condition, a randomized blocking strategy preserves the $0.401$ factor while achieving $O(T^{5/6})$ one-point bandit regret.

By Vaneet Aggarwal, Yiyang Lu