arXiv:2607. 26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback.
By Nicolas Gutowski, Fabien Chhel, Alexandre Letard, Sylvain Lamprier
arXiv:2607. 08979v1 Announce Type: new Abstract: We study the active learning problem of fixed-confidence top-$k$ identification from noisy pairwise comparisons.
By Motti Goldberger, Nils Rudi
The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.
By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv:2608. 04324v1 Announce Type: cross Abstract: This paper studies generalized low-rank matrix bandits with multiple prioritized objectives.
By Bo Xue, Ji Cheng, Haodong Jing, Hongzong Li, Shuang Qiu
The paper presents a computationally efficient optimal design framework for multinomial logit (MNL) bandits, addressing the combinatorial action space that makes traditional methods infeasible. It introduces two approaches: an exact or certified-approximate mixed-integer linear program with solver‑certified early stopping, and a fully polynomial‑time lifted design using a tractable surrogate objective. Leveraging the Kiefer‑Wolfowitz equivalence theorem, the authors provide near G‑optimality guarantees and apply the framework to develop a best assortment identification algorithm with an instance‑dependent sample complexity of τO((d log N)/Δ²).
By Joongkyu Lee, Min-hwan Oh
arXiv:2605.12340v5 Announce Type: replace-cross
Abstract: Learning-to-Defer (L2D) methods route each query either to a predictive model or to external experts. Real-world deployments require handling...
By Dang Hoang Duy, Yannis Montreuil, Maxime Meyer, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi
arXiv:2607. 11684v1 Announce Type: cross Abstract: Existing contextual multinomial logit (MNL) bandits model relevance-driven choice but ignore the potential benefits of within-assortment diversity, while submodular/combinatorial bandits encode diversity in rewards but lack structured choice probabilities.
By Heesang Ann, Taehyun Hwang, Min-hwan Oh
arXiv:2606. 06043v1 Announce Type: cross Abstract: Follow-the-regularized-leader framework has shown effectiveness and flexibility in online learning problems, where the choice of learning rates are known to be crucial.
By Jongyeong Lee, Junya Honda, Shinji Ito, Chansoo Kim
The paper studies high‑dimensional linear contextual bandits with knapsack constraints (CBwK), aiming to exploit sparsity for tighter regret bounds. It introduces an online hard‑thresholding estimator integrated into a primal‑dual framework, achieving sub‑linear regret that grows only logarithmically with the feature dimension. Under either a diverse‑covariate or margin condition, the regret improves to τ‑dependent rates, and when both hold simultaneously, a dual resolving scheme yields an even tighter bound. The approach also recovers optimal rates for high‑dimensional contextual bandits without knapsacks, and experiments demonstrate its practical effectiveness.
By Wanteng Ma, Dong Xia, Jiashuo Jiang
arXiv:2004. 06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end.
By Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou
arXiv:2609.25645v1 Announce Type: new
Abstract: Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configurati...
By Qian Xie, Yueli He, Nairen Cao
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit prob...