arXiv Machine Learning By Bo Xue, Zhi Hong, Jiayi Li, Yuanyu Wan, Ji Cheng, Shuang Qiu

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation

Read the original on arXiv Machine Learning →

arXiv:2608. 04333v1 Announce Type: new Abstract: Large language model (LLM) configuration evaluation is challenging due to limited evaluation budgets, varying costs, and multiple competing objectives.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv Machine Learning
Aug 27

Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification

The paper presents a computationally efficient optimal design framework for multinomial logit (MNL) bandits, addressing the combinatorial action space that makes traditional methods infeasible. It introduces two approaches: an exact or certified-approximate mixed-integer linear program with solver‑certified early stopping, and a fully polynomial‑time lifted design using a tractable surrogate objective. Leveraging the Kiefer‑Wolfowitz equivalence theorem, the authors provide near G‑optimality guarantees and apply the framework to develop a best assortment identification algorithm with an instance‑dependent sample complexity of τO((d log N)/Δ²).

By Joongkyu Lee, Min-hwan Oh