arXiv Machine Learning

CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM

arXiv:2606. 00846v1 Announce Type: new Abstract: Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties.

arXiv AI
4d ago

Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

arXiv:2609.37800v1 Announce Type: cross Abstract: Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two cha...

By Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ek\c{s}i, Rahul Sharma, Julia Mueller, Theresa Dombrowski, Jakob Karolus, Viktor Bengs, Eyke H\"ullermeier, Sebastian Vollmer
arXiv AI
6d ago

Cost-Aware Best-LLM Identification using Dueling Feedback

The paper introduces a new variant of the multi‑armed bandit problem that incorporates dueling feedback—pairwise comparisons of model responses—and heterogeneous sampling costs to identify the best large language model (LLM) from a set with varying query costs. Assuming a Condorcet winner, the authors propose a Track‑and‑Stop style algorithm that guarantees asymptotically optimal cost as the error probability approaches zero. Extensive experiments on synthetic and real‑world data show that this cost‑aware approach consistently outperforms both classical cost‑unaware algorithms and other cost‑aware extensions.

By Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
arXiv Machine Learning
Aug 10

Bootstrap-Conditioned Action Selection with Tabular Foundation Models

arXiv:2608. 06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts.

By Devansh Gupta, Shiv Tavker, Dmitry Efimov, Suchitra Sathyanarayana, Gitanjali Bhutani, Boris N. Oreshkin
arXiv Machine Learning
Aug 19

Latent Order Bandits

The paper introduces Latent Order Bandits (LOB), a new bandit framework that relaxes the strict assumptions of traditional latent bandits by only requiring a partial order of action preferences within each latent state. LOB allows instances sharing the same state to have different reward distributions as long as the action ranking remains consistent, making it suitable for scenarios like user groups on streaming services who agree on genre preferences but rate differently. The authors present an upper‑confidence bound algorithm for both total and partial latent orders, provide regret bounds, and propose a posterior‑sampling variant that empirically outperforms full‑prior latent bandits when reward scales vary across instances sharing the same latent state.

By Emil Carlsson, Newton Mwai, Fredrik D. Johansson
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv AI
Jun 2

ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning

arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.

By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause