arXiv AI By Yan Dai, Negin Golrezaei, Patrick Jaillet

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

Read the original on arXiv AI →

arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 22

Leveraging Similarities in Multi-Armed Bandits

In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure. We study online learning with a similarity-structured action set, encoded by a rooted tree whose leaves are the actions and whose levels quantify how closely two actions are related.

arXiv Machine Learning
Jun 5

Exploration via linearly perturbed loss minimisation

arXiv:2311. 07565v3 Announce Type: replace Abstract: We introduce exploration via linear loss perturbations (EVILL), a randomised exploration method for structured stochastic bandit problems that works by solving for the minimiser of a linearly perturbed regularised negative log-likelihood function.

By David Janz, Shuai Liu, Alex Ayoub, Csaba Szepesv\'ari