← Back to all news
arXiv Machine Learning September 10, 2026 By Wang Wei, Soumyabrata Pal, Koyel Mukherjee, Franck Dernoncourt, Ryan A. Rossi, Branislav Kveton, Hoda Eldardiry

Online Learning with LLM Experts from Limited Feedback

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 28

WISERouter: LLM Routing with Workload Budget Constraint

arXiv:2607. 23765v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale.

By Yifei Li, Zihui Gao, Laks V. S. Lakshmanan
llmsreinforcement-learning
More like this →
arXiv Machine Learning
Jun 30

Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits

arXiv:2602. 02061v2 Announce Type: replace Abstract: Explosive demands for LLMs often cause user queries to accumulate in server queues, requiring efficient routing (query-LLM matching) and scheduling (query prioritization) mechanisms.

By Seoungbin Bae, Junyoung Son, Dabeen Lee
llmsrag
More like this →
Hugging Face Trending Papers
Aug 17

Coverage-Maximizing Multinomial Subset Routing under Operational Constraints

We introduce Multinomial Subset Routing (MSR), a new online routing framework over $K$ experts in which the learner keeps a multinomial routing policy instead of a deterministic subset of experts. At each round, the learner samples $M$ experts i.

reinforcement-learning
More like this →
arXiv AI
Jun 16

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.

By Yan Dai, Negin Golrezaei, Patrick Jaillet
ragreinforcement-learningsafety
More like this →
arXiv AI
Jul 13

Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing

arXiv:2607. 09015v1 Announce Type: cross Abstract: We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing.

By Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi, Vijaykrishnan Narayanan
llmsreinforcement-learningbenchmarks
More like this →
arXiv AI
Aug 18

Coverage-Maximizing Multinomial Subset Routing under Operational Constraints

arXiv:2608. 16375v1 Announce Type: cross Abstract: We introduce Multinomial Subset Routing (MSR), a new online routing framework over $K$ experts in which the learner keeps a multinomial routing policy instead of a deterministic subset of experts.

By Quan Zhou, Yiyan Huang
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea