arXiv AI By Cheung Hao Lee, Patrick Wong

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

Read the original on arXiv AI →

The paper introduces Drift‑Aware Sparse Routing (DRS), a method for routing requests in a multi‑model language service while respecting compute, latency, memory, or cost budgets. DRS estimates reward and resource use from a rolling audit window, routes using pessimistic reward and optimistic cost estimates, updates resource shadow prices online, and applies a hard meter before commitment. The authors provide theoretical regret bounds that separate control from statistics, showing how the method adapts to non‑stationary prompt distributions and model changes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.