arXiv AI By Yifei Li, Zihui Gao, Laks V. S. Lakshmanan

WISERouter: LLM Routing with Workload Budget Constraint

Read the original on arXiv AI →

arXiv:2607. 23765v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.