The paper introduces SaveRouter, a sparse‑supervision framework for large language model routing that selectively gathers informative model feedback and shares capability information across related queries. By using only about 33–41% of available training feedback, SaveRouter achieves competitive or superior routing quality while reducing the break‑even deployment volume by 1.9–9.5× compared to conventional routers. The study also shows that the supervision level that minimizes serving cost may differ from the one that yields the earliest payback.
By Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye
arXiv:2609.37362v1 Announce Type: new
Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--eff...
By Guannan Lai, Han-Jia Ye
arXiv:2603.04445v3 Announce Type: replace-cross
Abstract: The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent mo...
By Yasmin Moslem, John D. Kelleher
FlexRouter is a routing framework for large language models that explicitly models model complementarity to maximize answer coverage. It formulates routing as a coverage-oriented subset selection problem and uses Determinantal Point Processes to capture both competence and redundancy. During inference, a greedy strategy based on marginal log-determinant gains allows the router to adaptively determine subset sizes without a fixed budget, achieving higher coverage with lower redundancy on the RouterEval benchmark.
By Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry
arXiv:2607. 20481v1 Announce Type: new Abstract: Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained routers or collaboration-aware finetuning that tie routing behavior to a particular operating regime.
By Evan Chen, Shiqiang Wang, Kevin S Chan, Su Wang, Christopher Brinton
arXiv:2609.23085v1 Announce Type: cross
Abstract: Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajector...
By Muhammad Abdur Rab Siddiqui, Daniela Rojas, Chen Yang, Wenqi Cui, Yuanyuan Shi, Yize Chen