arXiv Machine Learning By Guannan Lai, Haoran Hu, Long Chen, Zhenguo Li, Han-Jia Ye

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing

Read the original on arXiv Machine Learning →

arXiv:2606. 06924v1 Announce Type: new Abstract: Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

The paper introduces SaveRouter, a sparse‑supervision framework for large language model routing that selectively gathers informative model feedback and shares capability information across related queries. By using only about 33–41% of available training feedback, SaveRouter achieves competitive or superior routing quality while reducing the break‑even deployment volume by 1.9–9.5× compared to conventional routers. The study also shows that the supervision level that minimizes serving cost may differ from the one that yields the earliest payback.

By Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye
arXiv AI
3d ago

FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

FlexRouter is a routing framework for large language models that explicitly models model complementarity to maximize answer coverage. It formulates routing as a coverage-oriented subset selection problem and uses Determinantal Point Processes to capture both competence and redundancy. During inference, a greedy strategy based on marginal log-determinant gains allows the router to adaptively determine subset sizes without a fixed budget, achieving higher coverage with lower redundancy on the RouterEval benchmark.

By Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry
arXiv AI
Jul 24

Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating

arXiv:2607. 20481v1 Announce Type: new Abstract: Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained routers or collaboration-aware finetuning that tie routing behavior to a particular operating regime.

By Evan Chen, Shiqiang Wang, Kevin S Chan, Su Wang, Christopher Brinton