The paper introduces SaveRouter, a sparse‑supervision framework for large language model routing that selectively gathers informative model feedback and shares capability information across related queries. By using only about 33–41% of available training feedback, SaveRouter achieves competitive or superior routing quality while reducing the break‑even deployment volume by 1.9–9.5× compared to conventional routers. The study also shows that the supervision level that minimizes serving cost may differ from the one that yields the earliest payback.
By Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye
The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.
By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv:2605. 17106v2 Announce Type: replace-cross Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences.
By Aashna Garg, Siddharth Singha Roy, Jinu Jang, Federico Brancasi, Shengyu Fu
arXiv:2606. 02581v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) faces a fundamental three-way tension: deeper retrieval improves factual grounding but inflates token costs and end-to-end latency.
By Sanjay Mishra
arXiv:2603. 20895v3 Announce Type: replace-cross Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty.
By Tanay Varshney, Annie Surla, Michelle Xu, Gomathy Venkata Krishnan, Maximilian Jeblick, David Austin, Neal Vaidya, Davide Onofrio
arXiv:2608.28726v1 Announce Type: new
Abstract: The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant chal...
By Xinyuan Gui, Shaowen Wang, Sheng Sun, Zijian Wang, Zishu Yu, Zheming Yang
arXiv:2606. 07587v1 Announce Type: new Abstract: LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by dynamically selecting a model for each query.
By Yifan Lu, Qiyue Zhang, Shenrun Zhang, Zhibo Yu, Zhuang Wang, Hanjie Chen, Jiarong Xing
Dynamic LLM routers aim to reduce inference costs by directing each query to the cheapest capable model. In a study of six commercial routers across 14 settings and eight task categories, none surpassed a simple random router that selects between two well-chosen models at the same cost, with some underperforming by over 10 percentage points. The authors identify four common patterns—difficulty blindness, length reversal, semantic matching, and roster suboptimality—that explain this gap and propose a new evaluation method and a simple two-model router that mitigates these patterns, though its advantage over random routing remains modest.
By Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang
arXiv:2606. 18774v1 Announce Type: new Abstract: We present RouteJudge, an online pairwise preference evaluation framework for LLM routing systems, with a public platform available at https://routejudge.
By Guannan Lai, Haoran Hu, Han-Jia Ye
arXiv:2609.37362v1 Announce Type: new
Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--eff...
By Guannan Lai, Han-Jia Ye
arXiv:2608. 14641v1 Announce Type: new Abstract: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison.
By Kiran N. Kumar, Santhosh K. Saminathan
The paper introduces VDAR-Router, a routing framework for large language models that uses verbalized query difficulty analysis to guide model selection. It first generates an explicit difficulty profile for each query, retrieves historical examples with similar profiles, and then estimates model suitability to choose a model based on a reward function balancing performance and cost. Experiments on three datasets show that VDAR-Router consistently outperforms existing baselines in cost‑performance trade‑offs, and case studies confirm that explicit difficulty analysis improves example relevance and routing reliability.
By Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng, An-Zi Yen