arXiv AI By Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang

Dynamic LLM Routers are Often Misguided

Read the original on arXiv AI →

Dynamic LLM routers aim to reduce inference costs by directing each query to the cheapest capable model. In a study of six commercial routers across 14 settings and eight task categories, none surpassed a simple random router that selects between two well-chosen models at the same cost, with some underperforming by over 10 percentage points. The authors identify four common patterns—difficulty blindness, length reversal, semantic matching, and roster suboptimality—that explain this gap and propose a new evaluation method and a simple two-model router that mitigates these patterns, though its advantage over random routing remains modest.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

Most of the LLM routing gap is task type

The paper investigates why large‑language‑model (LLM) routers—systems that select the best model for each query—often fail to outperform a single best model. By evaluating 14 models on 294 questions across seven task types and three languages, the authors find that a simple static mapping of task type to model improves 21 of the 29 questions that routing could potentially solve, and that learned routers do not significantly exceed this performance. The study highlights that most routing gains stem from task‑type specialization rather than complex learned decision rules.

By Janghoon Lee
arXiv AI
6d ago

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

The paper introduces SaveRouter, a sparse‑supervision framework for large language model routing that selectively gathers informative model feedback and shares capability information across related queries. By using only about 33–41% of available training feedback, SaveRouter achieves competitive or superior routing quality while reducing the break‑even deployment volume by 1.9–9.5× compared to conventional routers. The study also shows that the supervision level that minimizes serving cost may differ from the one that yields the earliest payback.

By Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye
arXiv Computation and Language
Aug 27

VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval

The paper introduces VDAR-Router, a routing framework for large language models that uses verbalized query difficulty analysis to guide model selection. It first generates an explicit difficulty profile for each query, retrieves historical examples with similar profiles, and then estimates model suitability to choose a model based on a reward function balancing performance and cost. Experiments on three datasets show that VDAR-Router consistently outperforms existing baselines in cost‑performance trade‑offs, and case studies confirm that explicit difficulty analysis improves example relevance and routing reliability.

By Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng, An-Zi Yen