arXiv:2605.18859v3 Announce Type: replace-cross
Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...
By Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi
arXiv:2609.28322v1 Announce Type: new
Abstract: Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on t...
By Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez Rodriguez
arXiv:2607. 08665v1 Announce Type: new Abstract: Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle.
By Teng-Ruei Chen
HybridInfer is a thermal‑aware reinforcement‑learning router that selects among on‑device, edge, and cloud large language model tiers based on a phone’s thermal headroom and query complexity. Trained offline with a Q‑learning policy, it balances quality, latency, cost, and a thermal penalty, and includes a locality bonus that encourages on‑device execution. In real Android tests on 210 prompts, the learned router outperformed two hand‑tuned heuristics in quality while maintaining the lowest cost, and it proved more reliable and faster than always‑on‑device inference, especially for long queries.
By Simran Koul
arXiv:2608. 20316v1 Announce Type: new Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost.
By Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
arXiv:2607. 27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure.
By Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi