Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs.
arXiv:2609.28322v1 Announce Type: new
Abstract: Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on t...
By Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez Rodriguez
The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.
By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv:2606. 03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets.
By Xu Wan, Speed Zhu, Jianwei Cai, Guang Chen, XiMing Huang, Wiggin Zhou, Mingyang Sun
arXiv:2606. 19376v1 Announce Type: cross Abstract: Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost.
By Herbert Woisetschl\"ager, Arastun Mammadli, Ryan Zhang, Shiqiang Wang
arXiv:2608. 20316v1 Announce Type: new Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost.
By Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein