arXiv:2606. 03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets.
By Xu Wan, Speed Zhu, Jianwei Cai, Guang Chen, XiMing Huang, Wiggin Zhou, Mingyang Sun
The paper proposes an inference auction for large language model (LLM) APIs, enabling users to bid for faster service when compute demand exceeds capacity. The auction allocates priority efficiently without increasing latency, and includes fast algorithms for truthful bidding and an autobidding agent that adjusts bids within a user’s budget to maximize utility. Experiments show the auction improves system welfare while preserving the cache utilization and latency benefits of the SGLang inference framework.
By Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan
arXiv:2608. 13315v1 Announce Type: cross Abstract: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit.
By Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu
We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency.
arXiv:2606. 19376v1 Announce Type: cross Abstract: Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost.
By Herbert Woisetschl\"ager, Arastun Mammadli, Ryan Zhang, Shiqiang Wang
arXiv:2608. 13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time.
By Heming Fu, Shan Lin, Qianqian Xie, Guojun Xiong
The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.
By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv:2511. 00847v5 Announce Type: replace-cross Abstract: The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers.
By Yuhan Cao, Yu Wang, Sitong Liu, Miao Li, Yixin Tao, Tianxing He
arXiv:2608. 12719v1 Announce Type: cross Abstract: Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows.
By Haolong Chen, Zhengyuan Xin, Liang Zhang, Lei Xue, Guangxu Zhu
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs.
arXiv:2608. 20316v1 Announce Type: new Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost.
By Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
arXiv:2606. 02581v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) faces a fundamental three-way tension: deeper retrieval improves factual grounding but inflates token costs and end-to-end latency.
By Sanjay Mishra