arXiv AI

Inference Auctions

The paper proposes an inference auction for large language model (LLM) APIs, enabling users to bid for faster service when compute demand exceeds capacity. The auction allocates priority efficiently without increasing latency, and includes fast algorithms for truthful bidding and an autobidding agent that adjusts bids within a user’s budget to maximize utility. Experiments show the auction improves system welfare while preserving the cache utilization and latency benefits of the SGLang inference framework.

arXiv AI
Sep 24

Learning the Cost of Reliable Inference

arXiv:2609.28322v1 Announce Type: new Abstract: Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on t...

By Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez Rodriguez
arXiv Machine Learning
Aug 10

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

arXiv:2608. 06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.

By Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse
arXiv Machine Learning
Jul 22

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing

arXiv:2604. 05845v2 Announce Type: replace-cross Abstract: Auto-bidding services optimize real-time bidding strategies for advertisers under key performance indicator (KPI) constraints such as target return on investment and budget.

By Linghui Meng, Chun Gan, Shengsheng Niu, Chengcheng Zhang, Chenchen Li, Chuan Yang, Yi Mao, Xin Zhu, Jie He, Zhangang Lin, Ching Law
arXiv Machine Learning
Jul 31

PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective

arXiv:2607. 27265v1 Announce Type: new Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side Platform (DSP) bidding for advertisers, and Ad Exchange conducting auctions between them.

By Shengtian Yang, Yewen Li, Peng Jiang, Zhiyi Lyu, Bo An, Peng Jiang, Qingpeng Cai, Lei Feng
arXiv AI
2d ago

You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.

By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang