arXiv Machine Learning

Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption

arXiv AI
2d ago

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

The paper introduces Drift‑Aware Sparse Routing (DRS), a method for routing requests in a multi‑model language service while respecting compute, latency, memory, or cost budgets. DRS estimates reward and resource use from a rolling audit window, routes using pessimistic reward and optimistic cost estimates, updates resource shadow prices online, and applies a hard meter before commitment. The authors provide theoretical regret bounds that separate control from statistics, showing how the method adapts to non‑stationary prompt distributions and model changes.

By Cheung Hao Lee, Patrick Wong
arXiv Machine Learning
Jun 3

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).

By Oleg Miroshnichenko
arXiv Computation and Language
3d ago

LLP: LLM-Based Product Pricing in E-commerce

The paper introduces LLP, a Large Language Model–based generative framework for pricing second‑hand products on consumer‑to‑consumer platforms. LLP retrieves similar items to capture market dynamics, then uses LLMs to generate price suggestions, refined through supervised fine‑tuning and group relative policy optimization. A confidence‑based filter rejects unreliable predictions, and experiments show LLP outperforms prior methods, achieving higher static adoption rates when deployed on Xianyu.

By Hairu Wang, Sheng You, Qiheng Zhang, Xike Xie, Shuguang Han, Yuchen Wu, Fei Huang, Jufeng Chen
arXiv AI
Jun 10

A Theory of Training Profit-Optimal LLMs

arXiv:2605. 16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure.

By Sophie Hao, William Merrill
arXiv AI
Jul 14

Efficient Online Proportional Sampling with Applications to Smoothed Online Learning

arXiv:2607. 10963v1 Announce Type: cross Abstract: We study the problem of efficient online proportional sampling from a high-dimensional domain under a $\sigma$-smoothed adversary, where the sampling distribution is induced by a dynamically evolving weight function defined over a sequence of piecewise-structured partitions.

By Amirmahdi Mirfakhar, Maria-Florina Balcan, Hedyeh Beyhaghi