arXiv AI By Bhavtosh Rath

How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

Read the original on arXiv AI →

arXiv:2607. 25068v1 Announce Type: new Abstract: Routing decisions between a cheap heuristic and an expensive large language model (LLM) are typically framed as a difficulty problem: send the hard cases to the expensive path.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.

By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv AI
2d ago

OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework

The paper introduces OSCAR, an LLM‑based framework that translates business descriptions into accurate optimization models while verifying and improving them through a simulator, coder, and reviewer. OSCAR uses a cost‑ordered escalation strategy to select among LLMs of varying price and capability, achieving 95–100% accuracy on benchmark problems with local, open‑weight models. The framework also provides competitive guarantees and token‑cost advantages over existing LLMs like Codex and Claude Code.

By Jinzhi Bu, Haixin Tang, Huanan Zhang
arXiv AI
6d ago

PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.

By Pavel Kireyev