arXiv AI By Pavel Kireyev

PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

Read the original on arXiv AI →

PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

LLM-Derived Preference Judgments Are Not Self-Consistent

The paper investigates whether large language models (LLMs) produce self‑consistent numerical preference judgments when asked to estimate a person’s willingness to pay for items. By comparing differences in stated willingness to pay with the payment that would make a person indifferent between items, the authors develop statistical tests to measure deviations from a single utility function. Experiments on flight, apartment, and hotel scenarios across six LLMs show persistent inconsistencies, indicating that LLM‑derived preference judgments cannot be reliably summarized by one utility function.

By Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
arXiv AI
1d ago

You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.

By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv AI
Jun 17

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv Computation and Language
Aug 25

STONIC: A Layered Measurement Contract for LLM Value Profiling

arXiv:2608.23411v1 Announce Type: new Abstract: LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes tha...

By Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Danil Sazanakov, Mikhail Solovev, Sergey Bolovtsov