arXiv AI By Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier

LLM-Derived Preference Judgments Are Not Self-Consistent

Read the original on arXiv AI →

The paper investigates whether large language models (LLMs) produce self‑consistent numerical preference judgments when asked to estimate a person’s willingness to pay for items. By comparing differences in stated willingness to pay with the payment that would make a person indifferent between items, the authors develop statistical tests to measure deviations from a single utility function. Experiments on flight, apartment, and hotel scenarios across six LLMs show persistent inconsistencies, indicating that LLM‑derived preference judgments cannot be reliably summarized by one utility function.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.

By Pavel Kireyev
arXiv Machine Learning
Sep 22

Multiple latent orderings better predict language model preferences

The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.

By Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young
arXiv AI
2d ago

Evaluating LLM-Generated Preference Distributions

The paper evaluates how Large Language Models generate preference distributions for air travel, restaurants, and consumer products. It finds that while each model produces self-coherent outcomes that stabilize quickly, there is significant disagreement across different model families and scales, with little consensus even on the most probable preferences. These discrepancies persist across various decoding strategies, temperature settings, and prompt variations, indicating that the model choice itself has a larger impact than prompt wording.

By Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi
arXiv AI
Jun 17

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv AI
Jun 30

Can LLMs Rank? A Tale of Triads and Triage

arXiv:2606. 30412v1 Announce Type: cross Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources.

By Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das