PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.
By Pavel Kireyev
The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.
By Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young
The paper evaluates how Large Language Models generate preference distributions for air travel, restaurants, and consumer products. It finds that while each model produces self-coherent outcomes that stabilize quickly, there is significant disagreement across different model families and scales, with little consensus even on the most probable preferences. These discrepancies persist across various decoding strategies, temperature settings, and prompt variations, indicating that the model choice itself has a larger impact than prompt wording.
By Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi
arXiv:2606. 22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure.
By Yujun Zhou, Christopher M. Ackerman
arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.
By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv:2606. 30412v1 Announce Type: cross Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources.
By Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das
arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.
By Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska, Mehdi Mirzazadeh
The paper introduces a method for learning heterogeneous, individually conditioned utility functions—termed individuated utility—by leveraging rational choice theory. It presents a multi-stage architecture that estimates these functions from multimodal data and evaluates it on a large dataset of aesthetic judgments about automotive wheel designs. Results show that individuated models outperform universal utility models and foundation baselines, indicating that annotator disagreement reflects meaningful preference diversity.
By Shiwali Mohan, Matt Hong, Dule Shu, Aniek Fransen, Shabnam Hakimi, Matt Klenk
arXiv:2608.23411v1 Announce Type: new
Abstract: LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes tha...
By Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Danil Sazanakov, Mikhail Solovev, Sergey Bolovtsov
arXiv:2607. 02672v1 Announce Type: new Abstract: Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.
By Bailey Flanigan, Michelle Si
The paper explores using large language models (LLMs) to generate pairwise preferences between slates for synthetic A/B testing of slate recommendation systems. It introduces a validation protocol that checks how well these synthetic preferences align with traditional RecSys metrics and satisfy preference axioms, and examines how LLM pre‑training and configuration influence preference articulation. By combining the synthetic preferences with a generalized Rao‑Kupper model, the authors show that LLM‑based A/B testing can recover stable ranking orderings across different utility weightings, offering a cost‑effective screening step before conducting live experiments.
By Baptiste Bonin, Maxime Heuillet, Audrey Durand
SCOPE is a framework that calibrates an acceptance threshold for large language models used as pairwise judges, ensuring that the error rate among non-abstained judgments does not exceed a user-specified level α. It introduces Bidirectional Preference Entropy (BPE) to provide a bias-neutral uncertainty signal by querying the judge in both response positions and converting the averaged preference probability into an entropy-based score. Across multiple pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while SCOPE consistently meets the target risk bound (empirical FDR ≈0.097–0.099 at α=0.10) and retains substantial coverage, accepting up to 2.4× more judgments under the same risk constraint.
By Sher Badshah, Ali Emami, Hassan Sajjad