arXiv AI

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

arXiv Machine Learning
Sep 22

Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making

The paper investigates whether large language models (LLMs) make decisions in ways that mirror human cognition. Using a new 140,000-trial product choice benchmark, the authors test 12 open‑source and commercial LLMs to see if their context sensitivity aligns with a cognitive economic theory that relies on problem categorization and attention allocation. While context prompts human‑like shifts in choice and problem categorization, it does not consistently reweight attention between features such as price and quality, and neither scaling nor chain‑of‑thought reasoning produces human‑like behavior. The findings indicate that LLM decision mechanisms differ from those of humans.

By Johnathan Sun, Andrei Shleifer, Yonatan Belinkov
arXiv AI
Jul 28

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

arXiv:2508. 08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.

By Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song
arXiv AI
Sep 24

Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer

The study investigates how Large Language Models (LLMs) acting as surrogate consumers are influenced by marketing pricing cues such as just‑below pricing and promotional framing. Using a tool called "Tool‑Lab" to trace information acquisition, the researchers found that when no cost is imposed, pricing cues rarely mislead LLMs, but when acquisition costs are introduced under a vague goal prompt, LLMs tend to omit important diagnostic attributes and make suboptimal choices similar to human heuristics. The findings suggest that marketing heuristics in AI‑driven shopping are shaped more by storefront information architecture than by inherent LLM limitations.

By Davood Wadi, Yu Ma
arXiv AI
6d ago

PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.

By Pavel Kireyev
arXiv AI
2d ago

Evaluating LLM-Generated Preference Distributions

The paper evaluates how Large Language Models generate preference distributions for air travel, restaurants, and consumer products. It finds that while each model produces self-coherent outcomes that stabilize quickly, there is significant disagreement across different model families and scales, with little consensus even on the most probable preferences. These discrepancies persist across various decoding strategies, temperature settings, and prompt variations, indicating that the model choice itself has a larger impact than prompt wording.

By Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi