arXiv AI

LLM-Derived Preference Judgments Are Not Self-Consistent

arXiv:2608. 17644v1 Announce Type: new Abstract: Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.

arXiv AI
Jun 17

Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

arXiv:2602. 09802v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists.

By Manon Reusens, Sofie Goethals, Toon Calders, David Martens
arXiv AI
Jun 30

Can LLMs Rank? A Tale of Triads and Triage

arXiv:2606. 30412v1 Announce Type: cross Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources.

By Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das
arXiv AI
Jun 3

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.

By Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska, Mehdi Mirzazadeh
arXiv AI
Jul 21

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.

By Zachary Wojtowicz, Ayush Nayak, Jacob Andreas