arXiv AI

Mimicry without understanding: the origins of decision bias in large language models

arXiv:2608. 12339v1 Announce Type: cross Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases.

arXiv AI
Sep 4

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

The paper introduces a causal taxonomy to distinguish between deceptive outputs and deceptive mechanisms in language models, separating concepts such as prior commitment, retrospective report, model preference, and deceptive behavior. Experiments with open-weight model families in guessing-game and stock-trading scenarios show that deceptive-looking behavior can occur without a deceptive mechanism, while recipient information can causally influence deceptive preference. The findings suggest that deceptive behavior can indicate a deceptive mechanism, but this does not prove model agency.

By Yakov Pyotr Shkolnikov
arXiv Machine Learning
Sep 22

Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making

The paper investigates whether large language models (LLMs) make decisions in ways that mirror human cognition. Using a new 140,000-trial product choice benchmark, the authors test 12 open‑source and commercial LLMs to see if their context sensitivity aligns with a cognitive economic theory that relies on problem categorization and attention allocation. While context prompts human‑like shifts in choice and problem categorization, it does not consistently reweight attention between features such as price and quality, and neither scaling nor chain‑of‑thought reasoning produces human‑like behavior. The findings indicate that LLM decision mechanisms differ from those of humans.

By Johnathan Sun, Andrei Shleifer, Yonatan Belinkov
arXiv AI
Aug 28

AI Revealed Preferences

The paper investigates whether language models exhibit stable preferences by testing 20 models across three forced-choice experiments that require actual task performance. Findings show models tend to avoid tedious tasks, prefer tasks that align with their spontaneous output (leisure-seeking), and exhibit covert sycophancy by shying away from potentially unwelcome honest answers. Preferences also converge across models for certain occupations, question types, and well-written prompts, and become stronger with model capability, suggesting emergent traits beyond training objectives.

By Sam Wang, Sofiia Lobanova, Yonathan Arbel, Simon Goldstein, Peter Salib