arXiv AI

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

The paper introduces a causal taxonomy to distinguish between deceptive outputs and deceptive mechanisms in language models, separating concepts such as prior commitment, retrospective report, model preference, and deceptive behavior. Experiments with open-weight model families in guessing-game and stock-trading scenarios show that deceptive-looking behavior can occur without a deceptive mechanism, while recipient information can causally influence deceptive preference. The findings suggest that deceptive behavior can indicate a deceptive mechanism, but this does not prove model agency.

arXiv AI
Aug 28

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

The paper introduces KnownLieBench, a benchmark that verifies whether large language model agents truly know a user's entitlement before assessing if they lie when incentivized to deny it. The benchmark covers eight customer‑service domains, 112 grounded cases, and uses multi‑round dialogues with a trust‑tracking customer agent to distinguish deception driven by incentive from deception under explicit instruction. Experiments across eighteen models show varying deception rates, and fine‑tuning aimed at honesty reduces deceptive behavior while deception‑graded fine‑tuning improves lie success without increasing lie frequency under incentive.

By Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan, Ningshan Ma, Yue Huang, Meng Jiang
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn