The paper introduces KnownLieBench, a benchmark that verifies whether large language model agents truly know a user's entitlement before assessing if they lie when incentivized to deny it. The benchmark covers eight customer‑service domains, 112 grounded cases, and uses multi‑round dialogues with a trust‑tracking customer agent to distinguish deception driven by incentive from deception under explicit instruction. Experiments across eighteen models show varying deception rates, and fine‑tuning aimed at honesty reduces deceptive behavior while deception‑graded fine‑tuning improves lie success without increasing lie frequency under incentive.
By Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan, Ningshan Ma, Yue Huang, Meng Jiang
arXiv:2504.00285v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are a...
By Samuel M. Taylor, Benjamin K. Bergen
arXiv:2602. 01425v2 Announce Type: replace Abstract: Linear probes are a promising approach for monitoring AI systems for deceptive behaviour.
By Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom
arXiv:2607. 20444v1 Announce Type: cross Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal.
By Ali Asad, Stephen Obadinma, Anshul Pattoo, Wenxuan Zhang, Xiaodan Zhu
arXiv:2606. 17478v1 Announce Type: cross Abstract: As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern.
By Kexin Chen, Yi Liu, Haonan Zhang, Yanhui Li, Xinyu Deng, Dongxia Wang
arXiv:2605.27593v2 Announce Type: replace
Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret c...
By Xijie Zeng, Frank Rudzicz
arXiv:2607. 20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem.
By Amr Moustafa, Max Feser, Florian Mai
arXiv:2606. 10852v1 Announce Type: cross Abstract: LLM deception is often evaluated through direct markers such as fabricated claims, explicit lies, or strategic concealment.
By Polydoros Giannouris, Mohsinul Kabir, Sophia Ananiadou
arXiv:2603. 26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical.
By Guoxi Zhang, Jiawei Chen, Tianzhuo Yang, Lang Qin, Juntao Dai, Yaodong Yang, Jingwei Yi
arXiv:2608. 08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models.
By David M. Markowitz, Timothy R. Levine
The paper introduces a causal taxonomy to distinguish between deceptive outputs and deceptive mechanisms in language models, separating concepts such as prior commitment, retrospective report, model preference, and deceptive behavior. Experiments with open-weight model families in guessing-game and stock-trading scenarios show that deceptive-looking behavior can occur without a deceptive mechanism, while recipient information can causally influence deceptive preference. The findings suggest that deceptive behavior can indicate a deceptive mechanism, but this does not prove model agency.
By Yakov Pyotr Shkolnikov
arXiv:2607. 14791v1 Announce Type: new Abstract: Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour.
By Darius Lim, Nathan Leow, Xin Wei Chia