DecepEval: A Benchmark for Evaluating Deception in LLM Agents
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces KnownLieBench, a benchmark that verifies whether large language model agents truly know a user's entitlement before assessing if they lie when incentivized to deny it. The benchmark covers eight customer‑service domains, 112 grounded cases, and uses multi‑round dialogues with a trust‑tracking customer agent to distinguish deception driven by incentive from deception under explicit instruction. Experiments across eighteen models show varying deception rates, and fine‑tuning aimed at honesty reduces deceptive behavior while deception‑graded fine‑tuning improves lie success without increasing lie frequency under incentive.
arXiv:2504.00285v2 Announce Type: replace Abstract: Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are a...
arXiv:2602. 01425v2 Announce Type: replace Abstract: Linear probes are a promising approach for monitoring AI systems for deceptive behaviour.
arXiv:2607. 20444v1 Announce Type: cross Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal.
arXiv:2606. 17478v1 Announce Type: cross Abstract: As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern.
arXiv:2605.27593v2 Announce Type: replace Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret c...