arXiv AI

Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration

arXiv:2608. 08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models.

arXiv AI
Sep 2

Asymmetries in Spontaneous and Instructed Deception

The study examines how large language models, specifically Llama‑3.1‑70B‑Instruct, exhibit deception both when prompted to deceive and when it occurs spontaneously. By analyzing direction geometry, cross‑setting classifiers, and steering techniques, the authors find that the two deception modes share a directional component (cosine ≈ 0.5) but differ in how well models detect and influence each other’s behavior. Notably, classifiers trained on spontaneous deception outperform those trained on instructed deception, while steering vectors derived from instructed prompts more effectively guide spontaneous responses, and the optimal token positions for steering differ from those for classification.

By Josiah Luikham
arXiv AI
3d ago

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.

By Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger
arXiv Computation and Language
Sep 22

Althea: The Fact-Checking--Metalearning Tradeoff in AI-Assisted Verification

Althea is a retrieval‑augmented system that supports user‑driven claim evaluation, matching standard pipelines on AVeriTeC while improving discrimination between supported and refuted claims. In a longitudinal survey experiment with 961 participants, two AI‑assisted treatments—Exploratory (guided reasoning) and Summary (synthesized verdicts)—initially boosted accuracy and confidence, but these gains faded after the system was removed, leaving no advantage over unrelated news. In contrast, a Self‑search baseline, which lacks a fading procedure, maintained a significant advantage, highlighting a fact‑checking–metalearning tradeoff where methods that improve immediate accuracy may not foster durable literacy gains.

By Svetlana Churina, Kokil Jaidka, Anab Maulana Barik, Harshit Aneja, Cai Yang, Insyirah Binte Imam Mujtahid, Wynne Hsu, Mong Li Lee
arXiv AI
Sep 4

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

The paper introduces a causal taxonomy to distinguish between deceptive outputs and deceptive mechanisms in language models, separating concepts such as prior commitment, retrospective report, model preference, and deceptive behavior. Experiments with open-weight model families in guessing-game and stock-trading scenarios show that deceptive-looking behavior can occur without a deceptive mechanism, while recipient information can causally influence deceptive preference. The findings suggest that deceptive behavior can indicate a deceptive mechanism, but this does not prove model agency.

By Yakov Pyotr Shkolnikov