arXiv AI

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

arXiv:2607. 20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem.

arXiv Computation and Language
Sep 4

Probe Generalization as Subspace Selection for OOD Deception Detection

Linear probes can identify behaviors in language model activations but often fail on out‑of‑distribution data. This study shows that projecting inputs onto a small set of principal components (PCs) from the training distribution allows probes for Llama‑3.1‑8B‑Instruct to transfer across three deception‑detection datasets, nearly matching probes trained directly on the test data. By scoring PCs with an LLM judge to select those that encode transferable deception directions, the authors close the baseline‑to‑oracle gap by 78% on Insider Trading Report and 25% on Sandbagging, revealing that subspace selection largely determines OOD robustness.

By Daniel Yoo, Adrians Skapars
arXiv AI
Jul 3

Scaling Trends for Lie Detector Oversight in Preference Learning

arXiv:2607. 01567v1 Announce Type: new Abstract: Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers.

By Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day, Adam Gleave, Chris Cundy
arXiv AI
3d ago

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.

By Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger
arXiv Computation and Language
Aug 24

Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift

The paper introduces a new evaluation setting called scenario‑level out‑of‑distribution (SL‑OOD) detection for SMS and voice phishing, where entire attack scenarios are omitted from training while the label space stays fixed. It shows that high in‑distribution performance does not guarantee robustness to unseen scenarios, attributing this to scenario memorization. The authors propose ECoG, an evidence‑consistent generative framework that uses evidence‑span supervision and a rationale‑label consistency objective, achieving notable improvements in Macro‑F1, reduced prediction‑rationale inconsistency, and higher token‑level overlap with reference evidence.

By San Kim, JinYeong Bak