arXiv AI By Amr Moustafa, Max Feser, Florian Mai

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

Read the original on arXiv AI →

arXiv:2607. 20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 4

Probe Generalization as Subspace Selection for OOD Deception Detection

Linear probes can identify behaviors in language model activations but often fail on out‑of‑distribution data. This study shows that projecting inputs onto a small set of principal components (PCs) from the training distribution allows probes for Llama‑3.1‑8B‑Instruct to transfer across three deception‑detection datasets, nearly matching probes trained directly on the test data. By scoring PCs with an LLM judge to select those that encode transferable deception directions, the authors close the baseline‑to‑oracle gap by 78% on Insider Trading Report and 25% on Sandbagging, revealing that subspace selection largely determines OOD robustness.

By Daniel Yoo, Adrians Skapars
arXiv AI
Jul 3

Scaling Trends for Lie Detector Oversight in Preference Learning

arXiv:2607. 01567v1 Announce Type: new Abstract: Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers.

By Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day, Adam Gleave, Chris Cundy
arXiv AI
3d ago

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.

By Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger