arXiv AI

One Probe Won't Catch Them All: Towards Targeted Deception Detection

arXiv:2602. 01425v2 Announce Type: replace Abstract: Linear probes are a promising approach for monitoring AI systems for deceptive behaviour.

arXiv AI
Jul 3

Scaling Trends for Lie Detector Oversight in Preference Learning

arXiv:2607. 01567v1 Announce Type: new Abstract: Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers.

By Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day, Adam Gleave, Chris Cundy