arXiv:2607. 20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem.
By Amr Moustafa, Max Feser, Florian Mai
arXiv:2606. 12618v1 Announce Type: new Abstract: Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say.
By Alan Cooney, David Africa, Geoffrey Irving
arXiv:2602. 01425v2 Announce Type: replace Abstract: Linear probes are a promising approach for monitoring AI systems for deceptive behaviour.
By Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom
arXiv:2607. 20444v1 Announce Type: cross Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal.
By Ali Asad, Stephen Obadinma, Anshul Pattoo, Wenxuan Zhang, Xiaodan Zhu
arXiv:2606. 17478v1 Announce Type: cross Abstract: As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern.
By Kexin Chen, Yi Liu, Haonan Zhang, Yanhui Li, Xinyu Deng, Dongxia Wang
arXiv:2605. 06846v3 Announce Type: replace-cross Abstract: Recent work identifies secret loyalties as a distinct threat from standard backdoors.
By Alfie Lamerton, Fabien Roger