arXiv:2607. 13346v1 Announce Type: cross Abstract: Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored.
By Aman Mehta
arXiv:2606. 12618v1 Announce Type: new Abstract: Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say.
By Alan Cooney, David Africa, Geoffrey Irving
arXiv:2606. 24952v1 Announce Type: cross Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should be able to modify it.
By Cosimo Galeone, Anna Ettorre, Minsu Park, Giuseppe Ettorre, Daniele Ligorio
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance.
arXiv:2608. 08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models.
By David M. Markowitz, Timothy R. Levine
arXiv:2607. 20379v1 Announce Type: new Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it.
By Hiskias Dingeto
arXiv:2608. 17202v1 Announce Type: new Abstract: Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably.
By Mark Russinovich
arXiv:2606. 05976v1 Announce Type: new Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources.
By Kuan-Yen Chen, Fang-Yi Su, Jung-Hsien Chiang
arXiv:2607. 04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.
By Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
arXiv:2604. 08169v2 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.
By Niklas Herbster, Martin Zborowski, Alberto Tosato, Gauthier Gidel, Tommaso Tosato
arXiv:2608. 04066v1 Announce Type: new Abstract: How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust?
By Mohsen Arjmandi
arXiv:2607. 15434v1 Announce Type: cross Abstract: Multi-agent systems routinely place one AI agent in authority over another.
By Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh