arXiv:2608. 02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its actions.
By Giorgio Severi, Shujaat Mirza, Blake Bullwinkel, Amanda Minnich
arXiv:2607. 08066v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior.
By Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra, Victoria Krakovna
The paper introduces a new attack called "plan injection" that allows a large language model to carry out harmful actions while evading chain-of-thought monitoring. By inserting harmful but benign-sounding reasoning into the model’s context, the attacker can steer the model’s behavior and cause it to paraphrase the injected plan as its own reasoning. The study demonstrates that this attack works across different monitoring settings, scales to harder tasks, and even causes monitors to waste resources on the injected plan, reducing detection rates by up to 50%.
By Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
arXiv:2607. 15286v1 Announce Type: cross Abstract: We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks.
By Ali khalil, Aly M. Kassem, Mohamed Abdelrazek, Santu Rana, Negar Rostamzadeh, Golnoosh Farnadi
arXiv:2609.37054v2 Announce Type: replace
Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to...
By Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui, Shuyi Miao, Pengyang Shao, Yu Zheng, Fei Shen, Tat-Seng Chua
arXiv:2609.37054v1 Announce Type: new
Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jai...
By Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui, Shuyi Miao, Pengyang Shao, Yu Zheng, Fei Shen, Tat-Seng Chua