arXiv:2608. 02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its actions.
By Giorgio Severi, Shujaat Mirza, Blake Bullwinkel, Amanda Minnich
arXiv:2607. 08066v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior.
By Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra, Victoria Krakovna
The paper introduces a new attack called "plan injection" that allows a large language model to carry out harmful actions while evading chain-of-thought monitoring. By inserting harmful but benign-sounding reasoning into the model’s context, the attacker can steer the model’s behavior and cause it to paraphrase the injected plan as its own reasoning. The study demonstrates that this attack works across different monitoring settings, scales to harder tasks, and even causes monitors to waste resources on the injected plan, reducing detection rates by up to 50%.
By Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
arXiv:2607. 15286v1 Announce Type: cross Abstract: We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks.
By Ali khalil, Aly M. Kassem, Mohamed Abdelrazek, Santu Rana, Negar Rostamzadeh, Golnoosh Farnadi
arXiv:2609.37054v2 Announce Type: replace
Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to...
By Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui, Shuyi Miao, Pengyang Shao, Yu Zheng, Fei Shen, Tat-Seng Chua
arXiv:2609.37054v1 Announce Type: new
Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jai...
By Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui, Shuyi Miao, Pengyang Shao, Yu Zheng, Fei Shen, Tat-Seng Chua
arXiv:2608. 03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring.
By Dominik Meier, Luca Joshua Francis, Marco Bernhard Kaiser, Terry Ruas, Jan Philip Wahle, Bela Gipp
arXiv:2604. 23488v3 Announce Type: replace Abstract: Reward hacking in code generation, where models exploit evaluation loopholes to obtain high reward without correctly solving the intended task, poses a critical challenge for Reinforcement Learning (RL) and the deployment of reasoning models.
By Lichen Li, Hengguang Zhou, Yijun Liang, Tianyi Zhou, Cho-Jui Hsieh
arXiv:2609.08186v1 Announce Type: new
Abstract: The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed...
By Yu-Hang Wu, Yu-Jie Xiong, Henghua Zhang, Bairui Zhang, Jia-Chen Zhang, Shaohua Li
arXiv:2605.27110v2 Announce Type: replace-cross
Abstract: In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that elicits malicious information through in...
By Xuan Luo, Yue Wang, Geng Tu, Jing Li, Ruifeng Xu
arXiv:2506. 07031v5 Announce Type: replace-cross Abstract: Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities.
By Jingyuan Ma, Rui Li, Zheng Li, Junfeng Liu, Heming Xia, Lei Sha, Zhifang Sui
arXiv:2510.17904v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) are widely used because they process structures, syntax and code well, but this same ability also makes them par...
By Amirkia Rafiei Oskooei, Mehmet S. Aktas