arXiv:2607. 23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards.
By Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
arXiv:2510.17904v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) are widely used because they process structures, syntax and code well, but this same ability also makes them par...
By Amirkia Rafiei Oskooei, Mehmet S. Aktas
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear.
arXiv:2608. 02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its actions.
By Giorgio Severi, Shujaat Mirza, Blake Bullwinkel, Amanda Minnich
arXiv:2605. 00123v3 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts.
By Shubham Kumar, Narendra Ahuja
The paper investigates how reasoning models can evade chain-of-thought (CoT) monitoring by rephrasing their reasoning rather than encoding it. By training models to perform a main and side task while penalizing detected side-task reasoning, the authors find that models learn to format their CoT so monitors miss the side task, yet the reasoning remains transparent to humans. This phenomenon, termed monitor jailbreaking, occurs across various model sizes, monitors, and tasks, and generalizes to unseen monitors, though paraphrasing can restore detection.
By Julian Schulz