arXiv AI

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

arXiv:2602. 02600v3 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding, competitive generation quality, and initial evidence of improved jailbreak robustness.

arXiv AI
Jun 4

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

arXiv:2606. 04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs.

By Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue, Yue Zhao, Chaowei Xiao
arXiv AI
Jul 23

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

arXiv:2607. 19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates.

By Qingjia Huang, Jingyu Zhang, Jianguo Wu, Yakai Li, Weijuan Zhang, Yankai Rong, Junyi Yao, Shengzhi Zhang, Xiaoqi Jia
arXiv AI
2d ago

The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs

The paper investigates a continuation-triggered jailbreak in large language models, showing that moving an instruction suffix can markedly boost jailbreak success. By performing mechanistic interpretability at the attention‑head level, the authors reveal that the jailbreak arises from a competition between the model’s natural continuation drive and safety defenses learned during alignment. They introduce Head Competition Steering (HCS), an inference‑time technique that exploits this competition to suppress harmful outputs and distill the approach into a student model for efficient safety improvements.

By Yonghong Deng, Zhen Yang, Ping Jian, Xinyue Zhang, Zhongbin Guo, Chengzhi Li, Junxi Yin
arXiv Computation and Language
Sep 1

The Fragility of Jailbreak Robustness Across Operational States

The study shows that jailbreak robustness in language models is highly sensitive to operational-state changes. Even minor alterations to system prompts, not intended to affect safety, can dramatically shift attack success rates across seven aligned models and three jailbreak methods. The authors link these variations to changes in hidden representations along a refusal-related axis, which can predict jailbreak outcomes.

By Yuna Park, Hwang Youn Kim, Yujin Kim, Won Woo Ro, Suhyun Kim, Jae-In Hwang
arXiv AI
Jun 11

JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization

arXiv:2606. 11425v1 Announce Type: cross Abstract: Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many target queries.

By Ge Shi, Jun Yin, Donglin Xie, Fangyi Liu, Yucan Li, Menglin Liu