This paper investigates safety alignment in diffusion large language models (dLLMs), which generate text via iterative denoising instead of left‑to‑right decoding. By tracking token distributions and commitment decisions across denoising steps, the authors find that refusal signals are concentrated early in the denoising process and at leading response positions, and that early committed tokens strongly influence the final safety outcome. They introduce Refusal‑Aware Early Commitment (RAEC), a training‑free decoding method that preserves early refusal signals, and demonstrate that RAEC reduces attack success rates on LLaDA and Dream while largely maintaining utility.
By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
arXiv:2608. 07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood.
By Elena Dumitrescu, Gert Lek, Lydia Y. Chen, J\'er\'emie Decouchant
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
arXiv:2606. 04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs.
By Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue, Yue Zhao, Chaowei Xiao
arXiv:2606. 25182v1 Announce Type: cross Abstract: Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite safety training.
By Sofiia Nikolenko, Michele Papucci, Mina Rezaei, Shireen Kudukkil Manchingal
arXiv:2602. 02600v3 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding, competitive generation quality, and initial evidence of improved jailbreak robustness.
By Eliron Rahimi, Elad Hirshel, Rom Himelstein, Amit LeVi, Avi Mendelson, Chaim Baskin
arXiv:2608. 14392v1 Announce Type: new Abstract: Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.
By Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun
arXiv:2605. 19262v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored.
By Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou, Chengyu Huang, Pin-Yu Chen, Shengwei An
arXiv:2606. 28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood.
By Yanchen Yin, Dongqi Han, Linghui Li
arXiv:2607. 27386v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) offer a compelling alternative to autoregressive (AR) generation by enabling bidirectional context and iterative refinement.
By Saurabh Yadav, Badri Narayana Patro, Vijay Srinivas Agneeswaran
The paper investigates how fine‑tuning large language models with a small number of harmful examples can erode their refusal behavior, and explores whether localizing safety‑related behavior to specific layers or directions can provide robust defenses. Experiments across six checkpoints from four model families show that harmful and benign prompts remain linearly separable after attack, and that patching clean hidden states or freezing layers up to a transition depth can restore refusal. However, attackers can bypass these defenses by spreading updates or targeting singular directions, indicating that adaptive fine‑tuning can defeat localized repairs and highlighting the need for multiple defensive checks.
By Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim
arXiv:2510. 02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses.
By Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Hongsheng Hu, Xiaojun Jia, Di Wang, Zhan Qin, Kui Ren