The paper proposes a framework that explains why jailbreak attacks succeed against diffusion-based large language models (dLLMs) by viewing safety alignment as shaping a denoising energy landscape. It identifies two attack strategies—obscuring the query’s safety disposition at initialization or forcing the denoising path across an energy barrier mid‑trajectory—and introduces three training‑free detection signals that monitor initial safety disposition and kinetic energy in complementary subspaces. Experiments on several dense and sparse dLLMs show that these signals complement each other, and any attack that evades detection also fails to produce harmful content.
By Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran
arXiv:2606. 04236v1 Announce Type: cross Abstract: Discrete diffusion language models can generate text efficiently by updating multiple masked positions in parallel, but this parallelism introduces a quality-latency trade-off.
By Giries Abu Ayoub, Mario Barbara, Llu\'is Pastor-P\'erez, Tanja Bien, Aneesh Barthakur, Alaa Maalouf, Loay Mualem
arXiv:2607. 08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations.
By Ege \c{C}akar, Hannah Guan, Kayden Kehe
arXiv:2606. 04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens.
By Bochen Lyu, Yiyang Jia, Xiaohao Cai, Zhanxing Zhu
arXiv:2606. 22686v2 Announce Type: replace-cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque.
By Shivam Ratnakar, Kartikeya Vats
arXiv:2609.37391v1 Announce Type: new
Abstract: Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate...
By Kodai Kawamura, Kenji Kawaguchi, Anji Liu