arXiv AI By Eliron Rahimi, Elad Hirshel, Rom Himelstein, Amit LeVi, Avi Mendelson, Chaim Baskin

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

Read the original on arXiv AI →

arXiv:2602. 02600v3 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding, competitive generation quality, and initial evidence of improved jailbreak robustness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

The paper proposes a framework that explains why jailbreak attacks succeed against diffusion-based large language models (dLLMs) by viewing safety alignment as shaping a denoising energy landscape. It identifies two attack strategies—obscuring the query’s safety disposition at initialization or forcing the denoising path across an energy barrier mid‑trajectory—and introduces three training‑free detection signals that monitor initial safety disposition and kinetic energy in complementary subspaces. Experiments on several dense and sparse dLLMs show that these signals complement each other, and any attack that evades detection also fails to produce harmful content.

By Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran
arXiv AI
Jun 4

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

arXiv:2606. 04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs.

By Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue, Yue Zhao, Chaowei Xiao