arXiv AI By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang

Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

Read the original on arXiv AI →

This paper investigates safety alignment in diffusion large language models (dLLMs), which generate text via iterative denoising instead of left‑to‑right decoding. By tracking token distributions and commitment decisions across denoising steps, the authors find that refusal signals are concentrated early in the denoising process and at leading response positions, and that early committed tokens strongly influence the final safety outcome. They introduce Refusal‑Aware Early Commitment (RAEC), a training‑free decoding method that preserves early refusal signals, and demonstrate that RAEC reduces attack success rates on LLaDA and Dream while largely maintaining utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

The paper proposes a framework that explains why jailbreak attacks succeed against diffusion-based large language models (dLLMs) by viewing safety alignment as shaping a denoising energy landscape. It identifies two attack strategies—obscuring the query’s safety disposition at initialization or forcing the denoising path across an energy barrier mid‑trajectory—and introduces three training‑free detection signals that monitor initial safety disposition and kinetic energy in complementary subspaces. Experiments on several dense and sparse dLLMs show that these signals complement each other, and any attack that evades detection also fails to produce harmful content.

By Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran