arXiv AI

JUMP: Efficient Membership Inference on Fine-Tuned Diffusion Language Models

Hugging Face Trending Papers
Jun 10

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation. A notable limitation of the masked formulation, however, is that once a token has been unmasked it can no longer be revised, leaving dLLMs vulnerable to early sampling mistakes.

arXiv AI
4d ago

Less Uniform Discrete Diffusion is More Powerful and Scalable

The paper introduces Less Uniform Diffusion (LUDI), a framework that improves uniform diffusion language models by using a less uniform loss and per-token time embeddings to guide reverse transitions and enable confidence-based few-step sampling. Experiments demonstrate that LUDI provides cleaner supervision, enhances few-step generation, and scales to a 7B model (LUDI-7B) that achieves a 3-token-per-step speedup over autoregressive decoding while matching masked diffusion baselines. The work suggests that UDLMs still have untapped potential for complex generation tasks.

By Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi, Haili Bai, Jiacheng Sun, Yang Xiang
arXiv AI
Jun 10

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

arXiv:2606. 10829v1 Announce Type: cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled.

By Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
arXiv AI
Jun 2

SimSD: Simple Speculative Decoding in Diffusion Language Models

arXiv:2606. 02544v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding.

By Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang
arXiv AI
Jun 4

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

arXiv:2606. 04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs.

By Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue, Yue Zhao, Chaowei Xiao
arXiv Machine Learning
Sep 22

Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling

The paper introduces Independent Token Sampling (ITS), a query‑efficient method for detecting memorized training data in diffusion large language models (dLLMs). ITS selects token sets with weak internal dependency by approximating cumulative conditional mutual information using an attention‑derived pairwise dependency proxy and promotes diversity across sampling rounds. Experiments show ITS outperforms existing baselines, improving AUC by 0.18 on the ArXiv dataset while remaining effective under limited query budgets.

By Hongyao Yu, Tianqu Zhuang, Ziyuan Xu, Hao Fang, Jiaxin Hong, Bin Chen, Shu-Tao Xia
arXiv Machine Learning
Sep 25

Enabling Approximate Joint Sampling in Diffusion LMs

The paper introduces a lightweight single‑layer sampler that allows masked diffusion language models to approximate joint sampling of multiple tokens in a single full‑model forward pass. By training the sampler to mimic exact joint sampling from a frozen diffusion model, the authors enable parallel unmasking of tokens while maintaining a close match to the true joint distribution. Experiments on Dream‑7B and Llada‑7B models show that unmasking four tokens per denoising step yields a MAUVE score of 0.87, a substantial improvement over the marginal baseline of 0.31.

By Parikshit Bansal, Sujay Sanghavi