Hugging Face Trending Papers

Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs

Masked diffusion language models can generate outputs in any order, but recent findings suggest this flexibility may reduce diversity by delaying uncertain tokens. The study identifies low‑confidence remasking (LCR) as the main culprit, showing that its filtering of lower‑probability tokens suppresses diversity exponentially. Replacing LCR with top‑probability position selection (TPP) restores diversity, and adding Entropy‑Guided Initialization (EGI) further enhances rollout diversity and solution coverage, demonstrating the benefits of arbitrary‑order generation for diverse outputs.

Hugging Face Trending Papers
Jun 10

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation. A notable limitation of the masked formulation, however, is that once a token has been unmasked it can no longer be revised, leaving dLLMs vulnerable to early sampling mistakes.

arXiv Machine Learning
Sep 25

Enabling Approximate Joint Sampling in Diffusion LMs

The paper introduces a lightweight single‑layer sampler that allows masked diffusion language models to approximate joint sampling of multiple tokens in a single full‑model forward pass. By training the sampler to mimic exact joint sampling from a frozen diffusion model, the authors enable parallel unmasking of tokens while maintaining a close match to the true joint distribution. Experiments on Dream‑7B and Llada‑7B models show that unmasking four tokens per denoising step yields a MAUVE score of 0.87, a substantial improvement over the marginal baseline of 0.31.

By Parikshit Bansal, Sujay Sanghavi