arXiv AI By Zhihan Yang, Wei Guo, Jean-Marie Lemercier, Simon Welker, Yonggan Fu, Mohammad Mahdi Kamani, Sajad Norouzi, Julius Berner, Tomas Geffner, Karsten Kreis, Yongxin Chen, Molei Tao, John Thickstun, Pavlo Molchanov, Ante Juki\'c, Arash Vahdat, Morteza Mardani

Large Language Continuous Diffusion Models

Read the original on arXiv AI →

The paper introduces Sigma, a large-scale continuous diffusion language model (3B/8B parameters) that uses steerable, low-dimensional ODE/SDE latent trajectories to address the non-smoothness of discrete diffusion models. Sigma is trained blockwise via likelihood optimization, jointly denoises Gaussian-corrupted token embeddings, and learns an optimal embedding geometry, leveraging pre-trained autoregressive weights for faster training. During inference, classifier-free guidance and score temperature are identified as essential for high-fidelity reasoning and coding, and Sigma matches or exceeds discrete models on benchmarks such as GSM8K, Minerva, HumanEval, MBPP, MATH-500, and AIME, while also revealing unique structural benefits like embedding-space steering and graceful degradation for low NFEs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Less Uniform Discrete Diffusion is More Powerful and Scalable

The paper introduces Less Uniform Diffusion (LUDI), a framework that improves uniform diffusion language models by using a less uniform loss and per-token time embeddings to guide reverse transitions and enable confidence-based few-step sampling. Experiments demonstrate that LUDI provides cleaner supervision, enhances few-step generation, and scales to a 7B model (LUDI-7B) that achieves a 3-token-per-step speedup over autoregressive decoding while matching masked diffusion baselines. The work suggests that UDLMs still have untapped potential for complex generation tasks.

By Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi, Haili Bai, Jiacheng Sun, Yang Xiang
arXiv AI
4d ago

Hierarchical Continuous Diffusion Language Models

Hierarchical Continuous Diffusion Language Models (HC‑DLM) combine discrete token generation with a continuous latent trajectory in a single denoising process, addressing the independence bottleneck of parallel decoding in discrete diffusion models. The training objective is derived from a variational bound on token likelihood, and the latent state is the sole persistent generative element, with tokens read out and fed back at each step. Experiments on Sudoku, Countdown, and LM1B show HC‑DLM outperforming both discrete and continuous diffusion baselines in puzzle accuracy and generative perplexity.

By Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing
arXiv Machine Learning
Jun 29

Masked Language Flow Models

arXiv:2606. 27617v1 Announce Type: cross Abstract: Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in the few-step sampling regime where parallel generation ought to provide the greatest efficiency gains.

By Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang, Simon Vary, Patrick Rebeschini
arXiv Machine Learning
Sep 25

ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks

The paper introduces ELF-REG, a scaling of continuous diffusion language models (dLMs) for reasoning tasks. By aligning representations and using a frozen autoregressive teacher to supervise intermediate denoiser features, ELF-REG achieves higher pass@1 rates on GSM8K, MATH-500, HumanEval, and MBPP compared to prior dLMs. The method also enables strong low‑NFE performance through early‑stop decoding of intermediate predictions.

By Zeyu Michael Li, William Xingxu Chen, Bingshuo Qian, Jiayin Liu, Xiang Cheng