arXiv Machine Learning By Chen-Hao Chao, Wei-Fang Sun, Junwei Quan, Chun-Yi Lee, Rahul G. Krishnan

MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models

Read the original on arXiv Machine Learning →

MDM-Prime-v2 is a masked diffusion language model that improves upon the original MDM-Prime by addressing two key limitations: the high cross‑entropy loss caused by common Byte‑Pair‑Encoding tokenizers and the lack of guidance for choosing token granularity. The new design incorporates Binary Encoding and Index Shuffling, and the authors provide an analysis linking token granularity and sub‑token entropy to training objectives and downstream performance. When scaled to 1.1 B parameters, MDM-Prime‑v2 achieves higher average zero‑shot accuracy on eight commonsense reasoning benchmarks than comparable models such as GPT‑Neo, OPT, Pythia, Bloom, SMDM, and TinyLLaMA.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

arXiv:2609. 04010v1 Announce Type: new Abstract: Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation.

By Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu
arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv AI
4d ago

Less Uniform Discrete Diffusion is More Powerful and Scalable

The paper introduces Less Uniform Diffusion (LUDI), a framework that improves uniform diffusion language models by using a less uniform loss and per-token time embeddings to guide reverse transitions and enable confidence-based few-step sampling. Experiments demonstrate that LUDI provides cleaner supervision, enhances few-step generation, and scales to a 7B model (LUDI-7B) that achieves a 3-token-per-step speedup over autoregressive decoding while matching masked diffusion baselines. The work suggests that UDLMs still have untapped potential for complex generation tasks.

By Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi, Haili Bai, Jiacheng Sun, Yang Xiang