arXiv:2606. 27474v1 Announce Type: cross Abstract: How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding?
By Aditi Gupta, Neel Mishra, Kushagra Trivedi, Pawan Kumar
The paper introduces a new method for training data attribution in diffusion models called TID, which uses a local score discrepancy measure and can be estimated without retraining. It further distills this approach into TIDE, a forward‑only student that reproduces the teacher’s rankings using internal activations, achieving comparable accuracy at dramatically lower query cost. Experiments on CIFAR‑10, ArtBench‑10, and MS‑COCO show that TID outperforms existing methods and TIDE attributes samples in milliseconds, faster than generation itself.
By Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji
arXiv:2606. 06474v1 Announce Type: cross Abstract: Discrete diffusion language models generate text by iteratively denoising an entire response in parallel.
By Paul J\"unger, Justin Lovelace, Linxi Zhao, Dongyoung Go, Kilian Q. Weinberger
arXiv:2606. 10829v1 Announce Type: cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled.
By Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
arXiv:2606. 12232v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation.
By Stipe Frkovic, Metod Jazbec, Dan Zhang, Christian A. Naesseth, Ilija Bogunovic, Eric Nalisnick
arXiv:2606. 14620v1 Announce Type: new Abstract: Open diffusion language models are marketed as parallel, non-autoregressive decoders, yet the order in which a shipped checkpoint actually commits its tokens is almost never measured.
By Ali Asaria, Tony Salomone, Deep Gandhi
The paper introduces the first controlled benchmark for optimizers in discrete diffusion models, evaluating seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS‑M, Schedule‑Free) across four diffusion formulations: masked diffusion on text8, uniform diffusion on QM9 and LM1B, and Gaussian diffusion on CelebA‑64. Each optimizer undergoes the same search protocol and is retrained with full budget and multiple seeds, revealing that AdamW, while strong, is not universally optimal and that optimizers validated on autoregressive language models (Muon, MARS‑M, SOAP) can outperform tuned AdamW on certain tasks.
By Arman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev, Sebastian U. Stich, Maxim Panov, Eric Moulines, Peter Richt\'arik, Martin Tak\'a\v{c}
arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.
By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv:2601. 22947v2 Announce Type: replace-cross Abstract: Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive language models.
By Mengyu Ye, Keito Kudo, Ryosuke Takahashi, Jun Suzuki
The paper introduces Entropy-Valley (EV), a training‑free method for selecting target length in masked diffusion machine translation. EV evaluates candidate canvases by mean predictive entropy from all‑mask forward passes, choosing the length the model is best prepared to fill. Compared to a baseline that uses training‑corpus length statistics, EV recovers a substantial portion of the COMET‑22 gain across En→Zh, Zh→En, and En→De, and expert evaluation confirms adequacy improvements, especially for Zh→En.
By Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao
arXiv:2607. 01170v1 Announce Type: cross Abstract: Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces.
By Zhuoxuan Zhang (Yang), Kangqi Ni (Yang), Yuhang Chen (Yang), Mingfu Liang (Yang), Xiaohan Wei (Yang), Yunchen Pu (Yang), Fei Tian (Yang), Chonglin Sun (Yang), Frank Shyu (Yang), Adam (Yang), Song, Sandeep Pandey, Luke Simon, Tianlong Chen, Xi Liu
Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation. A notable limitation of the masked formulation, however, is that once a token has been unmasked it can no longer be revised, leaving dLLMs vulnerable to early sampling mistakes.