arXiv:2609.38176v1 Announce Type: new
Abstract: The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class,...
By Guangkuo Liu, Mert Okyay, Yifan F. Zhang, Fangjun Hu, Rahul Nandkishore, Xun Gao
The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).
By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data.
arXiv:2607. 08041v1 Announce Type: new Abstract: How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery.
By Henry Hunt, Mason Kamb, Surya Ganguli
DynHD is a method for detecting hallucinations in diffusion large language models (D‑LLMs) by focusing on token‑level uncertainty and its evolution during the denoising process. It introduces a semantic‑aware evidence construction module that filters out non‑informative structural tokens and highlights uncertainty in informative tokens, and a reference evidence generator that models the expected trajectory of uncertainty, enabling a deviation‑based detector to identify hallucinations. Experiments show DynHD outperforms existing baselines while being more efficient across various benchmarks and backbone models.
By Yanyu Qian, Yue Tan, Yixin Liu, Wang Yu, Shirui Pan
arXiv:2606. 15327v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have demonstrated strong scaling capacity as alternatives to autoregressive language models.
By Keyue Jiang, Yuxiang Wang, Yanan Zhao, Xiang Yu, Qifang Zhao, Bohan Tang, Baojian Zhou, Yanghua Xiao, Lin Qu, Xiaoxiao Xu