arXiv:2609.38176v1 Announce Type: new
Abstract: The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class,...
By Guangkuo Liu, Mert Okyay, Yifan F. Zhang, Fangjun Hu, Rahul Nandkishore, Xun Gao
The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).
By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data.
arXiv:2607. 08041v1 Announce Type: new Abstract: How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery.
By Henry Hunt, Mason Kamb, Surya Ganguli
DynHD is a method for detecting hallucinations in diffusion large language models (D‑LLMs) by focusing on token‑level uncertainty and its evolution during the denoising process. It introduces a semantic‑aware evidence construction module that filters out non‑informative structural tokens and highlights uncertainty in informative tokens, and a reference evidence generator that models the expected trajectory of uncertainty, enabling a deviation‑based detector to identify hallucinations. Experiments show DynHD outperforms existing baselines while being more efficient across various benchmarks and backbone models.
By Yanyu Qian, Yue Tan, Yixin Liu, Wang Yu, Shirui Pan
arXiv:2606. 15327v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have demonstrated strong scaling capacity as alternatives to autoregressive language models.
By Keyue Jiang, Yuxiang Wang, Yanan Zhao, Xiang Yu, Qifang Zhao, Bohan Tang, Baojian Zhou, Yanghua Xiao, Lin Qu, Xiaoxiao Xu
arXiv:2602.06155v2 Announce Type: replace
Abstract: Diffusion models generate samples through a sequence of learned denoising steps, and recent work has studied how semantic structure appears along t...
By Kuntian Chen, Wei Wei, Yizhou Zeng, Sophie Langer, Mariia Seleznova, Hung-Hsu Chou
The paper argues that diffusion models, originally developed for image synthesis, implicitly perform concept formation similar to the Cobweb cognitive model. Both models build hierarchical density structures using Gaussian prototypes, treat categorization as score‑following to reduce uncertainty, and exhibit a basic level of abstraction. The authors demonstrate this correspondence by extracting a diffusion hierarchy from MNIST and Fashion‑MNIST data and comparing its basic level to that of Cobweb, suggesting diffusion models can serve as a continuous, scalable instantiation of concept formation.
By Zekun Wang, Karthik Singaravadivelan, Christopher J. MacLellan
The paper revisits the continuous diffusion language model Plaid and introduces RePlaid, aligning its architecture with modern discrete diffusion models. RePlaid achieves a compute gap of only 20× compared to autoregressive models, surpasses Duo with fewer parameters, and outperforms MDLM in over‑trained settings. On OpenWebText, RePlaid sets a new state‑of‑the‑art continuous diffusion perplexity of 22.1 and demonstrates superior generation quality, while theoretical analysis links likelihood‑based training to linear cross‑entropy over time and structured embedding geometries.
By Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sekhar Sahoo, Yongxin Chen, Arash Vahdat, Morteza Mardani, John Thickstun
Uniform discrete diffusion models (UDMs) typically rely on explicit time conditioning, yet this study finds that such conditioning is often unnecessary in practice. While the population‑optimal UDM predictor generally depends on time—controlling how much the model should trust the observed context—the dependence becomes negligible in finite‑data language settings. Empirical results show that trained language UDMs exhibit limited time sensitivity across most of the diffusion trajectory, and time‑agnostic predictors can match or outperform time‑conditioned models on various datasets and training objectives.
By Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
The paper investigates why large language models (LLMs) fail to maintain continuous mixtures of token embeddings—used in latent-state reasoning—to preserve multiple reasoning paths. Through theory and experiments, it identifies three failure sources: transformer geometry distortion, amplification or contraction dynamics from softmax and autoregressive feedback, and the need for context-dependent corrections that scale with mixture size. Empirical results confirm the predicted transition between contraction and amplification and show pretrained models largely fall on the amplifying side.
By Ali Backour
arXiv:2602. 02908v2 Announce Type: replace-cross Abstract: Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed.
By Binxu Wang, Jacob Zavatone-Veth, Cengiz Pehlevan