arXiv Machine Learning

The Entropic Signature of Class Speciation in Diffusion Models

arXiv:2602. 09651v2 Announce Type: replace-cross Abstract: Diffusion models do not recover semantic structure uniformly over time.

arXiv Machine Learning
4d ago

Breakdown of Local Denoising as Semantic Speciation

arXiv:2609.38176v1 Announce Type: new Abstract: The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class,...

By Guangkuo Liu, Mert Okyay, Yifan F. Zhang, Fangjun Hu, Rahul Nandkishore, Xun Gao
arXiv AI
Sep 3

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).

By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
Hugging Face Trending Papers
Jul 9

An exact information theory of generalization phase transitions in Bayesian diffusion models

How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data.

arXiv Computation and Language
Aug 25

DynHD: Hallucination Detection for Diffusion Large Language Models via Denoising Dynamics Deviation Learning

DynHD is a method for detecting hallucinations in diffusion large language models (D‑LLMs) by focusing on token‑level uncertainty and its evolution during the denoising process. It introduces a semantic‑aware evidence construction module that filters out non‑informative structural tokens and highlights uncertainty in informative tokens, and a reference evidence generator that models the expected trajectory of uncertainty, enabling a deviation‑based detector to identify hallucinations. Experiments show DynHD outperforms existing baselines while being more efficient across various benchmarks and backbone models.

By Yanyu Qian, Yue Tan, Yixin Liu, Wang Yu, Shirui Pan
arXiv Machine Learning
Sep 14

Diffusion Models and Concept Formation

The paper argues that diffusion models, originally developed for image synthesis, implicitly perform concept formation similar to the Cobweb cognitive model. Both models build hierarchical density structures using Gaussian prototypes, treat categorization as score‑following to reduce uncertainty, and exhibit a basic level of abstraction. The authors demonstrate this correspondence by extracting a diffusion hierarchy from MNIST and Fashion‑MNIST data and comparing its basic level to that of Cobweb, suggesting diffusion models can serve as a continuous, scalable instantiation of concept formation.

By Zekun Wang, Karthik Singaravadivelan, Christopher J. MacLellan
arXiv Machine Learning
Sep 11

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

The paper revisits the continuous diffusion language model Plaid and introduces RePlaid, aligning its architecture with modern discrete diffusion models. RePlaid achieves a compute gap of only 20× compared to autoregressive models, surpasses Duo with fewer parameters, and outperforms MDLM in over‑trained settings. On OpenWebText, RePlaid sets a new state‑of‑the‑art continuous diffusion perplexity of 22.1 and demonstrates superior generation quality, while theoretical analysis links likelihood‑based training to linear cross‑entropy over time and structured embedding geometries.

By Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sekhar Sahoo, Yongxin Chen, Arash Vahdat, Morteza Mardani, John Thickstun
arXiv AI
6d ago

Does Uniform Discrete Diffusion Need Time?

Uniform discrete diffusion models (UDMs) typically rely on explicit time conditioning, yet this study finds that such conditioning is often unnecessary in practice. While the population‑optimal UDM predictor generally depends on time—controlling how much the model should trust the observed context—the dependence becomes negligible in finite‑data language settings. Empirical results show that trained language UDMs exhibit limited time sensitivity across most of the diffusion trajectory, and time‑agnostic predictors can match or outperform time‑conditioned models on various datasets and training objectives.

By Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
arXiv Machine Learning
Sep 3

The Dynamics of Continuous Mixture Collapse in Language Models

The paper investigates why large language models (LLMs) fail to maintain continuous mixtures of token embeddings—used in latent-state reasoning—to preserve multiple reasoning paths. Through theory and experiments, it identifies three failure sources: transformer geometry distortion, amplification or contraction dynamics from softmax and autoregressive feedback, and the need for context-dependent corrections that scale with mixture size. Empirical results confirm the predicted transition between contraction and amplification and show pretrained models largely fall on the amplifying side.

By Ali Backour