arXiv Machine Learning

Reasoning as Attractor Dynamics: Latent Memory Retrieval via Gibbs-Weighted Energy Minimization

arXiv:2606. 24543v1 Announce Type: new Abstract: Large Language Models (LLMs) are traditionally viewed as autoregressive generators.

arXiv AI
Sep 3

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).

By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
arXiv AI
Sep 15

How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition

The paper investigates how continuous latent states in large language models can store multiple reasoning steps through superposition. It challenges the intuition that retaining only the current reasoning frontier is optimal, showing that cumulative superposition of the full reasoning history can actually require fewer representational dimensions. The authors demonstrate that this approach preserves more valid evidence, improves downstream outcome discrimination, and delays unreliability, while also establishing that uniform cumulative weighting of memories is minimax‑optimal for future reasoning.

By Hongyu Gu, Chang Liu, Jingwen Fu
Hugging Face Trending Papers
5d ago

Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization

The paper investigates whether different forms of intermediate computation in large language models—such as token-based traces, pause tokens, and latent reasoning—rely on the same underlying mechanism. By training five variants of GPTNeoX on an extended multi-hop reasoning task, the authors find that while vanilla, Chain-of-Thought, and Pause Token models perform well on in-distribution data, they fail to generalize to longer-hop out-of-distribution problems. In contrast, latent-reasoning models exhibit better depth generalization, with causal analysis revealing a sparse recurrent search circuit that implements forward reachability propagation across the graph.

arXiv Computation and Language
Aug 25

DynHD: Hallucination Detection for Diffusion Large Language Models via Denoising Dynamics Deviation Learning

DynHD is a method for detecting hallucinations in diffusion large language models (D‑LLMs) by focusing on token‑level uncertainty and its evolution during the denoising process. It introduces a semantic‑aware evidence construction module that filters out non‑informative structural tokens and highlights uncertainty in informative tokens, and a reference evidence generator that models the expected trajectory of uncertainty, enabling a deviation‑based detector to identify hallucinations. Experiments show DynHD outperforms existing baselines while being more efficient across various benchmarks and backbone models.

By Yanyu Qian, Yue Tan, Yixin Liu, Wang Yu, Shirui Pan