arXiv AI

Free energy landscape of Dense Associative Memory

arXiv:2607. 19195v1 Announce Type: cross Abstract: Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories.

arXiv Machine Learning
Sep 11

Phases in a class of associative memories via hidden neurons

The paper investigates associative memory in a bipartite Hopfield–Krotov architecture, termed class H, where hidden neurons serve as the retrieval order parameter. Using the replica method, it derives replica‑symmetric phase diagrams and closed‑form capacities for polynomial load, showing that crosstalk statistics are similar for Ising and spherical visible neurons. With a softmax hidden layer, the load becomes exponential, mapping the thermodynamics onto a random‑energy‑model that exhibits paramagnetic, condensed, and frozen phases, and revealing that heating destabilizes retrieval through quantized attention reassignments while Gaussian patterns remain metastable at all loads.

By Toshihiro Ota, Masato Taki
arXiv Machine Learning
Sep 14

Hierarchical Prototype Emergence in Modern Hopfield Models

The paper studies how hierarchical correlations in data can be learned by a dense Hopfield network with polynomial activation. It analytically derives conditions for each level of a hierarchical memory model to be locally stable, meaning they correspond to local energy minima. Using prototype reconstruction as a minimal generalization test, the authors show that only a quasi‑polynomial amount of information is needed to generalize beyond specific memories or groups, and they observe a similar phase diagram for Fashion‑MNIST data.

By Aditya Cowsik, Adithya Sriram
arXiv AI
Sep 3

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).

By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
arXiv AI
Aug 7

The Impossibility Triangle of Long-Context Modeling

arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).

By Yan Zhou