Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-order and exponential extensions turn its retrieval update into softmax attention. The pol...
arXiv:2607. 19195v1 Announce Type: cross Abstract: Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories.
By Sumedha, Abhishek Singh
arXiv:2512.12767v2 Announce Type: replace-cross
Abstract: Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation...
By Thiparat Chotibut, Oleg Evnin, Weerawit Horinouchi
arXiv:2511. 02584v2 Announce Type: replace-cross Abstract: Associative memory, traditionally modeled by Hopfield networks, enables the retrieval of previously stored patterns from partial or noisy cues.
By Mark Bl\"umel, Andreas C. Schneider, Valentin Neuhaus, David A. Ehrlich, Marcel Graetz, Michael Wibral, Abdullah Makkeh, Viola Priesemann
arXiv:2605. 00366v4 Announce Type: replace-cross Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit strong storage capabilities, but the dynamical and geometric mechanisms underlying their stability remain poorly understood.
By Akira Tamamori
arXiv:2507.10383v5 Announce Type: replace-cross
Abstract: Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neu...
By Uri Cohen, M\'at\'e Lengyel
arXiv:2603. 26217v2 Announce Type: replace-cross Abstract: Generalized Hopfield models with higher-order or exponential interaction terms are known to have substantially larger storage capacities than the classical quadratic model.
By Matthias L\"owe, Franck Vermet
arXiv:2607. 26648v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven activity replaces dense multiply-accumulates with cheap accumulates.
By Zeyu Wang
The paper investigates how much learned memory is required to leverage additional data in autoregressive prediction models. It introduces a predictive‑energy spectrum that jointly governs data and memory scaling, proving a minimax law that links the number of prediction blocks and the size of the learned state to this spectrum. The authors demonstrate that optimal bit allocation and masked query‑key attention mechanisms realize this law, and they provide experimental evidence across multiple pretrained‑model scales.
By Chiwun Yang, Xiaoyu Li
Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven activity replaces dense multiply-accumulates with cheap accumulates. We argue the energy dividend of sparsity is not a property of SNNs but of the task.
arXiv:2609.16827v1 Announce Type: new
Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirica...
By Akira Tamamori
The study investigates the delayed transition from memorization to generalization—known as grokking—in two‑hidden‑layer MLPs trained on modular arithmetic. By exploring 384 hyperparameter configurations, the authors derive a power‑law scaling relation for the onset time of generalization, showing that data complexity dominates over model capacity. A clear phase boundary at weight decay around 1.0 separates grokking from non‑grokking regimes, and weight norm trajectories indicate implicit regularization during the transition.
By Anish Kataria