arXiv:2609.16805v1 Announce Type: new
Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit a "Ridge of Optimization" characterized by extreme stability and a...
By Akira Tamamori
arXiv:2609.10305v1 Announce Type: new
Abstract: Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transfor...
By Fang Li
arXiv:2609.16827v1 Announce Type: new
Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirica...
By Akira Tamamori
arXiv:2511. 02496v2 Announce Type: replace Abstract: We study latent geometry as an explicit component of representation quality in data-scarce learning.
By Ronald Katende
arXiv:2606. 24396v1 Announce Type: new Abstract: Large Transformer models function as Dense Associative Memories (DAMs), retrieving knowledge via high-dimensional attractor dynamics driven by the self-attention mechanism \citep{ramsauer2020hopfield, wu2024attention}.
By Kanishk Awadhiya
arXiv:2607. 07047v1 Announce Type: cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety.
By Szczepan Konior, Alexandre Quemy, Przemys{\l}aw Klocek, Gr\'egoire Cattan, Bart{\l}omiej Sobieski
arXiv:2605. 24942v2 Announce Type: replace-cross Abstract: Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation to nonlinear methods such as angular and kernelized steering, which define intervention transformations without learning an explicit geometry over paths in activation space.
By Narmeen Oozeer, Shivam Raval, Philip Quirke, Manikandan Ravikiran, Jeff Phillips, Shriyash Upadhyay, Amirali Abdullah
arXiv:2607. 03329v1 Announce Type: new Abstract: Conventional uniform convergence bounds and empirical risk minimization break down in massive over-parameterized models, such as large language transformers and biological sequence networks.
By Bing Cheng, Yi-Shuai Niu, Howell Tong, Shing-Tung Yau
arXiv:2607. 09889v1 Announce Type: cross Abstract: Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache that grows with the sequence.
By Siddharth Pal, Viktoria Rojkova
The paper investigates when language diffusion models, specifically Uniform-based Discrete Diffusion Models (UDDMs), shift from memorizing training data to generalizing to new data. It shows that UDDMs act as associative memories, forming basins of attraction around stored examples without requiring an explicit energy function. By measuring token recovery and conditional entropy, the authors identify a sharp transition governed by training set size, where memorization (vanishing entropy) gives way to generalization (finite entropy).
By Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
arXiv:2512. 12225v3 Announce Type: replace Abstract: Developing artificial agents that unify representation, memory, adaptation, and prediction remains a fundamental challenge in artificial intelligence.
By Laha Ale
arXiv:2608. 14556v1 Announce Type: new Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, material response, and anisotropic coupling must be learned from data.
By Dongzhe Zheng, Christine Allen-Blanchette