arXiv Machine Learning

On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

The paper investigates Associative Recall (AR) in the Mamba architecture, showing that Mamba implicitly learns linear hash functions to perform recall. It identifies the low‑level circuit responsible for this behavior and develops a theoretical framework—Recall Scaling Laws—based on similarity‑preserving hashing principles. The framework predicts embedding and state dimensions for perfect recall, recall success probability, and analyzes multi‑layer and multi‑head SSM patterns, with empirical results confirming its accuracy.

arXiv AI
Aug 7

The Impossibility Triangle of Long-Context Modeling

arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).

By Yan Zhou
arXiv Machine Learning
Sep 23

Latest Exact Match Attention

arXiv:2609.25802v1 Announce Type: new Abstract: We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends onl...

By Moritz Br\"osamle
arXiv AI
Jul 16

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

arXiv:2607. 13591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks.

By Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu
arXiv AI
Jul 14

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

arXiv:2607. 09889v1 Announce Type: cross Abstract: Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache that grows with the sequence.

By Siddharth Pal, Viktoria Rojkova
arXiv Computation and Language
Sep 4

Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations

Lngram v2 introduces a latent N‑gram memory system that decouples memory routes, memory dimension, and backbone width, enabling scalable memory capacity for transformers. It employs context‑aware grouped‑query attention, a zero‑value sink, and counterfactual surrogate gradients to improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments on vision‑language models up to 30B parameters show consistent performance gains, reduced memory parameters, and stable semantic structure in the discrete IDs.

By Yunao Zheng, Bin Wen, Xiaojie Wang
arXiv AI
Aug 20

rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation

The paper introduces rEDMRec, a method that compresses a large language model’s reasoning about user preferences and item comparisons into a compact, editable memory. This memory, organized into four channels—long‑term preference, short‑term context, item perception, and counterfactual hard‑negative comparisons—can be updated by an LLM controller and queried by a lightweight student LLM for ranking, eliminating the need to re‑run the expensive teacher model for each request. Experiments on ML‑1M, Amazon Beauty, and Steam datasets show that rEDMRec consistently outperforms zero‑shot, few‑shot, RAG, and GraphRAG baselines, achieving up to a 13.3% improvement in HR@1 on ML‑1M.

By Minh Hoang Nguyen, Tung Le, Huy Tien Nguyen