arXiv Machine Learning
Sep 10

On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

The paper investigates Associative Recall (AR) in the Mamba architecture, showing that Mamba implicitly learns linear hash functions to perform recall. It identifies the low‑level circuit responsible for this behavior and develops a theoretical framework—Recall Scaling Laws—based on similarity‑preserving hashing principles. The framework predicts embedding and state dimensions for perfect recall, recall success probability, and analyzes multi‑layer and multi‑head SSM patterns, with empirical results confirming its accuracy.

By Yuval Koren, Assaf Ben-Kish, Raja Giryes, Lior Wolf, Itamar Zimerman
arXiv AI
Aug 7

The Impossibility Triangle of Long-Context Modeling

arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).

By Yan Zhou
arXiv Machine Learning
Sep 23

Latest Exact Match Attention

arXiv:2609.25802v1 Announce Type: new Abstract: We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends onl...

By Moritz Br\"osamle
arXiv AI
Jul 16

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

arXiv:2607. 13591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks.

By Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu